[Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)
A developer has created a 5KB pure x86-64 assembly inference engine for Gemma-2B, achieving 4.5-4.7 tokens/s in FP16 on an older quad-core i5 desktop. This project, named PULSAR-ASM, uses FASM and has zero C/C++ or PyTorch dependencies, relying only on ctypes for OS functions. It aims to explore the minimal bare-metal footprint for running an autoregressive LLM and serve as a reference for micro-LLMs on resource-constrained microcontrollers.
This project offers a unique exploration into the minimal bare-metal footprint for LLMs, unlike other inference engines that rely on larger runtimes or compiler abstractions.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 4, 2026, 07:00 UTC
- Ingested
- Oct 4, 2026, 07:00
- Source type
- Dev community
Full text isn't available here.
Read at source →