Skip to content
RCreddit.com·
Not on the current live radar

[Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)

AI summary

A developer has created a 5KB pure x86-64 assembly inference engine for Gemma-2B, achieving 4.5-4.7 tokens/s in FP16 on an older quad-core i5 desktop. This project, named PULSAR-ASM, uses FASM and has zero C/C++ or PyTorch dependencies, relying only on ctypes for OS functions. It aims to explore the minimal bare-metal footprint for running an autoregressive LLM and serve as a reference for micro-LLMs on resource-constrained microcontrollers.

Why this one

This project offers a unique exploration into the minimal bare-metal footprint for LLMs, unlike other inference engines that rely on larger runtimes or compiler abstractions.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 4, 2026, 07:00 UTC

Ingested
Oct 4, 2026, 07:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com