RCreddit.com·
暂不在当前实时榜单
[Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)
A developer has created a 5KB pure x86-64 assembly inference engine for Gemma-2B, achieving 4.5-4.7 tokens/s in FP16 on an older quad-core i5 desktop. This project, named PULSAR-ASM, uses FASM and has zero C/C++ or PyTorch dependencies, relying only on ctypes for OS functions. It aims to explore the minimal bare-metal footprint for running an autoregressive LLM and serve as a reference for micro-LLMs on resource-constrained microcontrollers.
This project offers a unique exploration into the minimal bare-metal footprint for LLMs, unlike other inference engines that rely on larger runtimes or compiler abstractions.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年10月4日 07:00 UTC
- 收录
- 2026年10月4日 07:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →