跳到正文
RCreddit.com·
暂不在当前实时榜单

[Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)

AI 摘要

A developer has created a 5KB pure x86-64 assembly inference engine for Gemma-2B, achieving 4.5-4.7 tokens/s in FP16 on an older quad-core i5 desktop. This project, named PULSAR-ASM, uses FASM and has zero C/C++ or PyTorch dependencies, relying only on ctypes for OS functions. It aims to explore the minimal bare-metal footprint for running an autoregressive LLM and serve as a reference for micro-LLMs on resource-constrained microcontrollers.

为什么是这条

This project offers a unique exploration into the minimal bare-metal footprint for LLMs, unlike other inference engines that rely on larger runtimes or compiler abstractions.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月4日 07:00 UTC

收录
2026年10月4日 07:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com