Finally found a model my hardware can run at full precision: me
A developer, bored with optimizing local models, created a "tps counter for fingers" using real tokenizers. They achieved a personal best of around 2 t/s, which reportedly surpasses a 70B model on a laptop CPU and is approximately 76x slower than an 8B model running on a 4090. The developer also invited others to share their results.
Unlike typical model performance benchmarks, this report offers a human-centric comparison, pitting a developer's manual token generation speed against various AI models and hardware configurations.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 28, 2026, 21:00 UTC
- Ingested
- Sep 28, 2026, 21:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →