Breaking the 1.58-bit Barrier for Ternary LLMs
A new method called BITCOS has been introduced to improve the storage efficiency of Ternary Large Language Models (LLMs). Ternary LLMs conventionally store weights at $\log_2 3 \approx 1.585$ bits per weight, but current methods round this up to $1.625$ bits per weight. BITCOS leverages the high zero density in ternary LLMs, achieving $1.485$ bits per weight in sparse models. This leads to up to $1.28\times$ gain in matrix-vector multiplication and up to $1.18\times$ and $1.27\times$ decode throughput improvements on CPUs and GPUs, respectively.
This report introduces BITCOS, a new method that, unlike previous five-trit packing, achieves a lower bit-per-weight storage for ternary LLMs by adapting to zero density.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 16, 2026, 22:00 UTC
- Ingested
- Sep 16, 2026, 22:00
- Source type
- Research
- Basis
- Running about 2.1× the median of this source's recent listed items
- Triggering item
- Breaking the 1.58-bit Barrier for Ternary LLMs
- Metric comparison
- 167 vs median 80.5 (20 baseline samples)
- Detected
- 09/17, 05:00
Full text isn't available here.
Read at source →