Skip to content
HNHacker News·
Not on the current live radar

Breaking the 1.58-bit Barrier for Ternary LLMs

AI summary

A new method called BITCOS has been introduced to improve the storage efficiency of Ternary Large Language Models (LLMs). Ternary LLMs conventionally store weights at $\log_2 3 \approx 1.585$ bits per weight, but current methods round this up to $1.625$ bits per weight. BITCOS leverages the high zero density in ternary LLMs, achieving $1.485$ bits per weight in sparse models. This leads to up to $1.28\times$ gain in matrix-vector multiplication and up to $1.18\times$ and $1.27\times$ decode throughput improvements on CPUs and GPUs, respectively.

Why this one

This report introduces BITCOS, a new method that, unlike previous five-trit packing, achieves a lower bit-per-weight storage for ternary LLMs by adapting to zero density.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 16, 2026, 22:00 UTC

Ingested
Sep 16, 2026, 22:00
Source type
Research
Breakout verdict
Basis
Running about 2.1× the median of this source's recent listed items
Metric comparison
167 vs median 80.5 (20 baseline samples)
Detected
09/17, 05:00

Full text isn't available here.

Read at source →
Source·Hacker News·arxiv.org