Breaking the 1.58-bit Barrier for Ternary LLMs
A new method called BITCOS has been introduced to improve the storage efficiency of Ternary Large Language Models (LLMs). Ternary LLMs conventionally store weights at $\log_2 3 \approx 1.585$ bits per weight, but current methods round this up to $1.625$ bits per weight. BITCOS leverages the high zero density in ternary LLMs, achieving $1.485$ bits per weight in sparse models. This leads to up to $1.28\times$ gain in matrix-vector multiplication and up to $1.18\times$ and $1.27\times$ decode throughput improvements on CPUs and GPUs, respectively.
This report introduces BITCOS, a new method that, unlike previous five-trit packing, achieves a lower bit-per-weight storage for ternary LLMs by adapting to zero density.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月16日 22:00 UTC
- 收录
- 2026年9月16日 22:00
- 来源类型
- 研究
- 判定依据
- 热度约为该来源近期上榜条目中位水平的 2.1 倍
- 指标对比
- 167 vs 中位 80.5(20 条基线样本)
- 检出时间
- 09/17 05:00
本站未收录正文。
前往源站阅读 →