跳到正文
RCreddit.com·

I built ALHR: A tree based sparse attention system that achieves sub-quadratic inference while retaining accuracy. [P]

AI 摘要

ALHR, or Adaptive Learnable Hierarchical Routing, is a newly developed tree-based sparse attention system. It utilizes static binary trees and learnable functions to significantly reduce the number of keys that need to be read, achieving sub-quadratic inference while maintaining accuracy. This system boasts a KV Compression of 35.3x, meaning only 2.83% of keys are read. The ALHR repository is available on GitHub.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月9日 13:29 UTC

收录当时偏移:UTC+02026年10月9日 15:00 UTC

发布
2026年10月9日 13:29
收录
2026年10月9日 15:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

ALHR - Adaptive Learnable Hierarchical Routing, uses static binary trees and learnable functions to minimize the amount of keys to be read.

It does use a dense teacher while phase 1 of training however.

MQAR TEST AT 1024 TOKENS - Average keys read per query by dense - 512 Keys

Average keys read per query by ALHR - 30 keys

Top - 1 accuracy of dense - 94.9%

Top - 1 accuracy of ALHR - 92.1%

KV Compression of dense - 1x(100% read)

KV Compression of ALHR - 35.3x(2.83% read)

Peak VRAM of dense - 57 MB (Scales quadratically)

Peak VRAM of ALHR - 422 MB (scales linearly)

Cache compression of ALHR - 100%

The true log and Kaggle cell used to run it are in the logs folder in the repo

limitations: Full scale tests are still not completed, The training of this model would still be quadratic but the inference would be NlogN (as indicated in the logs in the repo)

Would love your opinions

ALHR Repository: https://github.com/vdev-ctrl/Adaptive-Learnable-Hierarchical-Routing-

来源·reddit.com