跳到正文
RCreddit.com·

CoWindow and MassAlloc Attention: collective causal coverage and distribution-adaptive compute [R]

AI 摘要

The authors of CoWindow (CoWA) and MassAlloc (MALA) Attention presented their research on reducing redundant computation in attention mechanisms. Their evaluations on models scaling from 0.6B to 14B, and continued-training experiments at 32B, showed significant FLOPs reduction. Specifically, CoWA decreased total training FLOPs by 28.5% and MALA by 23.1% at 14B with 32K context, while maintaining comparable model capabilities to FullAttn.

为什么是这条

This report details the specific FLOPs reduction percentages for CoWA (28.5%) and MALA (23.1%) at 14B with 32K context, unlike general claims of efficiency gains.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年9月29日 05:16 UTC

收录当时偏移:UTC+02026年9月29日 08:00 UTC

发布
2026年9月29日 05:16
收录
2026年9月29日 08:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

讨论趋势

暂无对比
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

I'm one of the authors of two recent papers exploring different sources of redundant computation in attention. I'd like to share the ideas and hear feedback from people working on long-context models and attention kernels.

CoWindow Attention (CoWA) distributes distant context across KV heads using complementary windows, while sharing local and prefix-sink windows. Each head attends sparsely, but the union of their visible positions covers the full causal history. The pattern is position-defined and requires no learned router or indexer.

Paper: https://arxiv.org/abs/2609.32704

MassAlloc Attention (MALA) retains full causal QK scoring, then uses attention's own softmax statistics to decide whether to execute subsequent computation for a tile. It reduces low-contribution post-score work, using a common tolerance across training and inference.

Paper: https://arxiv.org/abs/2609.32712

Both support training forward/backward and inference prefill/decoding. At 128K tokens on 8 H100 GPUs with TP=8, attention-operator speedups relative to FullAttn are:

Method Forward Backward Decode CoWA 7.4x 8.6x 3.0x MALA 2.2x 3.0x 1.6x These measurements are for the attention operators, not end-to-end model speedups.

We evaluated scaling from 0.6B to 14B and conducted separate continued-training experiments at 32B. At 14B with 32K context, total training FLOPs decreased by 28.5% for CoWA and 23.1% for MALA, with model capabilities comparable to FullAttn on the reported evaluations.

Two distinctions that matter: collective coverage does not imply identical head-wise interactions or outputs to FullAttn, and MALA still pays for full causal QK scoring. Neither result establishes universal lossless equivalence to dense attention.

I'd be interested in feedback on workloads that might stress collective coverage, or attention distributions where adaptive post-score allocation could be less effective. Happy to discuss implementation and evaluation details.

来源·reddit.com