跳到正文
RCreddit.com·

[Paper] DLoop: Looped Speculative Decoding

AI 摘要

DLoop is a novel looped speculative decoding method designed to accelerate autoregressive generation in large language models. It adaptively performs multiple drafting stages before a single verification, allowing the draft model to continue proposing tokens as long as it remains confident. This approach reduces the number of target-model forward passes needed for verification. DLoop improves wall-clock speedup by 5% to 41% across various speculative decoding methods like EAGLE-3, DFlash, Domino, and DSpark, while maintaining lossless decoding. The code will be released soon.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月9日 07:18 UTC

收录当时偏移:UTC+02026年10月9日 16:00 UTC

发布
2026年10月9日 07:18
收录
2026年10月9日 16:00
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at this https URL .

来源·reddit.com