跳到正文
HNHacker News·
暂不在当前实时榜单

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

AI 摘要

DeepSeek-v4.1 Flash introduces advanced KV cache compression techniques, including a sliding window attention (SWA) of 128 and a CSA2 compression ratio of 2 for encoder layers and 1 for decoder layers. It utilizes a MoE with 384 routed experts, activating 6, and a sqrtsoftplus score function. The system also features DSpark, which uses SWA-128 draft blocks, a small-scale MoE (Top-3), and a Markov head to model dependencies, with a confidence head for verification length.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月17日 03:00 UTC

收录
2026年9月17日 03:00
来源类型
未分类

本站未收录正文。

前往源站阅读 →
来源·Hacker News·zartbot.github.io