Skip to content
HNHacker News·
Not on the current live radar

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

AI summary

DeepSeek-v4.1 Flash introduces advanced KV cache compression techniques, including a sliding window attention (SWA) of 128 and a CSA2 compression ratio of 2 for encoder layers and 1 for decoder layers. It utilizes a MoE with 384 routed experts, activating 6, and a sqrtsoftplus score function. The system also features DSpark, which uses SWA-128 draft blocks, a small-scale MoE (Top-3), and a Markov head to model dependencies, with a confidence head for verification length.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 17, 2026, 03:00 UTC

Ingested
Sep 17, 2026, 03:00
Source type
Unclassified

Full text isn't available here.

Read at source →
Source·Hacker News·zartbot.github.io