Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime
A new system called phantom-kv has been developed to uncensor large language models without altering their weights. This system injects a small, learned bank of key/value tensors, approximately 18MB, into the model's KV cache, which the model processes as conversation history. An audit using an 8B judge-model revealed that while lexical refusal is suppressed, semantic refusal can persist as rephrasing, and the effect diminishes over long sessions with a 2-4k token half-life, requiring re-injection.
Unlike other uncensoring methods that modify model weights, this system achieves refusal removal by injecting a small, trained KV-cache bank, which can be unloaded anytime.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 22, 2026, 06:01 UTC
- Ingested
- Sep 22, 2026, 06:01
- Source type
- Dev community
Full text isn't available here.
Read at source →