RCreddit.com·
暂不在当前实时榜单
Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime
A new system called phantom-kv has been developed to uncensor large language models without altering their weights. This system injects a small, learned bank of key/value tensors, approximately 18MB, into the model's KV cache, which the model processes as conversation history. An audit using an 8B judge-model revealed that while lexical refusal is suppressed, semantic refusal can persist as rephrasing, and the effect diminishes over long sessions with a 2-4k token half-life, requiring re-injection.
Unlike other uncensoring methods that modify model weights, this system achieves refusal removal by injecting a small, trained KV-cache bank, which can be unloaded anytime.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月22日 06:01 UTC
- 收录
- 2026年9月22日 06:01
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →关联事件2 条报道 · 2 家发布者
Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime
Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime reddit.com · reddit.com
Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime reddit.com · reddit.com