跳到正文
RCreddit.com·
暂不在当前实时榜单

Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime

AI 摘要

A new system called phantom-kv has been developed to uncensor large language models without altering their weights. This system injects a small, learned bank of key/value tensors, approximately 18MB, into the model's KV cache, which the model processes as conversation history. An audit using an 8B judge-model revealed that while lexical refusal is suppressed, semantic refusal can persist as rephrasing, and the effect diminishes over long sessions with a 2-4k token half-life, requiring re-injection.

为什么是这条

Unlike other uncensoring methods that modify model weights, this system achieves refusal removal by injecting a small, trained KV-cache bank, which can be unloaded anytime.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月22日 06:01 UTC

收录
2026年9月22日 06:01
来源类型
开发者社区