HChuggingface.co·
暂不在当前实时榜单
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
Current safety alignment often treats harm as a topic property, using guard models like LlamaGuard-3 and benchmarks such as XSTest and OR-Bench. This can lead to models refusing safe prompts due to dangerous-looking words. Training on political refusal data, as shown with Qwen3-8B, significantly increases in-distribution political refusal from 9.47% to 84.75%. This also reduces the unsafe-response rate across HarmBench, StrongREJECT, and WildJailbreak from 26.26% to 0.14% when scored by LlamaGuard-3 in its strongest configuration.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月8日 16:00 UTC
- 收录
- 2026年9月8日 16:00
- 来源类型
- 官方发布
本站未收录正文。
前往源站阅读 →