跳到正文
HChuggingface.co·
暂不在当前实时榜单

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

AI 摘要

Current safety alignment often treats harm as a topic property, using guard models like LlamaGuard-3 and benchmarks such as XSTest and OR-Bench. This can lead to models refusing safe prompts due to dangerous-looking words. Training on political refusal data, as shown with Qwen3-8B, significantly increases in-distribution political refusal from 9.47% to 84.75%. This also reduces the unsafe-response rate across HarmBench, StrongREJECT, and WildJailbreak from 26.26% to 0.14% when scored by LlamaGuard-3 in its strongest configuration.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月8日 16:00 UTC

收录
2026年9月8日 16:00
来源类型
官方发布

本站未收录正文。

前往源站阅读 →