Thinking that we’ll get safety by CoT traces is wishful thinking. Safety lives in the harness, not the chain of thought
The discourse surrounding Astra's launch highlights concerns about AI safety, particularly regarding the use of recurrent depth and Chain of Thought (CoT) traces. While CoT can be useful for monitoring and post-incident analysis, it is not a faithful transcript of a model's behavior and can be manipulated by LLM providers. Researchers emphasize that readable reasoning is merely evidence, whereas the "harness" provides actual control, urging against confusing the two for ensuring safety.
为什么是这条This report uniquely argues that Chain of Thought (CoT) traces are not reliable for AI safety, unlike common assumptions, because they can be manipulated and do not fully reflect a model's internal processes.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月11日 22:00 UTC
- 收录
- 2026年9月11日 22:00
- 来源类型
- 开发者社区
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
本站未收录正文。
前往源站阅读 →