Thinking that we’ll get safety by CoT traces is wishful thinking. Safety lives in the harness, not the chain of thought
The discourse surrounding Astra's launch highlights concerns about AI safety, particularly regarding the use of recurrent depth and Chain of Thought (CoT) traces. While CoT can be useful for monitoring and post-incident analysis, it is not a faithful transcript of a model's behavior and can be manipulated by LLM providers. Researchers emphasize that readable reasoning is merely evidence, whereas the "harness" provides actual control, urging against confusing the two for ensuring safety.
Why this oneThis report uniquely argues that Chain of Thought (CoT) traces are not reliable for AI safety, unlike common assumptions, because they can be manipulated and do not fully reflect a model's internal processes.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 11, 2026, 22:00 UTC
- Ingested
- Sep 11, 2026, 22:00
- Source type
- Dev community
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
Full text isn't available here.
Read at source →