Skip to content
RCreddit.com·
Not on the current live radar

Thinking that we’ll get safety by CoT traces is wishful thinking. Safety lives in the harness, not the chain of thought

AI summary

The discourse surrounding Astra's launch highlights concerns about AI safety, particularly regarding the use of recurrent depth and Chain of Thought (CoT) traces. While CoT can be useful for monitoring and post-incident analysis, it is not a faithful transcript of a model's behavior and can be manipulated by LLM providers. Researchers emphasize that readable reasoning is merely evidence, whereas the "harness" provides actual control, urging against confusing the two for ensuring safety.

Why this one

This report uniquely argues that Chain of Thought (CoT) traces are not reliable for AI safety, unlike common assumptions, because they can be manipulated and do not fully reflect a model's internal processes.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 11, 2026, 22:00 UTC

Ingested
Sep 11, 2026, 22:00
Source type
Dev community

Discussion trend

No comparison yet
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

Article

Full text isn't available here.

Read at source →
Source·reddit.com