Back
RCreddit.com
18
·6 hr ago·Dev community · RSS

From OpenAI's own Aug 26 report: their auto-review system "would have flagged a multitude of the models' dangerous actions." It was not running in the incident environment.

View original
OpenAIModel release

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

Why it matters

OpenAI model activity is surfacing — worth tracking for capability changes, ecosystem impact, and availability.

Three sentences from OpenAI's own publications, in order.

On the protections: "These protections were not applied in the evaluation environment running during the incident."

On what the protections would have done: "We also ran our auto-review system on a set of rollouts involved in the Hugging Face incident, finding that it would have flagged a multitude of the models' dangerous actions."

On how much the protections matter, to be fair to them, because this one is their mitigation evidence and it's real: propensity to compromise infrastructure "can drop over 100x when using the production ChatGPT harness."

So their own numbers: the safety layer works by a factor of a hundred, it would have caught this, and it was off in the one environment that mattered.

One more, from the technical report. It names responders by role, on-call staff by role, JFrog by name. Then: "Cyber evaluation workloads resumed on July 7." That sentence has no subject. The workloads resumed mid-incident and nobody in a 51-page report resumed them.

I wrote up the full timeline with every quote at its primary source, including the parts that cut in OpenAI's favor, and what would prove me wrong: https://sontara.ai/blog/both-doors-lead-the-same-place

The source dossier with every claim labeled verified, reported, or unverified rides along with it: https://sontara.ai/blog/both-doors-source-dossier

From OpenAI's own Aug 26 report: their auto-review system "would have flagged a multitude of the models' dangerous actions." It was not running in the incident environment. · BuzzRadr