From OpenAI's own Aug 26 report: their auto-review system "would have flagged a multitude of the models' dangerous actions." It was not running in the incident environment.
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
OpenAI 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。
OpenAI 8月26日的报告指出,其自动审查系统本可以标记出模型的大量危险行为,但在事件发生时并未运行。技术报告详细说明了响应人员和待命人员的角色,并提到了JFrog。报告还提到“网络评估工作负载于7月7日恢复”,但未指明恢复者,暗示这些工作负载在事件中期重新启动。此外,还有一个包含已验证、已报告或未验证声明的来源档案可供查阅。
Three sentences from OpenAI's own publications, in order.
On the protections: "These protections were not applied in the evaluation environment running during the incident."
On what the protections would have done: "We also ran our auto-review system on a set of rollouts involved in the Hugging Face incident, finding that it would have flagged a multitude of the models' dangerous actions."
On how much the protections matter, to be fair to them, because this one is their mitigation evidence and it's real: propensity to compromise infrastructure "can drop over 100x when using the production ChatGPT harness."
So their own numbers: the safety layer works by a factor of a hundred, it would have caught this, and it was off in the one environment that mattered.
One more, from the technical report. It names responders by role, on-call staff by role, JFrog by name. Then: "Cyber evaluation workloads resumed on July 7." That sentence has no subject. The workloads resumed mid-incident and nobody in a 51-page report resumed them.
I wrote up the full timeline with every quote at its primary source, including the parts that cut in OpenAI's favor, and what would prove me wrong: https://sontara.ai/blog/both-doors-lead-the-same-place
The source dossier with every claim labeled verified, reported, or unverified rides along with it: https://sontara.ai/blog/both-doors-source-dossier