·
Archived topic · 归档话题,来源已停止追踪
GPT-Red: Unlocking Self-Improvement for Robustness
GPT-Red, an AI agent, is designed to enhance the robustness, alignment, and trustworthiness of future models. An early version of GPT-Red identified "Fake Chain-of-Thought" direct prompt injection attacks, which had success rates over 95% on GPT-5.1 but are now below 10% for GPT-5.6 Sol. GPT-Red has also saturated indirect prompt injection benchmarks for developer tools and browsing with over 97% accuracy. This indicates a self-improving safety flywheel, where current models contribute to making subsequent GPT releases safer through continuous algorithmic improvements and scaling of compute and data.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年7月15日 18:33 UTC
- 收录
- 2026年7月15日 18:33
- 来源类型
- 未分类