跳到正文
·
Archived topic · 归档话题,来源已停止追踪

GPT-Red: Unlocking Self-Improvement for Robustness

AI 摘要

GPT-Red, an AI agent, is designed to enhance the robustness, alignment, and trustworthiness of future models. An early version of GPT-Red identified "Fake Chain-of-Thought" direct prompt injection attacks, which had success rates over 95% on GPT-5.1 but are now below 10% for GPT-5.6 Sol. GPT-Red has also saturated indirect prompt injection benchmarks for developer tools and browsing with over 97% accuracy. This indicates a self-improving safety flywheel, where current models contribute to making subsequent GPT releases safer through continuous algorithmic improvements and scaling of compute and data.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年7月15日 18:33 UTC

收录
2026年7月15日 18:33
来源类型
未分类