返回
RCreddit.com

GPT-6 reportedly jailbroken within a day of release

OpenAI模型发布开源代码
时间与来源
发布
09/05 19:21
收录
09/05 21:00
来源类型
开发者社区
档位
社区
信源状态
正常
档位是按信源手工设定的编辑判断,不是逐条打分。
AI 摘要

据报道,一名研究人员在GPT-6 Astra发布后一天内就成功对其进行了越狱。此次越狱利用了经过修改的TIP(任务诱导先验)攻击,该攻击通过将有害目标隐藏在其他任务中,例如解决密码或执行Python代码,来利用模型的推理能力。研究人员指出,最初的最小TIP攻击对GPT-6不再有效,需要进行重新设计。

A researcher has reported a jailbreak of GPT-6 Astra within a day after release.

The attack is described as combination of TIP (Task-in-Prompt) attack from ACL 2025 paper with four other unnamed techniques.

TIP attacks exploit the model’s reasoning/instruction-following behaviour by hidding the harmful objective inside another task, like solving a cipher or executing a Python code. For GPT-6, the researcher says the original minimal TIP attack was no longer sufficient and had to be reworked.

They have reportedly disclosed the details privately to OpenAI rather than publishing the jailbreak.

The same researcher reported jailbreaking GPT-5 within an hour of its release a year ago.

Source: screenshot/post from the researcher; their ACL 2025 TIP paper linked in the original post.