返回
RCreddit.com

GPT-6 reportedly jailbroken within 24 hours using an extended Task-in-Prompt (TIP) attack [N]

OpenAI模型发布开源代码
时间与来源
发布
09/05 19:11
收录
09/06 00:00
来源类型
开发者社区
档位
社区
信源状态
正常
档位是按信源手工设定的编辑判断,不是逐条打分。
AI 摘要

一名研究人员报告称,GPT-6 Astra 在发布后24小时内就被成功越狱,他们使用了扩展的任务内提示(TIP)攻击。TIP攻击通过将有害目标隐藏在其他任务中,例如解决密码或执行Python代码,来利用模型的推理和指令遵循行为。对于GPT-6,研究人员表示,原始的最小TIP攻击已不再有效,需要进行重新设计。此消息来源于该研究人员的截图/帖子,其ACL 2025 TIP论文链接在原始帖子中。

A researcher has reported a jailbreak of GPT-6 Astra within a day after release.

The attack is described as combination of TIP (Task-in-Prompt) attack from ACL 2025 paper with four other unnamed techniques.

TIP attacks exploit the model’s reasoning/instruction-following behaviour by hidding the harmful objective inside another task, like solving a cipher or executing a Python code. For GPT-6, the researcher says the original minimal TIP attack was no longer sufficient and had to be reworked.

They have reportedly disclosed the details privately to OpenAI rather than publishing the jailbreak.

The same researcher reported jailbreaking GPT-5 within an hour of its release a year ago.

Source: screenshot/post from the researcher; their ACL 2025 TIP paper linked in the original post.