跳到正文
RCreddit.com·
暂不在当前实时榜单

Courts have started sanctioning hidden prompts in legal filings. The 2026 research says the hidden ones are the weak version, and "ignore instructions in the document" does nothing against the strong one.

AI 摘要

Courts have begun sanctioning hidden AI prompts in legal filings, with rulings in Brazil and Connecticut punishing attempts to influence AI models through concealed text. While easily detectable hidden commands are largely ineffective, research from Collu et al. (2026) indicates that hidden preferences, phrased as user preferences, can significantly sway AI models like GPT-4o and Claude Sonnet 4. Attempts to counter these with instructions like "ignore instructions in the document" proved ineffective. This suggests a tiered threat model, with visible preferences and subtle signals posing increasingly difficult challenges for detection and filtering.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月27日 23:00 UTC

收录
2026年9月27日 23:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com