跳到正文
OCopenai.com·

Our approach to EU text provenance rules

AI 摘要

OpenAI is addressing EU AI Act requirements by sharing its approach to text watermarking, which is part of its broader content provenance efforts. This initiative aims to help users understand content origins and creation methods, building on existing tools for identifying AI-generated images and audio. However, editing text can significantly weaken watermarks; for instance, replacing 10% of words in 400-token passages reduced detection from 92% to 66%. OpenAI acknowledges the practical limits of current technology and plans to continue improving detection and adapting its approach as technology and regulations evolve.

为什么是这条

Unlike previous announcements, OpenAI here details the specific impact of text editing on watermark detection rates, showing a 92% to 66% drop with just 10% word replacement.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年10月5日 15:00 UTC

收录当时偏移:UTC+02026年10月5日 16:00 UTC

发布
2026年10月5日 15:00
收录
2026年10月5日 16:00
来源类型
官方发布
档位
当事方
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

讨论趋势

→ 平稳
最近 24 小时与此前 24 小时的快照均值对比 · 7 天曲线

百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。

Content provenance helps people understand where content came from, how it was created or edited, and whether it contains signals associated with our models. We’ve already made tools publicly available to identify images and audio generated by our models. Today, we’re sharing our approach to text watermarking in response to the EU AI Act, and how it fits into our broader work.

The EU AI Act requires generative AI providers to make generated text identifiable in a machine-readable way. Text watermarking and detection remain early technologies with significant limitations, and views about their benefits and responsible uses are still developing. Our phased approach reflects both the EU AI Act requirements as well as the technology’s limitations, with an emphasis on transparency about what a text watermark can and cannot tell people:

- Starting today, API customers globally will be able to opt in to text watermarking for select models. Text watermarking will remain off by default in the API.

- Over the coming weeks, we will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union.

- We’re opening applications to access our text watermark detector. Access will initially be limited to approved researchers and expert organizations that can help us evaluate and improve the technology.

- The above only applies to text provenance—our verification tools for audio and images, including our openai.com/verify ⁠ web tool and our Content Provenance API ⁠ (opens in a new window) , will continue to be publicly accessible to organizations looking to understand whether an image or audio file was generated by one of our systems.

How our watermarking works and performs

Our text watermarking technology, textGrain, adds an invisible statistical signal to the model’s word choices. Our detector looks for that signal to assess whether a passage contains an OpenAI watermark. More details about how textGrain works can be found in our technical report ⁠ (opens in a new window) , which will be updated with additional details in the coming weeks. We also plan to make the technology available in open source so that others can build on it.

In our evaluations, textGrain matched or exceeded the performance of other approaches we tested, including SynthID for text. Even so, strong performance under ideal conditions does not guarantee reliable detection in everyday use.

Detectors can make two kinds of errors: they can report a watermark where none is present—a false positive—or miss a watermark that is present—a false negative. Our evaluations below illustrate some of the challenges:

- Shorter or more constrained text is harder to detect. At a target false positive rate of 1%, our detector identified watermarks in about 80% of 200-token passages, compared with about 95% of 400-token passages, for content such as psychology. Detection rates were substantially lower for content such as mathematics, where there is less flexibility in word choice.

- Editing can weaken the watermark. In an evaluation of 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%.

These limitations contribute to our decision to provide initial detector access only to approved researchers and expert organizations, who can help us evaluate reliability and responsible uses.

This chart shows results for watermarked responses to mathematics and psychology questions from the ELI5 dataset ⁠ (opens in a new window) at a target false positive rate of 1%. Detection improves with text length, but is substantially lower overall for content where there is less flexibility in word choice, such as mathematics.

Editing can substantially weaken the watermark signal. This chart shows how replacing 10% or 25% of the words in a passage affects detection. Results are based on watermarked English responses to questions from ELI5 ⁠ (opens in a new window) .

Impact of watermarking on output quality

Across the benchmarks we use to assess Astra, our latest frontier model, we do not see meaningful performance differences with and without watermarking.

Benchmark

Unwatermarked text (Astra, max)

Watermarked text (Astra, max)

Artificial Analysis Intelligence Index

49.57 points

49.76 points

AutomationBench

来源·openai.com