Promptea.
PolicyNotable

OpenAI will watermark ChatGPT text in the EU, and makes it opt-in for the API everywhere else

The textGrain method trades a measured slice of sampling randomness for a detectable signal. OpenAI's own tests show light editing weakens it, so the text detector stays gated.

Promptea Editorial4 min read

OpenAI said on Monday that it will start adding an invisible watermark to text that ChatGPT and Codex generate for users in the European Union, rolling out over the coming weeks across all plans. Separately, API customers anywhere in the world can switch the same watermark on for select models starting today. In the API it stays off by default, and OpenAI says it is not making text watermarking a global default at launch.

The trigger is regulatory. The EU AI Act's transparency obligations, which apply from August 2, 2026, require providers to mark AI-generated content in a machine-readable way. OpenAI's answer for text is a method it calls textGrain, described in a technical report co-written with researchers from the University of Pennsylvania and Yale and dated October 5.

A note on sourcing: OpenAI's blog post on openai.com refused our automated access, so the rollout details here come from the technical report and API documentation OpenAI published, plus independent reporting by TechCrunch and The Decoder, which both read the post. We could not independently confirm which API models support the opt-in or the name of the setting that turns it on.

How the watermark works

textGrain does not insert a symbol or hidden character. Each time the model picks the next token, a pseudorandom value derived from a secret key and the preceding few tokens nudges which part of the vocabulary gets sampled. Any one nudge is invisible; hundreds of them add up to a statistical pattern. A detector that holds the key can test a passage for that pattern using only the text, and the report says it does not need to know how strongly the watermark was applied.

The technical core is an explicit trade-off. Watermarking works by spending some of the randomness the model would otherwise use when it samples. The report formalizes this as an entropy budget: the operator chooses how much sampling randomness may be removed in exchange for a detectable signal, and the method is designed so that, averaged across keys, the model's token probabilities stay unchanged. The authors also address a known weakness of simpler schemes, where a fixed key can make repeated runs of the same prompt return identical answers, and say the method is compatible with speculative decoding, a common inference speed-up.

What the detector can and cannot do

OpenAI is unusually direct about the limits. According to the reporting, replacing 10% of words with synonyms cut detection on 400-token passages from about 92% to 66%, and The Decoder reports that replacing a quarter of the words brought it to about 17%. Short passages are harder: at 200 tokens The Decoder cites roughly 80% detection, with the detector tuned to a 1% false-positive rate. Math answers and translated text were also substantially harder to detect. These figures come from OpenAI's own testing; no independent evaluation is available yet.

OpenAI also frames what a positive result means narrowly. A detected watermark indicates that an OpenAI system generated or processed part of a text. It does not identify the user, measure how much a person contributed, establish ownership, or say whether the content is accurate. A negative result does not prove a human wrote it: the text may be too short, edited, translated, or produced by another company's model.

Because of those failure modes, the text detector is not public. OpenAI's API documentation says text verification is available only to approved organizations, including AI research and academic institutions, reviewed case by case. Image and audio checks, which rely on C2PA metadata and SynthID, remain available through the Content Provenance API and openai.com/verify.

How this compares with Claude

TechCrunch and The Decoder both note the contrast with Anthropic, which said about two months ago that it would watermark text generated by supported Claude models worldwide, including through its API. OpenAI's version is narrower: mandatory only in consumer ChatGPT and Codex in the EU, optional everywhere else. The Decoder also reports that OpenAI says textGrain matched or beat other methods, including Google's SynthID for text, in internal tests, found no significant benchmark differences with the watermark on, and plans to open-source the technology. None of those claims has been independently checked.

What changes for developers

For most API users, nothing changes unless they opt in. If you build products for EU users and need to show you mark AI-generated text, OpenAI now offers a vendor-side mechanism, though whether opting in satisfies your own obligations under the AI Act is a legal question, not a technical one.

The entropy-budget framing has a practical consequence worth keeping in mind. The signal lives in the choices the model makes when several tokens are plausible. Output with little freedom, such as math answers, short replies, or tightly constrained formats, carries less of it. That is consistent with OpenAI's own math results; it has not published numbers for code or structured JSON output. Anyone hoping to use the watermark to audit short, templated output should not count on it.

And the robustness numbers cut both ways. A watermark that light paraphrasing can weaken is a poor tool for enforcement, such as accusing a student or a writer, which is presumably why OpenAI is keeping the detector behind an application form for now. It has not said when access will widen.

Why this matters

  • It is OpenAI's first production text watermark, and it is shaped by EU AI Act transparency rules that apply from August 2, 2026.
  • API users get a choice: the watermark is opt-in and off by default, unlike Anthropic's approach, which reporting describes as applying to supported Claude models worldwide.
  • OpenAI's published limits make clear the detector is not a reliable way to prove who wrote a short or edited text.

Key takeaways

  • ChatGPT and Codex text generated in the EU will carry an invisible watermark, rolling out over the coming weeks.
  • API customers worldwide can opt in for select models starting October 5; it stays off by default.
  • In OpenAI's tests, replacing 10% of words with synonyms dropped detection from about 92% to 66%; short, math and translated text are harder to detect.
  • Text detection is limited to approved researchers and organizations; image and audio checks remain public.
Tags:
  • watermarking
  • content provenance
  • EU AI Act
  • API
  • ChatGPT
  • Codex
Companies:
  • OpenAI
  • Anthropic
Models:
  • ChatGPT
  • Codex

Get Promptea Weekly in your inbox

One email every Monday — the best AI stories of the week, verified and summarized.

OpenAI textGrain: EU text watermark, opt-in for the API · Promptea