Mastodon Skip to content
LIVE - NYSE/-/- CRYPTO/OPEN/24/7
BTC$85,938▼ 0.90%ETH$2,716▼ 0.35%SOL$120.93▼ 0.15%TOTAL CRYPTO$2.94T▼ 2.51%S&P 5007,773.95▲ 0.66%NASDAQ27,477.31▲ 1.05%DOW51,267.90▲ 0.18%GOLD4,173.40▲ 0.27%WTI89.60▼ 1.66%BRENT100.55▼ 1.66%EUR/USD1.1231▼ 0.21%USD/JPY157.84▲ 0.06%DXY102.10▲ 0.16%
AI

OpenAI Watermarks ChatGPT Text In EU Compliance Push

OpenAI rolled out invisible text watermarking in the EU to meet AI Act rules, with detection rates that fall sharply on edited text, its own numbers show.

Pexels – Andrew Neel

OpenAI started adding invisible watermarks to ChatGPT and Codex text output in the European Union this week, the company’s first response to EU AI Act rules that require machine-readable identification of generated text. The announcement from OpenAI landed on Oct. 5, alongside a technical report on the watermarking system called textGrain, which adds a statistical signal to the model’s word choices that a detector then looks for. OpenAI said text watermarking remains an early technology with significant limitations, and it did not hide those limitations in the numbers it published.

The rollout works in three parts. Starting Monday, API customers worldwide can opt in to text watermarking for some models, staying off by default in the API itself. Over the coming weeks OpenAI will add an invisible watermark to eligible ChatGPT and Codex output in the EU across all plans, regionally rather than globally at launch. Third, applications are open for access to the watermark detector, limited at first to approved researchers and expert organizations, since the company says it wants outside scrutiny before wider use.

Detection rates hold up only under ideal conditions

The company published detection numbers for the first time. At a target false positive rate of 1%, the detector identified watermarks in about 80% of 200-token passages, compared with about 95% of 400-token passages, for content in categories like psychology. Detection rates were substantially lower for content such as mathematics, where word choice leaves less statistical room for a signal to ride on.

Editing weakens the signal further. In an evaluation of 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%. On quality, OpenAI said benchmark scores for its Astra frontier model were essentially unchanged with and without watermarking, and a table in the announcement shows a 49.57 to 49.76 point spread on the Artificial Analysis Intelligence Index, so the signal costs almost nothing in model performance.

What a watermark cannot tell you

The company laid out a list of what a watermark does not establish, a list longer than the detection rates. A watermark does not measure human contribution, so it cannot separate a paragraph a model generated from one a person wrote. It does not establish ownership or legal responsibility, does not tell you whether a passage is accurate or misleading, and does not identify the user who generated it. And OpenAI said that the absence of a detected watermark does not prove human authorship, since text generated with OpenAI tools can be too short, edited, translated, or produced by an unsupported model or a rival system that carries no watermark of its own.

“The watermark can indicate that an OpenAI system generated or processed part of a text. It cannot tell who used the system, how much a person contributed, who owns the text, or whether the content is accurate.”

The release notes summary of the announcement adds that text watermarking stays off by default in the API, so businesses can choose how watermarking fits their products. OpenAI also said it plans to release the watermarking technology in open source so others can build on it, a step that could let rival detection tools emerge rather than stay locked to one company’s criteria.

EU rules made this necessary, not a market pull

The EU AI Act requires generative AI providers to make generated text identifiable in a machine-readable way. OpenAI’s response highlights the technical hard part of that requirement rather than just the compliance step. Machine detection of text is a statistical problem, not a metadata problem, because anyone who copies the text and rewrites parts of it or translates it can break the signal entirely. That is different from image and audio provenance, where watermark insertion and C2PA Content Credentials are harder to strip out because the file itself carries more room for embedding.

What other providers do

Google’s SynthID, the tech it already applies to images and audio, is a reference point for the space, and OpenAI said in this announcement that textGrain matched or exceeded the performance of other approaches it tested, including SynthID for text. A content provenance API is already publicly accessible for checking whether an image or audio file came from one of OpenAI’s systems, and that part of the toolkit remains open to organizations rather than restricted to researchers. Text detection stays narrower at launch, with no public detector endpoint announced for text.

Why so much hedging

OpenAI framed the whole rollout as limited in scope on purpose. “Text watermarking and detection remain early technologies with significant limitations, and views about their benefits and responsible uses are still developing,” the company said in the announcement. That framing matches what the detection numbers show: high reliability only on longer, unedited text in categories where word choice is flexible. For a short email, a translated paragraph, or a math solution, the system cannot reliably say whether a model wrote it, and OpenAI has not claimed otherwise.

Regulators will read this differently. The EU AI Act’s requirement is that generated text be machine-identifiable, which the company now meets in a narrow way on text that has not been edited. Compliance and utility are two different bars here, and OpenAI’s data does not claim the second. Neither does anyone else yet, and the detector access restriction to vetted researchers means it will be some time before the broader ecosystem gets a clear read on how well this works outside the lab. The open source plan, if it lands, is the piece most likely to settle that debate, since independent implementations can test robustness without asking the company’s permission.

SourcesOpenAI announcement on EU text provenance (Oct. 5, 2026); textGrain technical report from OpenAI; ELI5 dataset evaluations cited by OpenAI; release notes compiled by Releasebot.
Share: X