AI Text Watermarks: Checkmate for Cheaters? What Anthropic’s New Labelling Commitment Can Do – and What It Can’t
Anthropic has committed to embedding an invisible watermark in text generated by Claude, a response to the EU labelling obligation in force since 2 August 2026. The retrofit must be in place by 2 December. But what does such a watermark actually prove, and where do the competitors stand?
The key facts in brief
- Anthropic has signed the EU transparency framework for AI-generated content: Claude models released in the EU from 2 August 2026 onwards are to carry invisible watermarks in text and signed provenance data in files.
- Every Claude model currently on offer predates that cut-off date and runs without an active watermark; for systems placed on the market before the cut-off, a statutory deadline runs until 2 December 2026.
- Among the three large providers, only Google states that it already marks generated text today; OpenAI and Anthropic deny doing so for ChatGPT and Claude.
What exactly has Anthropic announced?
In early August 2026, Anthropic described in its help centre how the company will label what Claude produces. The basis is Article 50(2) of the EU AI Act. It obliges providers of generative AI systems to mark their outputs as artificially generated in a machine-readable and detectable form. The obligation has applied since 2 August 2026. To meet it, Anthropic has joined the Code of Practice on Transparency of AI-Generated Content, a voluntary framework from the EU Commission. It shows companies how to satisfy the obligation with legal certainty.
By the end of July 2026, according to the EU Commission, around 190 organisations had signed up, among them Anthropic, Google, Microsoft, OpenAI, Meta, Mistral, Cohere, Aleph Alpha and Black Forest Labs. Anthropic’s pledge in detail: Claude models released in the EU from 2 August 2026 onwards are to embed a watermark in generated text. File formats such as SVG, PNG and JPG additionally receive signed provenance data under the open C2PA standard. The Coalition for Content Provenance and Authenticity uses it to document across the industry where a digital item came from.
According to Anthropic, the labelling is to apply worldwide, not only to users in the EU, and to cover every access route: the Claude platform (API), claude.ai, Claude Code, Claude Cowork, Claude Tag and the cloud partners AWS, Google Cloud and Microsoft Foundry.
By when does the watermark have to arrive?
For older systems the law grants a transition period. Anthropic calls the retrofit of its models released before the cut-off ongoing work, without naming a date. The law names one.
Sonnet 5, the default model for the free and Pro tiers, appeared on 30 June 2026, the more capable Opus 5 on 24 July 2026 – both before the cut-off. Anthropic’s own release notes show that no new model has followed Opus 5. The later entries concern platform features and the retirement of the older Opus 4.1. No Claude model currently on offer therefore carries a confirmed active watermark.
Anthropic does not have unlimited time for this. The Digital Omnibus, an amending package to the AI Act, grants systems already on the market a four-month grace period. According to the EU Commission, their providers must have implemented the labelling by 2 December 2026. For systems entering the market from 2 August onwards, the obligation applies immediately.
One ambiguity remains. The law attaches to AI systems placed on the market before the cut-off. Anthropic, by contrast, describes its pledge along the lines of individual models. Whether a new model that moves into an existing system such as claude.ai or the API after 2 August counts as newly placed on the market is unresolved. A second deadline follows in any case: by 2 February 2027, signatories of the Code of Practice must make detection of their watermarks interoperable with one another, for instance through a shared query interface or a note readable within the content indicating which detector is responsible.
How does the principle work – and what remains open?
Anthropic does not reveal how the watermark works technically. The help centre states only that the method weaves a watermark imperceptible to humans directly into the text, without altering meaning, quality or readability. Because it forms part of the text itself, it travels along when the text is copied and survives some editing. The principle becomes more concrete at another provider.
For its watermarking technology SynthID, Google publicly describes how a language model works as it writes. It does not choose words at random but calculates a probability for every possible next token. A token is the smallest unit of text into which a language model breaks down language as it generates. SynthID shifts these probabilities in a consistent pattern that no reader notices.
A detector later demonstrates this pattern statistically. Decisive for the whole debate: marking and detecting are two separate steps. The watermark arises as the model writes. Whether anyone can check it later, and who, is decided by the provider through a separate tool. The Code of Practice explicitly permits this separation. That Anthropic works on the same principle is plausible, but not confirmed. The company has disclosed neither its own mechanism in detail nor tools with which users or third parties could check a watermark themselves. Both are said to be in preparation, with technical documentation to follow.
A detected watermark proves only that a text may have been processed by Claude. It does not prove that Claude wrote it
Will AI texts be easier to spot in future?
Even if the technology works as announced, limits remain. Anthropic names them in the same document. A detected watermark proves only that a text may have been processed by Claude. It does not prove that Claude wrote it. Anyone who has a text of their own merely edited, translated or summarised also receives marked text. The thinking behind it still comes from a human.
The reverse also holds: the absence of a detectable watermark does not necessarily mean the text came from a human. Anthropic names several cases in which the marking effectively disappears. These include heavy editing, translation and very short passages. With files, the metadata is lost through format conversion, re-saving or screenshots.
The Code of Practice requires providers in principle to use two layers of marking: signed metadata and a watermark. Plain text cannot carry metadata – a raw character string without a file format, on a web page or in a chat window, for example. For it, the Code declares a single layer sufficient, namely the watermark. In the most widely discussed case of all, copying out of the chat, only the mechanism applies that the document itself rates as the less reliable one.
There is also a lower threshold. Texts below 200 tokens, roughly 150 words, count as very short in the Code and cannot be marked even with basic reliability at the current state of the art. Above that, providers must mark. But the Code notes that this becomes reliable only with very long texts. Social networks recently carried the question of whether such labelling exposes AI-written coursework. The answer is neither a plain yes nor a plain no.
Anyone who thoroughly rewrites an AI text will most likely push it below the detection threshold. And even those who do not will not be checked by everyone. The Code allows providers to restrict watermark detection in plain text to verified professional users, because the results can be misleading. The document counts among that group public authorities, law enforcement, media, fact-checkers and educational and research institutions. Checking is free of charge in principle; only providers with fewer than one million users per month may levy fees at very high request volumes, and not even then towards the groups named. Universities are therefore included, the general public is not.
Watermarks at ChatGPT, Gemini & co.: who stands where?
Anthropic is neither alone with the announcement nor out in front. Google is furthest along, at least by its own account. According to Google DeepMind, SynthID marks text created in the Gemini app and the web interface. Google is the only one of the three providers to state that it marks text at all – and the only one to explain the underlying principle publicly.
The lead has two gaps. The first concerns coverage: by Google’s description, the text watermark covers the Gemini app and the web interface, not the API through which developers embed Gemini in their own applications. That statement, however, dates back to DeepMind’s 2024 announcement. Whether it reflects the current position, Google does not say.
The second gap concerns detection. Google runs a verification portal, the SynthID Detector, to which, by its own description, image, video and audio files can be uploaded; Google does not list text there. The portal is also in a trial phase, to which journalists, media professionals and researchers gain access via a waiting list. Unlike with image, video and audio, nobody outside the company can confirm that Gemini text actually carries a watermark.
OpenAI starts with image and sound, not text. The company has attached C2PA provenance data to generated images since 2024 and extended that approach in May 2026 with SynthID – for images from ChatGPT, Codex and the API. On 31 July 2026, two days before the EU obligation took effect, audio followed, according to OpenAI’s own account. The public verification tool that OpenAI first presented in May for images has since also recognised audio files; an interface for automated provenance checks was added.
For text the labelling is missing, and for a stated reason. In an August 2024 addendum to










