TL;DR
Anthropic says future Claude models will embed an invisible statistical watermark by using a secret key when choosing among similarly suitable words. The mark can indicate probable Claude involvement, but it cannot establish authorship and may weaken after extensive editing.
Anthropic has explained how future Claude models will watermark generated text, using a secret key to shape low-stakes word choices without adding hidden characters or extra tokens. The company is adopting the system to meet European Union AI transparency rules, giving authorized detectors a way to estimate whether Claude was involved in producing text while stopping short of proving authorship.
Claude generates text by repeatedly choosing the next word or token from a set of candidates. When several choices are similarly appropriate, the model normally uses randomness to select one. Under Anthropic’s watermarking method, a secret key and preceding words become the source of that randomness. Across a sufficiently long passage, those choices form a statistical pattern that a detector holding the key can compare with the sequence Claude would be expected to produce.
Anthropic said its implementation is a version of Google DeepMind’s SynthID-Text approach, described in a peer-reviewed 2024 Nature paper. The mark does not insert metadata, invisible spaces or hidden characters into prose. Anthropic also said it adds no billable tokens, has a negligible effect on speed and carries no information identifying a user, organization or conversation. Those performance and quality statements are company findings; Anthropic has not published an independent evaluation of its specific implementation.
Supported models will apply the mark at the model level across Claude, the Claude API, Claude Code, Claude Cowork and Claude Tag, including access through major cloud partners. Anthropic plans worldwide coverage because it said it currently lacks a durable regional restriction method. Supported PNG, JPG and SVG files will use a separate system: cryptographically signed C2PA provenance metadata attached to the file.
How Claude’s Text Watermarking Works
Future supported Claude models will use a secret key to guide low-stakes word choices, creating an invisible statistical pattern. The signal can indicate probable Claude involvement—but it cannot prove authorship.
A statistical fingerprint built word by word
Claude already selects each next token from several plausible candidates. Watermarking changes the source of randomness when those candidates are similarly suitable, allowing a hidden pattern to accumulate across a sufficiently long passage.
Generate candidates
The model identifies multiple words or tokens that fit the preceding text.
Apply the secret key
The key and preceding words determine the controlled randomness used for selection.
Build the pattern
Repeated low-stakes choices form an imperceptible statistical signature.
Compare and score
A detector holding the key estimates whether Claude was likely involved.
Invisible in the prose, useful in the aggregate
Anthropic describes its implementation as a version of Google DeepMind’s SynthID-Text approach. The published research supports the general technique, while performance claims for Claude’s specific implementation remain company findings.
Nothing is inserted
The watermark is not a visible label, an invisible space, embedded metadata, or a hidden string. It emerges from ordinary vocabulary choices.
Minimal operating cost
Anthropic says the method adds no output tokens, has a negligible effect on speed, and should create no watermark-related usage charge.
No user identity encoded
The pattern carries no personal, organizational, or conversation information. It signals possible Claude involvement, not who entered the prompt.
Copying can preserve it
Copying and pasting leaves word choices intact. Minor edits may also preserve enough of the pattern for detection.
Images use another system
Supported PNG, JPG, and SVG files will use cryptographically signed C2PA provenance metadata rather than the text watermark.
Accuracy remains undisclosed
Claude-specific thresholds, false-positive rates, false-negative rates, and independent evaluation results have not been published.
Editing changes what a detector can see
The illustration below summarizes Anthropic’s qualitative claims, not measured accuracy. Detectability depends on passage length, the proportion of Claude-selected words, and how heavily the text is transformed.
Relative signal retention
Conceptual scale based on the announced behavior; no public detection percentages are available.
“Nothing is added to the text and there are no hidden characters.”
Anthropic“A watermark can only determine that Claude was likely involved.”
Anthropic“We will soon be offering a watermark detection API.”
AnthropicWhat the result means—and what it does not
The crucial distinction is involvement versus authorship. Claude may generate, translate, rewrite, or lightly proofread material, and each workflow can leave a different amount of detectable signal.
| Question | What the watermark can indicate | What it cannot establish |
|---|---|---|
| Was Claude involved? | YESProbable involvement at some stage | Which exact workflow or prompt was used |
| Who wrote the document? | ≈Claude selected enough words to leave a pattern | Human authorship, ownership, or legal responsibility |
| Is an unmarked text human? | NONo reliable conclusion from absence alone | That AI was not used or that the mark was never present |
| Can edits remove the signal? | YESHeavy rewriting, paraphrasing, or mixing may weaken it | A universal editing threshold that guarantees removal |
| Does short text work? | ≈Only if enough statistical choices accumulate | Dependable detection from every short passage |
From generation to interpretation
EU transparency rules, global deployment
Anthropic linked the initiative to Article 50 of the EU AI Act and the transparency Code of Practice it signed in July 2026. Because regional restriction is not yet durable, supported models are planned to apply the mark worldwide.
Model-level coverage
Supported models will mark output across Claude, the Claude API, Claude Code, Claude Cowork, Claude Tag, and access through major cloud partners.
New and older models
Covered models launched in the EU from August 2, 2026 are expected to support marking. Anthropic plans to add older supported models over the coming months.
Detector documentation
Future guidance is expected to explain API access, supported platforms, thresholds, and the correct interpretation of probability-based results.
Bottom line: the announcement describes future and supported Claude models. It is not evidence that every Claude response currently carries a detectable watermark, and a detector result should never be treated as conclusive proof of authorship.
Detection Gains With Clear Limits
The system could give publishers, educators and compliance teams a provider-backed signal that is different from conventional AI-detection software, which generally searches for stylistic patterns without Anthropic’s key. Copying and pasting marked text should preserve the statistical signal, and minor edits may leave it detectable.
The distinction between involvement and authorship matters. A translation can carry a mark because Claude selects every word, while proofreading may leave little signal when most words remain human-written. A positive result means Claude was probably involved at some stage; it does not prove that Claude originated the ideas, wrote the full document or bears legal responsibility. Anthropic said ownership rights remain unchanged.
As an affiliate, we earn on qualifying purchases.
EU Rules Drive Global Marking
Anthropic tied the change to Article 50 of the EU AI Act and the related Code of Practice on Transparency of AI-Generated Content, which the company signed in July 2026 alongside other providers. The marking requirements began applying on August 2, 2026 to covered providers serving the European market.
Claude models launched in the EU on or after that date are expected to support marking from release. Anthropic said it is also adding support to models released before August 2, with deployment planned over the coming months. The announcement concerns future and supported models, not proof that every Claude response currently carries a detectable mark.
“Nothing is added to the text and there are no hidden characters.”
— Anthropic
AI generated text verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Detection Accuracy Is Still Undisclosed
Anthropic has not yet published detection thresholds, false-positive rates or false-negative rates for its Claude implementation. It also has not explained how its version differs from the published SynthID-Text method or who will receive access to the secret-key detection system.
Reliability will vary with the text. Very short passages may lack enough choices for a strong signal, while heavy editing, paraphrasing, translation or mixing with other writing can weaken or remove the mark. A missing mark cannot establish that material is human-written, and a detected mark is not conclusive proof that Claude wrote the original.
As an affiliate, we earn on qualifying purchases.
Detection API and Older Models
Anthropic plans to release a watermark detection API, publish more technical guidance and add marking to older Claude models over the coming months. Its documentation is expected to clarify detector access, supported platforms and how users should interpret probability-based results.
Source: Anthropic
As an affiliate, we earn on qualifying purchases.
Key Questions
Can readers see Claude’s text watermark?
No. Anthropic describes it as an imperceptible statistical pattern created through ordinary word choices. It contains no visible label or hidden characters.
Does the watermark identify the Claude user?
No. Anthropic said the mark contains no personal, organizational or chat information. It is designed to signal possible Claude involvement, not identify who submitted the prompt.
Can editing remove the watermark?
Light editing may leave the signal intact, according to Anthropic, while a near-complete rewrite can remove it. Detection may also fail when a passage is too short.
Does a positive result prove Claude wrote the text?
No. It indicates that Claude probably generated or processed part of the material. It cannot separate original generation from translation, rewriting or extensive editing.
Will watermarking make Claude more expensive?
Anthropic said it creates no additional output tokens and has a negligible effect on speed, so the system should cause no watermark-related usage charge.
Source: Anthropic