How Does Claude Watermark AI-Generated Text? What Actually Changed
Anthropic quietly started embedding invisible watermarks into everything Claude writes. Here's the actual mechanism, straight from Anthropic's own documentation, not the shorthand version that's been circulating.
In mid-August 2026, Anthropic confirmed that Claude models now embed a machine-readable watermark into the text they generate. It's invisible, it doesn't change what Claude writes, and unless you went looking for the announcement, you'd have no way of knowing it was there. The change moved fast and generated a fair amount of confusion and pushback online, so it's worth actually explaining what's happening, rather than repeating the shorthand version that's been circulating.
Why this happened now
The trigger is regulatory, not a product decision Anthropic made independently. The EU AI Act's Article 50 requires providers of generative AI systems to mark AI-generated content in a machine-readable way. In July 2026, Anthropic joined roughly 190 other signatories, including several other major AI labs, in signing the EU's Code of Practice on Transparency of AI-Generated Content. Claude models launched on or after August 2, 2026 support this marking from day one. Models released before that date are covered by a transition period, and Anthropic has said it's working to extend watermarking to them over the following months.
Two details matter here that got lost in a lot of the initial coverage. First, this isn't unique to Claude, every major provider that signed the same Code is implementing its own version. Second, Anthropic has said it's applying the watermark globally at launch, not just to EU traffic, because it doesn't yet have a reliable way to scope the behavior by region. For a look at how those same providers actually compare on writing quality rather than policy compliance, see this breakdown of Claude, GPT, and Gemini for writing in 2026.
How the watermark actually works
The mechanism is a version of SynthID-Text, a technique Google DeepMind published in a peer-reviewed Nature paper in 2024, itself building on an approach first proposed by computer scientist Scott Aaronson in 2022. Understanding it requires a quick look at how language models generate text in the first place.
A model like Claude produces text one word at a time, and at many points in a sentence there's more than one reasonable next word. Take the sentence "the weather today was cold and...". The next word is very unlikely to be something nonsensical, but "overcast" and "grey" are both perfectly good choices, and it doesn't really matter to a reader which one gets picked. Ordinarily, a choice like that gets settled by an arbitrary random number.
Watermarking changes where that randomness comes from. Instead of an arbitrary random number, the choice is made using a cryptographic key combined with the words that came just before it. The word Claude picks is still effectively random from the reader's point of view, but the specific sequence of choices now follows a pattern that's only checkable by someone holding that key. Enough of these low-stakes choices, repeated across a long piece of text, add up to a statistically detectable signal, without changing the meaning, quality, or feel of a single sentence.
Anthropic uses a Monopoly analogy in its own explainer that's worth repeating: imagine replacing dice rolls with a sequence of digits from pi, starting at some arbitrary point. The moves are still effectively random to the players and don't change how the game plays out. But if you knew which digits of pi were used, you could later check whether a given game's sequence of moves was consistent with that source. The watermark works the same way, just applied to word choices instead of dice rolls.
Does it change what Claude writes?
According to Anthropic, no measurable difference. Google DeepMind's original SynthID-Text paper described an A/B test run on live Gemini traffic comparing thumbs-up and thumbs-down rates between watermarked and unwatermarked responses, finding no statistically significant difference. Anthropic has said it ran its own controlled study with human raters comparing watermarked and unwatermarked answers side by side, also finding no perceptible quality difference. The method also adds no extra tokens, so it doesn't cost more or run slower to serve.
What survives editing, and what doesn't
Because the watermark lives in the actual sequence of word choices rather than in hidden characters or separate metadata, it travels with the text through copy and paste, and can survive light editing. A full rewrite where every word gets replaced removes it entirely, though at that point, Anthropic itself notes, it's fair to question whether the result should still be described as AI-generated at all.
The watermark also isn't evenly distributed through a piece of text. It depends on there being a genuine choice to make. Highly factual or precise passages, where there's essentially one correct answer, leave the watermark little to work with. Code behaves the same way for the same reason, correctness usually requires an exact token, not one of several equally good options, so watermarking has less room to operate there too. Light proofreading of someone else's writing is a related edge case: if Claude only fixes a handful of words in an otherwise human-written passage, there may simply not be enough Claude-chosen words left to produce a reliable signal.
What a detected watermark can and can't tell you
This is the part most coverage has flattened into something more definitive than it actually is. A detected watermark indicates a probability that Claude was involved in producing a piece of text, not certainty, and not authorship in a legal or ownership sense. Specifically, per Anthropic's own documentation:
- It can't tell the difference between "Claude wrote this from scratch" and "Claude lightly edited someone else's writing."
- It carries no information that could identify a specific user, account, or organization.
- It becomes unreliable on very short passages, since there simply aren't enough word choices to build a confident signal.
- It says nothing about content generated by a different AI system, even one using its own watermark, since that would use an entirely different key.
- Its absence doesn't prove text wasn't AI-generated. Older models, heavy editing, translation into another form, or an unsupported product surface can all mean a genuine Claude output carries no detectable mark.
Anthropic has also said a detection tool for checking text against this watermark is coming, though as of this writing the implementation details haven't been published yet.
How this differs from tools like Pangram or GPTZero
Third-party AI detectors work on a completely different principle. They don't have Anthropic's key, so they can't check for this watermark at all. Instead, they look for statistical patterns in phrasing, the stylistic habits that tend to show up in AI-generated text more often than in human writing. Anthropic's own explainer points out a couple of examples: a fondness for constructions like "this isn't X, it's Y," and heavier-than-expected use of the word "quietly." This is the same category of pattern-matching this blog has covered before in the context of why AI writing has a recognizable feel to it, but it's a fundamentally different mechanism from cryptographic watermark detection, and it comes with the same reliability problems third-party detectors have always had: false positives on actual human writing, and false negatives on AI text that's been sufficiently edited.
How people have reacted
Coverage since the announcement has been mixed, and it's worth representing that fairly rather than only presenting Anthropic's framing. Multiple outlets reported real user frustration, particularly around the idea that lightly-edited or proofread personal writing could still carry a detectable mark, and around the fact that the policy applies with no opt-out. Anthropic's own materials are fairly direct about this tradeoff: watermarking is being applied globally, without a regional toggle, because there isn't yet a reliable way to scope it by geography, and it's framed as a transparency measure rather than a preference either individual users or organizations can turn off.
What this actually means if you write with AI
Practically, very little changes day to day. The watermark doesn't affect what Claude produces, doesn't identify you personally, and doesn't change who owns the output or who's responsible for it, Anthropic has stated all three explicitly. If you're already in the habit of substantially editing AI drafts for voice and specificity, rather than publishing the first generation unchanged, that habit was always good practice for the writing itself, and it happens to also be the scenario where a watermark has the least to work with, since heavy rewriting is exactly the case Anthropic describes as removing it.
Whatever your reasons for wanting a first draft rewritten into something that actually sounds like yours, Unzap.app is built for exactly that pass. Paste it in, pick a style, and get a version back worth publishing.
Try Unzap.appThe bigger picture
This is very new, and it's going to keep evolving. Detection tooling isn't public yet, older Claude models aren't covered yet, and other AI providers are rolling out their own versions of the same regulatory requirement on their own timelines. Worth treating anything written about this topic, this post included, as a snapshot of a specific moment rather than a settled state of the technology. If you're trying to understand what a watermark can actually prove about a piece of text, the honest answer today is: a probability, not a certainty, and a much narrower claim than most of the discourse around it suggests.