- Home
- The Thinking Wire
- We Said Text Could Not Carry a Watermark. Anthropic Shipped One Anyway
We Said Text Could Not Carry a Watermark. Anthropic Shipped One Anyway
In August 2026, Anthropic shipped statistical word-choice watermarking in all new Claude models. We had published the argument that a text watermark mandate was unenforceable. The vendor most associated with safety engineering just tested that claim in production, and the honest move is to grade our own thesis against the release. One half of it did not survive. The other half was confirmed by the mechanism itself.
The regulatory driver is explicit. Anthropic cites the EU Code of Practice on Transparency of AI-Generated Content, which has roughly 190 signatories, as the reason the feature exists. This matters for how you read everything below: the watermark was built to satisfy a transparency commitment, and the engineering choices follow from that goal rather than from an adversarial threat model.
What Actually Shipped
The mechanism biases word choices statistically, invisible to a reader, detectable by a key holder. Three properties define its scope, all as reported by Daring Fireball:
- It applies to text over 200 tokens.
- It marks text the model generated and text the model merely processed, which includes proofreading and summarization of a human draft.
- Only Anthropic holds the detection keys. No third party can run verification independently.
On the quality question, the evidence from the adjacent Google deployment is reassuring. Google’s SynthID user study measured a thumbs-up rate delta of 0.01% and a thumbs-down delta of 0.02% between watermarked and unwatermarked models. Statistical word-choice marking, at least at these scales, does not degrade output in a way users notice.
So the mark is real, it is cheap, and it rides along with normal usage. The question that decides whether any of this is useful for governance is what it takes to remove it.
The Half of Our Thesis That Died
Our original post argued that text was too malleable a medium to hold a mark at all. That claim did not survive the release.
Against casual copying, the mechanism holds up better than we predicted. The strongest single number comes from the mechanism analysis at declaude.org: text that a human paraphrases becomes re-detectable after roughly 800 tokens, around 600 words. Read that operationally. An employee who pastes Claude output into a document and rewords it by hand, sentence by sentence, still produces a detectable artifact once the passage is long enough. The statistical signal accumulates across the whole text, so local edits dilute it without erasing it.
For the honest-user population, which is most of any company, the watermark works. A student lightly editing a generated essay, a contractor reshaping a generated report, a marketer touching up generated copy: at realistic document lengths, detection is likely. Whatever we wrote about the impossibility of marking prose, the shipping product falsifies it at this tier.
The Half That Was Confirmed
The same source carries the collapse data. Under full re-composition, where the text is rewritten rather than edited, roughly 0.5% of marked windows survive, and detector accuracy falls from an AUC of 0.99 to approximately 0.5. An AUC of 0.5 is a coin flip. The detector at that point carries no information.
One disclosure before you weigh those numbers: declaude.org’s author, James Padolsey, builds a watermark-removal tool, so he has a stated commercial stake in the watermark looking weak. His figures come from MarkLLM KGW/EXP data rather than from his own product benchmarks, which is why we cite them, but the stake belongs next to the citation.
The stake itself is the second confirmation. A removal tool existed before the public detection tooling did. The economics of the arms race are visible in the launch sequence: the party who wants to strip the mark shipped first, and the parties who want to check for the mark still cannot, because only Anthropic holds the keys. Our original claim, that a text watermark is unenforceable against deliberate deception, now has the vendor’s own mechanism as supporting evidence. Anyone motivated to hide AI authorship runs one full rewrite pass and walks through the detector at coin-flip odds.
Detection Without Attribution
The scope rule deserves more attention than the robustness numbers, because it is where compliance programs will hurt their own people first.
The watermark applies to processed text, over 200 tokens, whether Claude wrote it or merely touched it. A found mark therefore means the text passed through Claude. It does not mean Claude wrote it. An employee who drafts a report from scratch and asks Claude to proofread it produces a marked document. A detector-driven policy that treats “mark found” as “AI-written” will misclassify exactly the employees who used the tool in the most defensible way.
This is the same detection-versus-provenance distinction we walked through in the PhotoDNA precedent for provenance infrastructure: a signal that something passed through a system is a different object from a claim about who authored what, and the whole failure mode lives in conflating them. The key-holder design compounds it. Because only Anthropic can run detection, you cannot audit a positive result, reproduce it, or contest it with your own tooling. Any policy consequence you attach to the mark rests on a verification step you can neither perform nor inspect.
The Control That Catches the Honest and Misses the Adversary
Put the two halves together and the shape is familiar to anyone who has run a compliance program. The watermark reliably catches the casual, honest user: the person who pasted, lightly edited, or had a draft proofread. It structurally misses the deliberate deceiver, who rewrites and drops detection to chance. The population you could already govern through policy and culture is the population the instrument sees. The population you built the instrument for is the population it cannot see.
That inversion is not a flaw specific to Anthropic’s engineering. It falls out of the physics the collapse data describes: the signal lives in word choice, and full re-composition replaces every word. Any statistical text watermark inherits the same asymmetry. Which is why the right response is to fix your policy language now, before a vendor’s watermark surfaces inside your own content pipeline through a proofreading step nobody logged.
Do This Now
Three concrete moves, none requiring new tooling.
First, grep your AI policy and any vendor contract for language that treats watermark detection as proof of AI authorship. Rewrite it to “processed by an AI system,” which is what the mark actually attests. The 200-token processed-text rule guarantees false attributions under the stricter reading.
Second, decide in advance what a positive detection triggers, and make it a conversation rather than a sanction. The mechanism’s own numbers say positives will skew toward honest users, and the verification path runs through a single vendor you cannot audit.
Third, keep the watermark out of your fraud and deception controls entirely. AUC 0.5 under rewriting means it contributes nothing against a motivated adversary, and a control that only appears to cover a risk is worse than a documented absence, because it stops you from building the control that would.
We were wrong that text could not carry a mark. We were right that the mark cannot enforce honesty. Both updates are now on the record, which is where a thesis belongs after the evidence ships.
This analysis synthesizes Anthropic’s Watermark Text Adulteration in Claude Is a Perversion of Writing (Daring Fireball, August 2026) and How AI Text Watermarking Works (declaude, August 2026).
Victorino Group helps organizations write AI-use controls that survive contact with real detection mechanics. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation