Cover image Claude AI Watermark: 7 Holes In The Concept

  • Aug 18

On The Claude AI Watermark: 7 Holes In The Concept

Anthropic says Claude output now carries an invisible watermark, and the commentary says your writing is tracked. The clauses, the research and the removal-tool market say something more complicated. Seven holes in the concept, from a detector nobody outside Anthropic can run to cheap forgery.

Anthropic announced on 11th of August 2026 that Claude output will carry an invisible watermark, and the internet went bananas. Influential social accounts told us that Claude now tracks any writing we produce, proofread or edit with it. There wasn't much you could do about it, they said. In fact, nobody outside Anthropic could even perceive or measure the watermark itself. It may be no surprise, but a lot of the commentary was simply hype and clickbait. The facts of the matter are a little more complicated and bring with them a significant dose of perhaps.

The sections below cover the marking method, the exact legal duties and their exemptions with the clause text quoted. They also cover what the watermark can and cannot establish about a piece of text. I checked every claim against the AI Act and Anthropic's own documentation on 13 August 2026, and re-verified the removal-tool market on 18 August. Full disclosure, I used Claude models and other models such as GLM 5.2 to prepare for this article.

In This Article

  • What Anthropic Announced

  • What The Law Requires, In The Act's Own Words

  • Where Disclosure Is Required, And Where It Is Not

  • The Gap Between The Discourse And The Clauses

  • Hole 1. Detection Falls Short Of The Standard, On The Regulator's Record

  • Hole 2. Watermark Removal Is Possible, And Forgery Is Worse

  • Hole 3. Nobody Outside Anthropic Can Run The Test (At The Moment)

  • Hole 4. The Category Error, Provenance Read As Authorship

  • Hole 5. The Obligation Has No Definition Behind It

  • Hole 6. The Secret Key Is The Barrier To Interoperable Detection

  • Hole 7. The Mark Lands Where The Duties Are Lightest

  • Conclusion

What Anthropic Announced

Claude models launched on or after 2 August 2026 embed an imperceptible statistical watermark in generated text, and supported file types carry C2PA provenance metadata. The marking applies across the Claude Platform, Claude, Claude Code, Claude Cowork, Claude Tag, and cloud access through AWS, Google Cloud and Microsoft Foundry. Anthropic implements EU AI Act Article 50(2) worldwide, for every user.

C2PA is a standard from the Coalition for Content Provenance and Authenticity, with backing from Adobe, Microsoft, the BBC, Sony, Nikon and others. It attaches cryptographically signed provenance metadata to a digital file. A C2PA manifest carries assertions about the asset's creation and edits, a structured bundle of those claims, and a cryptographic hash. The hash binds the manifest to the specific bytes of the file. Any tampering with the content invalidates the manifest, and any tool that reads C2PA can verify it independently. That independence is the crucial difference from the text watermark. Anthropic applies C2PA signed metadata to supported file types including .svg, .png and .jpg, and the failure mode is structural. A file's metadata strips through format conversion, re-saving, or screenshots, because the metadata lives beside the content and travels with the file wrapper. Copy text out of any wrapper and only the in-content statistical watermark survives, which is why text needs a different mechanism entirely.

The model encodes the text signal in the sequence of word choices during generation. The mark survives copy and paste because the words themselves carry it, with no hidden characters and no metadata tag. At each generation step the model has many acceptable next words. A pseudorandom function, keyed by a secret and seeded by the words already produced, splits the vocabulary into a favoured group and an unfavoured one. Sampling then leans towards the favoured group by a small fixed amount, too small to affect readability because the alternatives were interchangeable. Detection re-derives the favoured words at each position, counts how many appear, and runs a statistical test against what chance would produce. Human text sits near chance, and watermarked text sits far above it.

There is no signature string anywhere in the document, and the mark is a statistical bias spread across hundreds of word choices. Inspection therefore cannot find it, and detection needs both the key and a decent length of text. Kirchenbauer and colleagues at the University of Maryland opened this research family in 2023, and Google's SynthID-Text is the reference implementation. Anthropic has published nothing on its own method, so the description above draws on the published research and the technical reporting.

"One false positive is enough to defeat the inference, in the same way a single black swan defeats the claim that all swans are white."

What The AI Act Requires

Article 50(2) states that "providers of AI systems, including general-purpose AI systems, generating synthetic audio, image, video or text content, shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated". The duty binds providers, meaning Anthropic, OpenAI, Google and Meta. It does not bind the person holding the text.

The exemption in the same clause is where the discourse goes wrong. The obligation "shall not apply to the extent the AI systems perform an assistive function for standard editing or do not substantially alter the input data provided by the deployer or the semantics thereof". No numeric threshold exists anywhere in the Act, and the test is semantic.

Article 50(4) covers deployers, and it requires that "deployers of an AI system that generates or manipulates text which is published with the purpose of informing the public on matters of public interest shall disclose that the text has been artificially generated or manipulated". An exemption applies where the content "has undergone a process of human review or editorial control and where a natural or legal person holds editorial responsibility for the publication of the content".

Article 3(4) defines a deployer as a person or body using an AI system under its authority, "except where the AI system is used in the course of a personal non-professional activity". So the professional-use boundary is the Act's own wording, and a person drafting their own emails in a personal capacity, it seems, falls outside the definition. Penalties for transparency breaches run to EUR15 million or 3% of global turnover under Article 99(4), and enforcement of these rules began on 2 August 2026.

Where Disclosure Is Required

For published text, the disclosure duty has two cumulative conditions. Most business and marketing writing fails the second one, so no duty arises at all. This section covers text outputs only, meaning blog articles, social posts and syndicated articles.

The publisher must be a deployer, meaning the use is professional. The text must also be published with the purpose of informing the public on matters of public interest. Neither Article 50(4) nor Recital 134 defines that phrase, and no illustration appears anywhere in the Act, which leaves the boundary to future guidance. Then the exemption switches the duty off even when both conditions hold. It applies where human review or editorial control has happened and a named natural or legal person holds editorial responsibility.

A business must disclose where it publishes AI-generated commentary on an election, a public health question or an economic policy debate straight to its blog, with no human review and nobody holding editorial responsibility. The same duty catches an automated content operation posting AI-written articles about current affairs at volume, unreviewed, on a site presented as informational.

Marketing copy, product and service pages, promotional social posts and sales newsletters carry no disclosure duty. They fail gate two because their purpose is commercial promotion. A person posting on their own social accounts in a personal capacity fails gate one. Client reports, proposals, internal documents and email never reach the public at all. A reviewed article with a named person behind it carries no disclosure duty, and a marketing post carries none either. For the overwhelming majority of business publishing, the watermark question and the disclosure question never meet.

The Gap Between Discourse And The Act

The version of the story circulating online says your text is watermarked so you are exposed. It fails on every element when set against the clauses above. The mark creates no duty for the holder of the text. Disclosure duties arise from what is published and by whom, and the mark is irrelevant to that analysis.

The proofreading case shows how little the mark says about authorship. Anthropic marks more than the law requires, since Article 50(2) exempts assistive editing, yet a document Claude merely proofread can carry the mark. A mark on a text therefore says almost nothing about who wrote it. The discourse frightens people using AI at work, and those same people are the ones least affected by the mark itself.

Hole 1. Detection falls short of the standard, on the regulator's record

Detection is probabilistic and asymmetric, so the mark cannot do the evidential work the discourse assigns to it. Detection returns a probability score, and a score is a piece of evidence to weigh. A positive result cannot distinguish author from editor, per Anthropic's own "may have been processed" wording.

False negatives are trivial, since Anthropic's page says substantial editing, paraphrasing, translation and mixing with human-written material weaken or remove the signal. Statistical detection needs roughly 100 tokens or more to return a meaningful result. Social posts, subject lines, captions and brief email replies fall below that threshold and carry no usable signal at all. Signal strength also depends on how much freedom the model had, since the mark accumulates only where genuine alternatives existed. Long prose, essays and creative writing appear to carry it well. Terse factual output, code, tables of figures and anything where the wording is largely deterministic carry it weakly or not at all.

The EU's own compliance instrument, the Code of Practice on Transparency of AI-Generated Content published on 10 June 2026 with roughly 190 signatories, records the same weakness. Forensic detection is "not deemed mature enough to comply with the quality requirements set in Article 50(2)", access to detection tools for free-form text "may be restricted to verified expert users" to compensate for "lower reliability", and text under 200 tokens "cannot be watermarked, even with a basic level of reliability". Providers sign that document to show compliance, and it carries the drafters' own record that text detection does not currently meet the standard the Act sets.

Hole 2. Watermark removal is possible, and forgery is worse

The removal-tool market that appeared within days of the announcement appears to confirm the conditional claim. An open-source project, watermarks-remover, released on 11 August, offers to strip AI provenance marks from text and files across Claude, OpenAI and Gemini. Sabrina Ramonov shipped a free no-signup web tool that cleans hidden marks from text, documents and images. Both attack the mark through heavy rewriting, and both concede the statistical rewrite is best-effort and cannot certify that a vendor detector will fail. The GitHub project separates the two jobs in its own documentation, since one layer strips metadata and invisible characters deterministically while a second layer attempts the statistical mark through model rewriting, and the repository states that no tool can certify that rewritten text fails the official check. Harel-Canada and colleagues (ACL 2025) found automated removal succeeded in 26% of attempts, falling to 10% when the output had to survive human quality review, with traces surviving hundreds of edits. Removal is cheap, it seems, when the attacker disregards quality, and hard when the result must still read well.

Forgery is arguably a more serious problem, and it is absent from all mainstream coverage that I have found. Jovanović and colleagues (ICML 2024) reverse-engineered watermarking schemes by querying an API and both forged and stripped marks for under 50 quid with over 80% success. A follow-up study of Google's SynthID-Text found it resists forgery well, at 4% success, but is easier to strip, with paraphrasing alone above 90%. Removal yields AI text that reads as human, and forgery yields human text that reads as AI-generated. If somebody can use an application to plant a fake watermark on text a human has written then the watermark tells you nothing substantial or reliable. The detector can't identify what is true and what is forged.

Hole 3. Nobody outside Anthropic can run the test (at the moment)

Only Anthropic holds the key, and nobody outside the company can run the test, reproduce a result, or examine the method, at least for the moment. You might see a future time where Anthropic licenses the detector key to institutions perhaps. The failure is Anthropic's own implementation, since Kirchenbauer and colleagues passed peer review at ICML 2023 and SynthID-Text appeared in Nature with an open-source release. Anthropic's own scheme remains undisclosed and fails the trust measure.

Reliable measurement methodologies earn their trustworthiness through a lengthy period of scrutiny by peers and by adversaries. The Kirchenbauer paper went through peer review at ICML, and independent research groups have cited, tested, attacked and extended it over three years. The SynthID-Text paper went through Nature's review process and carries an open-source release that lets anyone run the detector and verify the claims. Anthropic's scheme has been through none of that. The company has published no method, no detector, no false-positive rate, no detection threshold, and no independent evaluation. The watermark exists, but whether it is statistically reliable is a question nobody outside Anthropic can answer, and the company has not given anyone the means to ask it. When only the manufacturer asserts an instrument's reliability and nobody else can verify it, the instrument has not earned trust.

Hole 4. The category error, provenance read as authorship

Researchers built watermarking to answer a question about provenance, and the current noise online is making it a question about authorship, integrity and transparency on the part of the author. These two questions are different, and the statistical power, or absence thereof, of the test does not connect them. What AI output can and cannot establish about human work is a question I have written about before, and the watermark sits squarely in that gap.

The test never reads the sentence, does not know what words or sentences are, and it has no understanding or care for meaning. It reads token IDs and works on statistical pattern matching of these tokens. It reads the integers the model uses internally and counts which of two secretly assigned bins each one fell into. Semantic drift cannot disturb it, because the count was never semantic. The word "bad" can mean deficient or excellent depending on the sentence around it, and the tally never notices. Meaning in human language is retroactive and contextual, and the model's counting process is indifferent to all of it by design.

So the instrument is well-formed and answers a question of no human consequence. It establishes that a string of tokens passed through this particular machine, that's all. Authorship, effort, research, understanding and responsibility are outside its reach. Anthropic concedes this in its own documentation with "may have been processed by Claude", which presents a fair dose of ambiguity. The wording is the company stating that the rating does not address the question the public may be asking. In fact, it's merely a statistical probability score, not an assessment of truth. The watermarking method does not deliver on the question of truth, it's not yes or no, it's only probability.

"It establishes that a string of tokens passed through this particular machine, that's all. Authorship, effort, research, understanding and responsibility are outside its reach."

What the test tests

The structure repeats a problem familiar in the field of psychometrics. There, the measuring procedure defines the construct, and the practitioner reports the procedure's output as the thing itself. An IQ test, for example, purports to measure human intelligence. It measures timed performance on a bounded and discrete set of cognitive tasks, and then presents a calculated number as a person's intelligence. The world has many intelligent people by this measure, yet very few of them are wise or ethical. Therefore, this discrete measure is looking through a very narrow aperture. The question of whether a narrow test can capture something as complex as intelligence is one I have examined elsewhere, and it applies here with equal force.

The same problem exists for watermarking which measures correlation not proof. The field of psychometrics recognised this problem which is what construct validity and epistemic circularity attempt to capture. IQ tests are published, and that is where the watermark fares worse. A century of adversarial literature exists, and the instrument improved because attackers could reach it. The watermark is sealed, so it carries the narrowness of psychometrics and loses the transparency that made psychometrics trustworthy.

Consider a university that uses a watermark detector to screen assignment submissions. Even at a false-positive rate of three in a hundred thousand, the detector will flag some genuinely human-written submissions as AI-generated, and it cannot identify which flagged results are wrong. One false positive is enough to defeat the inference, in the same way a single black swan defeats the claim that all swans are white. The test gives you a score and nothing that tells you whether this particular score is one of the errors. Then add forgery, since anyone can plant a mark on human writing for under fifty quid. A positive result then has two available explanations that the score cannot separate.

I am not arguing the watermark won't work. The statistical signal appears on the face of things to be legitimate. The point is that a working detection still doesn't tell the reader what they want to know. It answers a question about provenance (and not very well or accurate in my humble opinion) when the reader is asking a question about authorship. What a machine can produce and what a person can produce is not as clean as the discourse assumes. Then again, it may not necessarily be about verifying authorship of a third-party for AI firms. Perhaps convenient, then, that they can simply apply the watermark across the board.

Hole 5. The obligation has no definition behind it

Article 50(2) requires marking solutions that are "effective, interoperable, robust and reliable", and the Act defines none of those four words anywhere. An obligation with no definition defies assessment, so nobody can say whether a scheme readable by a given AI company complies or fails. The Act also hedges its own requirement with the phrase "as far as this is technically feasible, taking into account the generally acknowledged state of the art". That phrase is an explicit concession that the method may not be mature enough to meet the standard it sets.

Article 3 of the Act carries 68 defined terms, and effective, interoperable, robust and reliable appear not to be among them. Neither is "watermark", "marking" as a standalone term, nor "machine-readable format". Recital 133 lists acceptable techniques as "watermarks, metadata identifications, cryptographic methods for proving provenance and authenticity of content, logging methods, fingerprints or other techniques". It asks that solutions be "sufficiently reliable, interoperable, effective and robust as far as this is technically feasible". An open-ended list of techniques and an undefined standard together amount to the Act declining to specify interoperability. That really speaks for itself and if this is the trend, I wonder how effective the Act will really be.

Hole 6. The secret key is the barrier to interoperable detection

When we consider an interoperable detector it seems that differing tokenisers across models may present a problem. A tokeniser is the system that chops text into numbered sequences, and every model family has its own with its own numbering. So a mark planted by one model may be unreadable by another even with the identical algorithm and the identical key. But that is not the real barrier, because a watermarking scheme that does not depend on tokenisers at all already exists and has passed peer review. SemStamp (NAACL 2024) marks sentence meaning and leaves token IDs alone, so it works regardless of which model produced the text. If the tokeniser were the only thing stopping interoperable detection, SemStamp would have solved it.

The secret key may be what blocks cross-provider detection. A watermark readable by everyone is arguably removable by anyone, so the security of the scheme depends on the key staying secret. Interoperable detection means letting other parties verify. The EU Code of Practice has set a date, 2 February 2027, for an interoperability solution. Its "publicly readable signpost" option is referral, and a signpost routes the query onward, which concedes that each provider's detector remains the only thing able to read its own mark. I will cede to someone with more expertise in this area but it seems that watermarking will have a problem here.

Hole 7. The mark lands where the duties are lightest

Article 50(2) of the AI Act exempts assistive editing from marking, and Article 50(4) exempts human-reviewed text from disclosure. Yet Anthropic's own help page confirms a mark can appear when Claude merely proofreads, formats or translates your own writing. The Claude watermark, therefore, has the potential for negative consequences where its involvement has been the least. I can see academics, students, and writers for example, reconsidering the level with which they use AI in their work. These professions are quite sensitive to questions of plagiarism and authenticity. That said, most academics are scared shitless of AI and arguably are the least knowledgeable amongst us in this regard.

Claude watermarking is remarkable really when you consider these AI models are themselves producing content which is a consequence of global plagiarism at an unprecedented level. By the end of 2025, plaintiffs had filed over 70 copyright infringement lawsuits against AI companies, according to the Copyright Alliance's 2025 year in review. The same review records that Anthropic alone settled for $1.5 billion after downloading 482,460 pirated books from shadow libraries to train Claude. Reuters' legal analysis of the 2025 court decisions confirms the same pattern across OpenAI, Meta, Stability AI, Google and others.

Conclusion

I use AI for research and drafting. It helps me develop a process and simplify some of them more complex work involved in writing content related to my profession. I'd love to say that it dramatically shrinks the time takes the produce an article, but quite honestly it doesn't. I've been most of the day working on this one, for example. To go from idea to finished article takes a lot of time and a lot of thinking and a lot of effort. I will say however, AI has made the process easier. I can get it to create dozens of social posts including unique social images and then schedule a whole lot for the next 30 days. I don't have to touch my social media accounts for this, just execute the workflow and the AI does the rest - brilliant!

At the same time, I don't want Anthropic's grubby little fingers all over my work. It may or may not happen depending on how much of the original text that it gave me in the outline is actually included in the final article. These sentences you're reading now, for example, I produced using voice to text. I'm speaking my idea through a software package the prints on the screen. Very handy indeed, and a lot quicker I'm typing. If I was to simply give the AI an idea and let it write everything from start to finish, I'd be like the craftsman who promises you a handmade chair and has a machine out the back spitting them out every few minutes. If an AI creates the article, then I've made myself redundant, and what's the point of that. The process that I enjoy is going from concept to finish product and I've got to stay involved the whole way through. That's the whole point of it.

Strip out the undefined terms from the Act and the responsibility outlined reads more like "mark your output in a machine-readable format so it can be detected as artificially generated, and sure you can decide yourself how that's done." Anthropic seems to have met this requirement and it won't be long before the others follow. However, if the AI Act demands for machine-readable transparency, it seems that what we have so far doesn't really stack up. The law's own interoperability wording asks for more and until Anthropic, and others for that matter, publish their mechanism, an AI watermark tells us nothing substantial. Especially when there are clever folks creating tools that can remove it, or at least bringing the robustness of the method into question.

The ones who care will continue to create the real thing. The ones that don't will use the machine. If we are too extrapolate out to its extreme then we can just have a machine live on our behalf and remove us from the equation altogether.

Sources

All sources verified 13 August 2026, with the removal-tool market re-verified 18 August 2026.

0 comments

Joinor login to leave a comment