How Anthropic plans to watermark Claude’s AI-generated text

How Anthropic plans to watermark Claude’s AI-generated text

AI

It may soon become easier to identify AI-generated content, even if it’s not the usual type of “It’s not X, it’s Y” posts found on LinkedIn and other social networks.

As you may know, the EU now requires AI companies serving its market to label their AI-generated content so that it can be more easily identified.

Anthropic and several other major AI providers have agreed to comply with the EU code of conduct. Anthropic is one of the first companies to share details about how watermarks will be implemented in Claude.

Picture

Anthropic has also confirmed that a regular user cannot see the watermark.

According to the company, it has no practical impact on the quality or content of Claude’s output, including creativity and readability.

For those who don’t know, some AI-generated images already use invisible watermarking and provenance systems, and text-based output now follows a similar concept, although the underlying implementation is different.

While the change is being rolled out to comply with EU AI law, Anthropic says the watermark will initially be applied to texts generated by Claude globally.

“We’re applying watermarks globally at launch because we don’t yet have a permanent way to segment them by region,” says Anthropic explained in a blog post.

Anthropic says future Claude models will generate watermarked text. Models released before August 2, 2026 will fall under the EU’s transition period, and Anthropic says it is working to watermark these models in the coming months.

Claude’s watermark does not add hidden characters

Anthropic says its implementation is based on Google DeepMind’s SynthID text approach and explains that it works with certain exceptions during generation.

As you may know, AI models generate text by repeatedly selecting which token could reasonably come next. Rather than adding characters after the fact or changing the finished answer, Claude’s watermark changes the source of randomness used in some of these decisions.

“Watermarking involves making simple decisions like this, which occur multiple times in a generated section of text, to leave a pattern in Claude’s responses. But this pattern is not apparent to the reader Is “It is discoverable to anyone who has a key that encodes it,” Anthropic explained.

“When watermarking is used, the decisions are still made randomly, but the source the randomness is different. Instead of using an arbitrary random number generator to select the next word, watermaking uses the key and a few words before it to determine which word the model should select.”

“That is, the words Claude chooses are still random, but now you can check the order of the words and see if it matches the choices Claude would make if he used the key. If that’s the case, you can assign a probability that the text was generated by Claude.”

I also read Research work on topic, and here is an excerpt explaining how generative watermarking works:

Generative watermarking works by carefully modifying the next token sampling process to introduce subtle, context-specific changes into the generated text distribution. Such changes introduce a statistical signature into the generated text; During the watermark detection phase, the signature can be measured to determine whether the text was actually generated by the watermarked LLM. A key advantage of the approach is that the detection process does not require computationally intensive operations or even access to the underlying LLM (which is often proprietary).

The paper goes into more detail and gives more examples, but the important part is that Anthropic doesn’t add any visible marks or hidden characters to Claude’s answer.

Google paper
Google’s research paper explains how watermarking works
Source: Google DeepMind

If Claude instead has several reasonable choices about what to generate next, the watermarking system uses a secret key and some of the preceding words as part of the randomness used to make the decision.

These individual decisions should look completely normal to the reader, but over a long enough text they leave a statistical pattern.

A detector that has the Anthropic key can examine the word order and determine how consistent it is with the choices Claude would have made when using the watermark, thereby estimating the likelihood that Claude was involved in writing the text.

According to Anthropic, internal testing found no impact on the creativity, readability, or content of Claude’s answers.

The company also states that watermarking does not require additional tokens and has a negligible impact on generation speed.

“Nothing is added to the text and there are no hidden characters,” Anthropic noted. “Adding watermarks does not require additional tokens and will not be more expensive.”

Code and factual answers may be less watermarked

As mentioned earlier, there are certain exceptions to watermarking, and for good reasons.

For factual statements where only one answer is correct, the watermark does not affect the selection, according to Anthropic.

The same principle also applies to code where replacing one term with another could affect the output.

“Where and Exactly An output is required – if there is no choice and something would be factually incorrect or a section of code would break if a different term were chosen – the watermark will not be applied.”

“For example, once the model has written “2 + 2 =”, there is a very clear best choice for the next token (when the model completes the sum, there is no answer as good as “4”; when it comes to George Orwell). Nineteen eighty-four“There is no answer as good as ‘5’,” the company noted.

“The “nudge” of the watermark would not apply here. For the same reason, code – which in many cases must be accurate – generally has fewer watermarks than some other forms of text.”

Anthropic notes that watermarks can still be used in parts of the code where there are arbitrary choices, such as comments, but says this should only have a negligible impact on the actual code produced.

This is consistent with Google’s SynthID-Text paper, which states:

There are two main factors that affect the detection performance of the scoring function. The first is the length of the text X: Longer texts contain more evidence of watermarks, giving us greater statistical confidence when making decisions. The second factor is the amount of entropy in the LLM distribution when it generates the watermarked text X. For example, if the LLM distribution has very low entropy, meaning that it almost always returns exactly the same answer to the given prompt, tournament sampling cannot select tokens that score higher G Features. In short, like other generative watermarking, tournament sampling is more powerful when the LLM distribution contains more entropy and less effective when there is less entropy.

It is also worth noting that light proofreading of human-written texts may leave too little Claude-generated material for reliable detection.

Anthropic says the watermark only applies to words Claude actually selects, so a few grammar or punctuation changes may not provide enough evidence.

Anthropic says that a translation created by Claude is watermarked because Claude selects every word in the translated output.

Anthropic is developing an API to detect Claude watermarks

It turns out there is an easier way to detect watermarks, as Anthropic plans to offer a watermark detection API.

The API will be able to estimate the likelihood that Claude was involved in writing a text, but Anthropic emphasizes that this is not the same as proving who wrote it.

A Claude watermark also cannot detect whether the text was written by another AI model, as other providers may use different watermarking methods and different keys.

“A watermark can only determine that Claude was likely involved in the content at some point. It cannot differentiate between ‘Claude wrote this’ and ‘Claude heavily edited this.’

“A light edit probably won’t completely remove the watermark; a complete rewrite, replacing every word, will.”

With small samples, detection also becomes less reliable because the detector has less choice of words to analyze.

Anthropic takes a different approach to generated PNG, JPG, and SVG files.

Claude appends cryptographically signed C2PA provenance metadata indicating that the file was created or processed with Claude, rather than modifying the file itself with an embedded watermark.


Item image

Overall prevention scores can hide what happens after the first access. Once attackers use valid credentials, prevention drops sharply.

The 2026 Blue Report measures defense technology for technology in 338 million simulations conducted in customer production environments.

Get the report

Leave a Reply

Your email address will not be published. Required fields are marked *