Claude is getting more ambitious with the watermarks and I can smell trouble a mile away

Claude is getting more ambitious with the watermarks and I can smell trouble a mile away

Anthropic wants to make AI-generated text easier to identify, and on paper I have very little to complain about. The company is experimenting with one invisible watermark which can be baked directly into the text generated by Claude.

Sounds like a reasonable idea. AI-generated text is everywhere, and knowing where something is coming from can definitely help. Plus, Anthropic doesn’t just hide a marker somewhere inside a document. His approach changes the way Claude selects words to create a statistical pattern that can be detected later.

But there is one detail that bothers me. Anthropic tests how durable this watermark is, even after the text has been modified.

I can already smell trouble there.

Claude touched my writing. Did he really write that?

Think about the translation for a moment. Suppose someone writes an entire essay in Spanish himself and asks Claude to translate it into English. The ideas are theirs. The research is theirs. The arguments are theirs. Claude’s only work is translation.

However, the resulting text may still bear Claude’s watermark.

The same question applies to proofreading. What if someone writes something themselves and asks Claude to correct the grammar? What about shortening a paragraph, changing its tone, cleaning up dictation, or simply making an awkward sentence easier to read?

These are no longer external uses of artificial intelligence. People are increasingly turning to assistants like ChatGPT, Gemini, and Claude for everyday tasks that have little to do with creating original work. A watermark can tell you that Claude was involved in a passage. It is not known whether Claude actually wrote it. Anthropic makes the same claim, saying the watermark shows Claude’s involvement, not who created the original work.

Now imagine explaining this difference to a professor after your essay has just been marked by your detection software.

We already know how messy AI detection can be

I wouldn’t worry nearly as much if our achievements so far in AI recognition they were particularly good. It isn’t.

MIT Sloan’s Guide is fairly straightforward about existing AI detectors. It says it is high error rate and may lead instructors to falsely accuse students of misconduct.

We have already seen what this looks like in practice. Students found themselves defending work they say they wrote after being identified by automated systems as generated by artificial intelligence. In one case he documented it The Guardiana student’s essay was flagged as entirely AI-generated, even though the student said only approved spelling and grammar help was used. The appeal was ultimately accepted.

To be clear, Claude’s watermark is fundamentally different. Traditional AI detectors look at the writing and essentially estimate if an AI could have produced it. Anthropic intentionally plants a detectable signal in Claude’s output. In theory, this would make your system significantly more reliable. But reliability is not the only problem here. The interpretation is.

We use AI to prove that we haven’t used artificial intelligence

Things have reached a somewhat ridiculous point.

Students are worried about the detection of artificial intelligence turning to so-called AI humanizerswhich specifically rewrite the text to make it less likely that the sensors will work. Some students even use these tools in their own written work because they are concerned about false positives. Detector companies are naturally developing methods to identify humanizers.

Read it again.

A human can write something, worry that an AI thinks it was written by an AI, pass it over to another AI to make it look more human, and then use another system to determine if the AI ​​looked human.

It’s a technological ouroboros.

Making Claude’s watermark flexible enough to survive editing and translation is technically impressive. Previous research has shown that translation can defeat some text-watermarking techniques, so addressing this weakness would be a significant step forward.

I don’t think that making the sign more difficult to remove will solve the more important problem.

A watermark needs context

There are good reasons for watermarking AI-generated content. It can help identify mass-produced misinformation, unpublished synthetic texts, or AI-written material that will later end up in educational datasets.

The problem is that AI assistants do a lot more than create content from scratch. People use them to translate text, proofread documents, summarize research findings, help with code, improve accessibility, or simply clean up an email before sending it. In this context, detecting the involvement of artificial intelligence does not automatically tell you who actually created the work.

All of these interactions affect AI to wildly different degrees. If Claude writes an essay from scratch, it’s useful to know. If Claude translates an essay someone spent three weeks doing their own research and writing, it says a lot less knowing that Claude was involved.

The watermark is perfectly capable of answering “Did Claude touch this?” Are you worried about what happens when people start treating the answer as “Claude wrote this” evidence?

Anthropic can build the world’s smartest watermark. Unless the people using it understand this difference, I suspect we’re going to have some problems.

\

Leave a Reply

Your email address will not be published. Required fields are marked *