AI-Written Text Is Getting an Invisible Fingerprint
Anthropic will be watermarking text generated by future Claude models. The technology is clever, probabilistic, and carefully limited in what it claims to prove.
The problem may begin when humans start using it.
In order to comply with the EU AI Act, Anthropic announced that future Claude models will generate text containing a machine-detectable watermark designed to indicate the likelihood that Claude was involved in producing the text.
At first glance, that sounds straightforward enough. AI generates text. The text carries a detectable fingerprint. Someone checks the fingerprint later. Problem solved.
Except that isn't really what the technology does. And once schools, employers, publishers, search engines, content platforms, and other institutions begin using it, things could get considerably more complicated.
First, What Is the Watermark?
It is not metadata. It does not contain hidden characters. Nothing extra is secretly inserted into the text. It does not identify the user, the company, the Claude account, or the conversation that produced the text.
Instead, the watermark exists in the statistical pattern of the words Claude chooses. Large language models generate text one token at a time. At many points during that process, several different words might work equally well.