// ARS TECHNICA — INTELLIGENZA ARTIFICIALE
Claude's new Scarlet Letter watermark is invisible — for now
The mark flags anything Claude processed, even human writing it only edited.
Anthropic has revealed that it will soon watermark content that is processed (not just generated!) by any of its models. In a support article, Anthropic explained that it was rolling out machine-readable watermarks to comply with the European Union’s AI Act, which requires all AI system providers to watermark AI-generated or manipulated audio, image, text, and video outputs. The law applies to any AI model released after August 2 and provides a grace period until December 2026 for providers to update previously released models.
Anthropic confirmed that moving forward, all new models offered globally—not just in the EU—will mark AI-generated content “from day one.” Text outputs will “carry embedded watermarks,” invisible to the user, and other “generated files will include digitally signed provenance metadata where supported,” Anthropic said.
Notably, Anthropic is deploying a “nuke it from orbit” approach, applying the watermarks to all processed content where supported, even though the EU does not require it for cases where an AI system performs “an assistive function for standard editing” (the guidance’s own example is grammar correction), or where it doesn’t “substantially alter” the user’s text or its meaning. A watermark applied at the model level can’t tell wholesale generation from a comma fix, so Claude may end up stamping exactly the content the law was written to leave alone. How thoroughly it truly watermarks will not be known until Anthropic releases a detection tool that can be tested. The company said that it plans to eventually share details on how to detect marks in order to offer technical support that the EU’s law requires.
Anthropic also noted that the watermarks won’t work on “some platforms or features” that don’t support them. For non-text content, Anthropic will use the C2PA metadata approach to record provenance.
The approach described by the EU and implemented by Anthropic is unfortunately trivially easy for bad actors to bypass, while potentially punishing users who trust the system to accurately label their outputs. Text watermarks work by biasing the model’s word choices in a pattern spread across the entire document, only detectable in aggregate by the right tool. The catch is that “invisible” can also mean the model occasionally trades the best word for a slightly worse one, just to keep the signal intact. Anthropic noted that those marks “will travel with the text when it’s copied and pasted elsewhere, and may persist through some editing.” But if watermarked text is pasted into another chatbot system that edits the text, the watermark could be destroyed. With image and video content, screenshotting/recording or using any decent metadata editing tool will suffice to remove this information, too. And once Anthropic tells the world how to identify these watermarks, building a system to remove them would be trivial.
Further, the potential for misinterpretation seems high; the watermark is not particularly informative. Anthropic explained that a “detected mark provides a signal that content was processed by Claude, but is not fully conclusive.” The only real message the mark sends is that “the content may have been processed by Claude,” Anthropic said, and the mark may even appear on content that was not generated by Claude. On top of this, you have the general public, who may not grasp the difference between processed text and wholly generated text. If the system watermarks human-authored text simply because it was edited in a workflow that touches Claude, suddenly it carries the same denotation as wholly generated text does. And all of this, in a system where the “Lack of a detected mark doesn’t mean the content wasn’t AI-generated or processed,” Anthropic said.
To its credit, Anthropic acknowledges that it may be marking some content that the AI Act does not require to be labeled: “people often use Claude to proofre