Post

AT
Ars Technica

Claude's new Scarlet Letter watermark is invisible — for now

Anthropic has revealed that it will soon watermark content that is processed ( not just generated!

The mark flags anything Claude processed, even human writing it only edited.

Anthropic confirmed that moving forward, all new models offered globally—not just in the EU—will mark AI-generated content “from day one.” Text outputs will “carry embedded watermarks,” invisible to the user, and other “generated files will include digitally signed provenance metadata where supported,” Anthropic said.

Notably, Anthropic is deploying a “ nuke it from orbit ” approach, applying the watermarks to all processed content where supported, even though the EU does not require it for cases where an AI system performs “an assistive function for standard editing” (the guidance’s own example is grammar correction), or where it doesn’t “substantially alter” the user’s text or its meaning.

A watermark applied at the model level can’t tell wholesale generation from a comma fix, so Claude may end up stamping exactly the content the law was written to leave alone. How thoroughly it truly watermarks will not be known until Anthropic releases a detection tool that can be tested. The company said that it plans to eventually share details about how to detect marks in order to offer technical support that the EU’s law requires.

Anthropic also noted that the watermarks won’t work on “some platforms or features” that don’t support them. For non-text content, Anthropic will use the C2PA metadata approach to record provenance.

By Ashley Belanger
Tweet media