AI 安全與內容來源
Anthropic Adds Model-Level Text Watermarks to New Claude Outputs, Uses C2PA Signatures for Files
Anthropic says Claude models released starting August 2 will embed invisible markers in text and add signed provenance data to supported images and documents. Technical specifications and third-party detection tools have not yet been released, so the markers cannot currently be treated as definitive proof of AI generation.

Anthropic’s newly disclosed content-marking scheme goes beyond attaching an “AI-generated” label in the Claude web interface: it places signals within the model output itself. The company says Claude models released on or after August 2, 2026 that support the feature will weave invisible watermarks into text generated across all products, APIs, and partner platforms. Because the marker forms part of the text, it may survive copying and pasting and may also withstand some editing. The feature will apply globally, rather than only to users in the European Union.
Files such as images will use a different mechanism. Anthropic plans to add digitally signed provenance metadata to supported formats, including PNG, JPEG, and SVG, using C2PA Content Credentials. C2PA records a file’s origin and processing history, allowing signatures to be used to verify whether it has been modified. This is not equivalent to an in-content signal such as a text watermark, and its persistence characteristics also differ when a file is resaved, transcoded, or stripped of metadata.
The change responds to the transparency obligations under Article 50 of the EU AI Act, which take effect on August 2. The provision requires generative systems, where technically feasible, to provide machine-readable, detectable, and interoperable markings. However, Anthropic has not yet explained its text-encoding method, false-positive rate, resistance to rewriting, or effects on outputs in different languages and in code. A third-party detection interface also remains pending in future technical documentation. The company explicitly warns that finding a marker indicates only that content may have been processed by Claude, while failing to find one cannot prove that the content was created by a human. Engineering teams therefore should not immediately connect the system to blocking or disciplinary rules. The next steps should be to monitor the detection API, key and revocation mechanisms, transition arrangements for older models, and real-world reliability after formatting, translation, and rewriting by multiple models.