Anthropic to Launch Watermark Detection API for Identifying AI-Generated Claude Text
Anthropic is rolling out a watermark detection API that lets third-party developers check whether text was generated by Claude. The move stems from EU AI Act compliance requirements and uses a variant of Google DeepMind's SynthID Text method.
Anthropic will soon offer a watermark detection API that lets third-party developers plug AI text detection into their own applications, allowing them to check whether content was generated by Claude.
The company uses a variant of Google DeepMind's SynthID Text method, published in Nature in 2024, adapted for Claude. The technique alters the randomness source during word selection, embedding a traceable statistical pattern into generated text. According to Anthropic, this has no effect "on the content, level of creativity, or readability of Claude's text."
How the watermark works — and where it fails
The watermark degrades in reliability under several conditions, according to Anthropic's published FAQ:
- Short texts and fact-heavy passages offer fewer alternative phrasings, weakening the traceable pattern.
- Code faces the same limitation, since syntax constrains word choice.
- Pure corrections, where a human selects every word, carry no watermark at all.
- Translations do carry the watermark, since Claude selects all words in that process.
- Heavy rewriting by a human can strip the watermark entirely.
Critically, the watermark can only indicate that Claude was likely involved in producing a text. It cannot determine whether Claude wrote the entire piece or made minor edits, and it cannot distinguish human-written text from content generated by a different AI model.
EU AI Act drives global rollout
Anthropic is adding watermarking to comply with the EU AI Act. The company is among roughly 190 signatories to the EU Code of Practice on transparency for AI-generated content, signed in July 2026.
Because there is currently no technical mechanism to restrict the feature by region, Anthropic says watermarking will roll out worldwide rather than exclusively in the EU. The company says it is exploring additional options on an ongoing basis, with further updates expected.
All Claude models released after August 2, 2025 support watermarking natively. Older models will receive the capability in the coming months, according to Anthropic. For files, Anthropic is using the open C2PA standard, which attaches provenance metadata without altering the underlying file.
How this differs from existing AI detectors
Anthropic's approach is architecturally distinct from third-party AI detection tools like Pangram. Those services lack access to Anthropic's watermarking keys and instead rely on statistical pattern-matching — scanning for telltale phrasing, word overuse, or stylistic quirks characteristic of AI-generated text.
Anthropic's watermark detection, by contrast, checks for a specific cryptographic-style signal embedded during generation. This is a fundamentally different detection method and should prove more reliable than pattern-based heuristics, since it doesn't depend on inferring AI authorship from surface-level stylistic cues that can vary or be mimicked.
What this means
This is a compliance-driven feature rollout, not a new model. The watermark detection API doesn't change Claude's outputs or capabilities — it adds a verification layer that regulators, platforms, and third parties can use to flag likely AI-generated content.
The EU AI Act is clearly the forcing function here: without a binding regulatory requirement, voluntary watermarking adoption across the industry has been slow. Anthropic joining ~190 other signatories signals the Code of Practice is gaining real traction, and OpenAI, Google, and Meta likely face similar pressure to expose detection APIs for their own models.
The honest disclosure of limitations — short texts, code, translations, heavy rewrites — matters more than the announcement itself. Watermarking is not a foolproof detection system; it's a probabilistic signal that degrades under common editing behavior. Anyone treating a "no watermark detected" result as proof of human authorship will be wrong often enough to matter, particularly for academic integrity or content moderation use cases that demand higher certainty than this tool can provide.
Related Articles
Anthropic Releases Claude Sonnet 5.5, Now Powering Free Tier on Claude.ai
Anthropic released Claude Sonnet 5.5, claiming it runs 30%+ faster and costs up to 30% less than Sonnet 5 while beating it on benchmarks, at the same price. The model now powers the free tier on claude.ai, giving Anthropic a notably stronger free offering than OpenAI's ChatGPT.
Anthropic Releases Claude Sonnet 5.5: 30% Faster, 30% Cheaper Than Sonnet 5
Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family following last week's Opus 5.5. The model runs more than 30% faster and costs up to 30% less for most work while keeping Sonnet 5's per-token pricing.
Anthropic Python SDK 1.9.0 Adds Reference to Unreleased 'claude-sonnet-5-5' Model ID
Anthropic's anthropic-sdk-python v1.9.0 release adds a reference to an unannounced 'claude-sonnet-5-5' model ID, a new between_tools thinking type, and the ability to run tool calls while a reply streams. No pricing, context window, or benchmark data for the model has been disclosed.
Anthropic Launches Claude Sonnet 5.5, Claims 30% Faster Performance at Lower Cost Than Predecessor
Anthropic has released Sonnet 5.5, the latest version of its mid-tier Claude model, claiming 30% faster performance and significantly lower token costs than its predecessor. The company says the model now outperforms Opus 5.5 on agentic coding tasks and carries cyber capabilities comparable to Opus 5.
Comments
Loading...