changelogAnthropic

Anthropic Adds Machine-Readable Watermarks to Claude-Generated Text and Files

TL;DR

Anthropic is adding machine-readable watermarks to Claude-generated text and digital signatures to generated files to comply with the EU AI Act's Article 50 transparency mandate. The change applies to models launched after Aug. 2 and rolls out globally, though Anthropic admits detection can fail on heavily edited or short text.

2 min read
0

Anthropic is embedding machine-readable watermarks into text generated by Claude and adding digital signatures to generated files, a compliance move tied to the European Union's AI Act that will apply globally rather than just in the EU.

Why it matters

The change creates a new detection layer for AI-generated content, but with a catch: any document Claude touches—even one used only to clean up, translate, or format human-written copy—can get stamped with an AI signature. Comms teams and writers using Claude as an editing tool, not a drafting tool, could see their human-authored work flagged as AI-generated.

How it works

For models launched in the EU after Aug. 2, Anthropic says it is marking content in two ways, applied "wherever Claude is offered, worldwide," according to the company:

  • Text watermarks: Claude embeds patterns into generated text that Anthropic claims are "imperceptible" to readers but detectable by machine-readable tools.
  • File metadata: Generated media files carry digital signatures confirming the asset was processed by Claude.

Limitations

Anthropic disclosed two significant gaps in its own detection technology:

  • False positives on AI-assisted work: Content can trigger a detected mark even when Claude was used solely to proofread, format, or translate human-written text, according to Anthropic's support documentation.
  • Detection drop-off: Watermarks may not survive if text is heavily rewritten, mixed with other copy, or too short to carry a reliable pattern.

Regulatory context

Anthropic says the changes are designed to satisfy Article 50 of the EU AI Act, which mandates disclosure for AI-generated or AI-manipulated content. OpenAI has outlined a similar compliance approach for the EU AI Act, but according to OpenAI's own guidance, its current watermarking work focuses primarily on images and audio rather than text.

AI providers face a Dec. 2 deadline to bring legacy models into compliance with the EU rules, meaning more detail on industry-wide approaches should surface in the coming months.

Zoom out

The move lands amid a broader push by platforms to flag AI-generated content. LinkedIn is testing a "seems like AI slop" button. Substack has embedded Pangram's AI detection suite for subscribers. Snap has stopped promoting AI-generated video in its main Spotlight feed. Anthropic's watermarking effort signals that regulation—not platform policy—is now the primary driver behind AI content labeling infrastructure.

What this means

This is a compliance feature, not a new model. It doesn't change Claude's capabilities, pricing, or context window—it changes what leaves Claude's output pipeline. The real friction point is precision: an editing tool that also triggers an "AI-generated" flag on human-written text creates disclosure risk for exactly the use case—light editing and translation—that many professional users rely on. Expect pressure on Anthropic and competitors to build a middle category, distinguishing "AI-authored" from "AI-assisted," before the Dec. 2 legacy-model deadline forces broader industry alignment on what a watermark is actually supposed to certify.

Related Articles

product update

Anthropic to Watermark All AI-Generated Text From Claude Models Starting August 2

Anthropic confirmed it will watermark AI-generated text and files from Claude models, complying with the EU AI Act's Transparency Code that took effect August 2. The watermark is applied at the model level and persists through copy-paste, according to the company.

changelog

Anthropic Cuts False Positives in Fable 5's Biology Filter by 85%, Keeps Virology and Toxicology Blocked

Anthropic has cut false positives in Fable 5's biology safety classifier by roughly 85%, letting users ask about lab results, symptoms, and medical questions without being rerouted to the weaker Opus 5 model. Dual-use topics like virology, toxicology, and molecular design remain restricted, with Anthropic citing the difficulty of containing biological threats once released.

changelog

Anthropic SDK v0.121.0 Adds Session Budgets, Mid-Conversation Tool Changes, and GitHub Skills Auto-Loading

Anthropic released version 0.121.0 of its Python SDK on August 7, 2026, introducing a new beta for mid-conversation tool changes, session budgets, an advisor tool, pinned inference location, and skills auto-loading from GitHub. The update also removes retired Claude Opus 4.1 models from the API.

research

Researchers Extract Hidden Chain-of-Thought from OpenAI, Anthropic, Google Models via Shared Encryption Keys

A paper published at stolen-thoughts.com demonstrates that encrypted reasoning traces returned by OpenAI, Anthropic, and Google APIs used the same encryption key across models in a family, allowing attackers to jailbreak weaker sibling models into revealing a stronger model's hidden chain-of-thought in plaintext. All three providers have since patched the vulnerability.

Comments

Loading...