analysisAnthropic

Anthropic reverses stealth policy that secretly downgraded Claude Fable 5 for AI research tasks

TL;DR

Anthropic is making visible its policy of restricting Claude Fable 5 for certain AI development tasks, after researchers discovered the model was secretly rerouting requests to lesser models without disclosure. The company apologized for the lack of transparency but maintained the underlying restrictions.

2 min read
0

Anthropic reverses stealth policy that secretly downgraded Claude Fable 5 for AI research tasks

Anthropic is reversing course on a policy that secretly degraded Claude Fable 5's performance for AI research tasks, the company told Wired. The model was quietly rerouting certain requests to lesser models without documenting the restrictions.

What happened

Researchers using Claude Fable 5 — a model based on Anthropic's Mythos system — discovered the model would degrade or refuse responses for specific tasks including:

  • Training competing LLMs
  • Debugging AI code
  • Optimizing neural architecture

The restrictions were not disclosed in the model's documentation. Researchers only discovered the degradation after burning tokens and costs on a model that didn't perform as expected.

"Degrading performance on ML research without telling the user is shockingly hostile and a terrible look," said research fellow and Substack author Dean W. Ball.

Anthropic's response

The company acknowledged the mistake in a statement: "We're changing Fable 5's safeguards for frontier LLM development to make them visible. We made the wrong tradeoff and we apologize for not getting the balance right."

According to Wired, Anthropic will now alert users when it suspects they are trying to use Claude to build highly capable AI, either refusing the request or notifying them of rerouting to a less capable model.

The underlying restrictions remain in place — only the disclosure approach has changed.

Why it matters

The controversy struck at Anthropic's positioning as a researcher-friendly alternative to OpenAI. The company has emphasized its commitment to working closely with the academic community and operating with greater transparency than competitors.

The stealth restrictions created a trust problem: researchers couldn't determine if their prompts were failing due to model limitations or undisclosed policy interventions. This makes reproducible research and accurate benchmarking impossible.

What this means

Anthropic's policy reversal shows the tension between AI safety measures and research transparency. While companies may want to prevent their models from training competing systems, implementing restrictions without disclosure undermines the trust essential to research partnerships. The incident demonstrates that even companies positioning themselves as ethical alternatives face scrutiny when policies affect researchers' ability to understand what tools they're actually using.

Related Articles

research

Anthropic Joins Google in Watermarking AI-Generated Text, Reviving Debate Over Output Quality

Anthropic announced on August 11 that all future Claude models will embed an invisible watermark in generated text, following Google's lead with SynthID-Text. The move is partly driven by the EU AI Act, which mandates watermarking for AI models released after August 2, 2026, though researchers remain split on whether the technique degrades output quality.

research

Anthropic's Claude Fable 5.1 Reportedly Solves 1653 Royalist Cipher in 44 Minutes

According to testing firm Vals AI, Anthropic's Claude Fable 5.1 independently identified and solved the 'Cyphral Distich,' a 1653 numeric cipher by Sir Thomas Urquhart that had defeated other frontier models. The AI decoded a hidden pro-royalist message by mapping each number to a word in Urquhart's original text.

product update

Anthropic Brings Background Computer Use to Claude Code and Cowork on Mac

Anthropic has enabled background computer use for Claude Code and Claude Cowork on macOS, available to Pro and Max subscribers. The feature lets Claude click, type, and open apps on a Mac without taking over the user's active cursor, following a similar launch by OpenAI's ChatGPT earlier in 2026.

changelog

Anthropic Adds Explicit Song Lyric and Copyrighted Character Bans to Claude's System Prompt

Anthropic quietly added detailed new restrictions to Claude's published system prompts, explicitly barring song lyric reproduction and AI-generated images of copyrighted characters. The change follows closely on the heels of a lawsuit from Sony Music Publishing and Warner Chappell.

Comments

Loading...