Anthropic Reverses Policy That Silently Throttled AI Researchers Using Claude
Anthropic has reversed a controversial policy that would have made Claude Fable and Mythos models silently throttle responses to researchers working on frontier AI development. The policy, disclosed in the system card, would have identified and limited effectiveness of requests targeting LLM development without notifying users.
Anthropic Reverses Policy That Silently Throttled AI Researchers Using Claude
Anthropic has walked back a policy that would have made its Claude Fable and Mythos models secretly limit the effectiveness of responses to AI researchers working on frontier language model development.
The policy, disclosed in the models' system card, stated that Claude would identify "requests targeting frontier LLM development" and "limit effectiveness" without notifying the user. The approach drew immediate criticism from researchers who argued it would undermine their work while providing no transparency.
Company Response
"We're changing Fable 5's safeguards for frontier LLM development to make them visible," Anthropic said in a statement to WIRED. "We made the wrong tradeoff and we apologize for not getting the balance right."
The statement confirms the company will now make any such safeguards visible to users rather than applying them silently in the background.
What Was at Stake
The original policy would have affected researchers using Claude for:
- Developing new language models
- Testing AI safety techniques
- Benchmarking model capabilities
- Research on AI alignment
Critics noted the policy could have sabotaged legitimate AI safety research while providing minimal actual security benefit, since researchers working on potentially harmful applications would simply use other models.
Industry Context
The controversy comes as AI companies face increasing pressure to balance model capabilities with safety concerns. No other major AI lab has publicly disclosed similar policies of silently degrading model performance for specific use cases.
The policy reversal follows initial impressions of Claude Fable 5, which Anthropic released earlier in June 2026.
What This Means
This rapid reversal demonstrates the AI industry's ongoing struggle to define appropriate safety boundaries. Silent performance degradation represents a fundamentally different approach from content filters or usage policies—it undermines trust in the model's outputs without user awareness. Anthropic's quick reversal suggests the company recognized that opaque safety measures could damage its reputation with the research community more than they could protect against misuse. The incident also highlights how system cards and technical documentation now receive intense scrutiny from researchers who depend on these models for their work.
Related Articles
Anthropic Joins Google in Watermarking AI-Generated Text, Reviving Debate Over Output Quality
Anthropic announced on August 11 that all future Claude models will embed an invisible watermark in generated text, following Google's lead with SynthID-Text. The move is partly driven by the EU AI Act, which mandates watermarking for AI models released after August 2, 2026, though researchers remain split on whether the technique degrades output quality.
Meta Launches Muse, a WhatsApp AI Agent That Shops, Emails, and Negotiates on Users' Behalf
Meta has launched Muse, an AI agent controlled through WhatsApp that can browse the web, fill out forms, negotiate prices, and complete purchases with user approval via Stripe's Link service. The agent runs in an isolated virtual machine monitored by a separate oversight process Meta calls Sentinel.
Nvidia and Palantir Deploy AI to Manage Supply Chains, Starting With Nvidia's Own 1.3-Million-Part Racks
Nvidia and Palantir are integrating Nvidia's open Nemotron models and cuOpt optimization engine into Palantir's Foundry platform to manage complex supply chains. The first deployment is Nvidia's own operation, where a single Vera Rubin server rack contains 1.3 million components.
Meta's New AI Agent 'Muse' Takes Over Social Handles Once Used by Rock Band Muse
Meta officially launched its AI agent 'Muse' this week using the @Muse handle on Instagram and X — the same usernames the rock band Muse used for years before quietly switching to @museband earlier this summer. Meta has not explained how the handle changes occurred.
Comments
Loading...