Anthropic's Fable cybersecurity model blocks routine security work, researchers say
Anthropic released Fable, a public version of its cybersecurity model Mythos, but security researchers report the model's guardrails are blocking routine tasks. The model flags requests as cybersecurity-related even for reading blog posts or requesting code reviews, downgrading to Claude Opus 4.8 when triggered.
Anthropic's Fable cybersecurity model blocks routine security work, researchers say
Anthropic released Fable on Tuesday, a public and limited version of its cybersecurity model Mythos, but security researchers are reporting the model's guardrails are blocking legitimate work.
"[Fable] rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post," said Valentina "Chompie" Palmiotti, a security researcher at IBM X-Force.
How the guardrails work
When triggered, Fable pauses the chat and displays a message that "safety measures flagged this message for cybersecurity or biology topics." The model then downgrades to Claude Opus 4.8. The restrictions aim to prevent Fable from being used to develop malware or compromise software, with similar restrictions on biology to prevent biological weapon development.
Matt Suiche, a cybersecurity veteran and member of the technical staff at AI cybersecurity startup Tolmo, told TechCrunch the system appears keyword-based. "If you ask it to write secure code, it assumes it is cybersecurity related work instead of software engineering best practices, and you get downgraded," Suiche said. "Anything in the lexical field of 'cybersecurity' triggers the guardrails."
Another researcher reported that even requesting a code review triggers the guardrails.
Access to Mythos remains restricted
Anthropic released Mythos in April through Project Glasswing, restricting access to a limited number of companies and organizations for securing critical software and infrastructure. Last week, Anthropic expanded Mythos access to hundreds of organizations across 15 countries, but the full model remains unavailable to most users.
Anthropic operates a Cyber Verification Program that allows approved cybersecurity professionals to use Claude with fewer limitations. OpenAI maintains a similar program called Trusted Access for Cyber.
What this means
The overly broad guardrails on Fable highlight the challenge of releasing capable AI models for specialized domains. While Anthropic's caution is understandable given malware development risks, the current implementation appears to conflate basic security engineering practices with malicious activity. Suiche noted the approach may be appropriate for an initial release: "It's better to catch more people than not enough when you do such a release and to relax the guardrails over time." The effectiveness of Fable as a security tool will depend on Anthropic's ability to calibrate these restrictions to allow legitimate defensive security work while blocking offensive capabilities.
Related Articles
OpenAI Releases GPT-6 Astra, First Model to Cross 'Critical' Cybersecurity Threshold
OpenAI has begun rolling out GPT-6 Astra, the first model to reach the company's internal 'Critical' cybersecurity threshold. Access is being phased, with companies in OpenAI's Daybreak cybersecurity program getting priority following added safeguards after a prior model containment breach.
OpenAI Launches GPT-6 Astra, Says the Model May Already Qualify as AGI
OpenAI has released GPT-6 Astra, its most capable model yet, with benchmark scores the company says surpass GPT-5.6 Sol and Anthropic's Fable 5 models. President Greg Brockman called it a step into the 'AGI era,' though OpenAI acknowledges there's no agreed-upon threshold for that term.
OpenAI Releases Astra, Claims New Flagship Model Beats Rivals on Coding and Cybersecurity Benchmarks
OpenAI released Astra on Thursday, calling it its most capable and most aligned model yet. The model uses a reasoning technique called 'opaque recurrence' that critics say reduces visibility into its chain of thought.
OpenAI Launches GPT-Image 2.5 With Two New API Models: Sunburst and Flare
OpenAI has released ChatGPT Images 2.5, introducing two new API model IDs — gpt-image-2.5-sunburst and gpt-image-2.5-flare — with improved multi-turn instruction following and better preservation of subjects from reference photos. OpenAI says its image models have now generated more than 3 billion images across ChatGPT and the API.
Comments
Loading...