Sakana AI Claims Fugu Ultra v1.1 Router Beats Anthropic's Fable 5 Without Including It in the Pool
Sakana AI has released Fugu Ultra v1.1, an update to its multi-model router, claiming performance gains up to 7.9 points over v1.0 and results that beat Anthropic's Fable 5 despite Fable 5 not being part of the router's model pool. All benchmark figures come from Sakana itself and remain unverified.
Fugu Ultra v1.1 — Quick Specs
Sakana AI has released Fugu Ultra v1.1, an update to its AI model router that distributes each query across a pool of publicly available top-tier language models. According to Sakana, the new version delivers performance gains of up to 7.9 points over v1.0, with the largest improvements on ProgramBench and TerminalBench 2.1.
The company claims Fugu v1.1 now beats Anthropic's Fable 5 across most benchmarks — a notable claim given that Fable 5 is not included in Fugu's selection pool of underlying models. Sakana has not published the full benchmark tables or independent verification for these results. All performance figures currently come from Sakana's own testing.
What Fugu does
Fugu Ultra works as a router rather than a standalone model: it evaluates each incoming query and dispatches it to whichever model in its pool it judges best suited to handle that specific request, then returns the result. This approach lets Sakana claim state-of-the-art performance by leveraging whichever frontier model currently performs best on a given task category, rather than training a single foundation model from scratch.
Pricing remains unchanged at $5 per million input tokens and $30 per million output tokens. Sakana says it takes roughly two weeks of training and evaluation before any new top-tier model gets added to the router's pool — which explains, in the company's telling, why a model as recent as Fable 5 hasn't been incorporated yet despite the router allegedly outperforming it.
Sakana has published a technical report describing the router's architecture, though the underlying selection and orchestration mechanisms have not been independently audited.
New in this update
Fugu v1.1 adds a Claude Code-compatible endpoint, allowing developers to call the router directly from the terminal alongside existing coding workflows. Fugu has been available since launch on platforms including OpenRouter and Vercel.
The first version of Fugu received a lukewarm reception. Reviewers and users flagged high token consumption, slow response times, and inconsistent output quality — criticisms Sakana has not directly addressed in this update's announcement beyond the claimed benchmark improvements.
Sakana AI continues to exclude the EU and EEA from Fugu's availability, citing GDPR and unspecified "EU-specific regulations."
What this means
Router-based products like Fugu sidestep the cost and complexity of training frontier models by instead orchestrating access to models built by others — but that also means their competitive claims are entirely dependent on which models sit in the pool and how quickly new ones get added. A two-week onboarding lag means Fugu's benchmark comparisons will regularly involve outdated competitor lineups, which raises questions about how meaningful a claim of beating an excluded model actually is.
Until Sakana publishes verifiable benchmark data or third parties replicate the results, the performance claims for v1.1 should be treated as marketing rather than established fact. The persistent complaints about token usage and latency from v1.0 are also unresolved in this release, which matters more for developer adoption than head-to-head benchmark wins against models not actually available for direct comparison.
Related Articles
Anthropic Releases Claude Fable 5.1 and Mythos 5.1, Cuts Cache Pricing 75% But Output Tokens Jump 70%
Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, claiming the top spot on Artificial Analysis's Intelligence Index at 66. Cache-read pricing dropped 75% to $0.25 per million tokens, but a 1.7x increase in output token usage pushes net per-task cost up 20%.
Meta Releases Muse Spark 1.3, Cheapest Model in Its Performance Class at $0.55 Per Task
Meta has released Muse Spark 1.3, its fourth model in five months, with an xhigh tier available now and a more powerful max tier in limited preview. The model improves sharply on agentic benchmarks and costs $0.55 per index task—cheaper than any rival at the same performance level—but still trails Claude Fable 5.1 on most tests.
Google Launches Gemini 3.8 Flash, Warns It May Use More Tokens Despite Unchanged Pricing
Google released Gemini 3.8 Flash just weeks after Gemini 3.7 Flash, keeping the same per-token pricing of $0.75/$3.75 per million input/output tokens but warning it may consume more tokens overall. The model also ships with a cyber-focused variant restricted to a new government partner program called Fairwind.
Anthropic Adds Explicit Song Lyric and Copyrighted Character Bans to Claude's System Prompt
Anthropic quietly added detailed new restrictions to Claude's published system prompts, explicitly barring song lyric reproduction and AI-generated images of copyrighted characters. The change follows closely on the heels of a lawsuit from Sony Music Publishing and Warner Chappell.
Comments
Loading...