model release

Sakana AI Releases Fugu Ultra: Multi-Agent Orchestration System with 1M Context Window at $5/$30 per Million Tokens

TL;DR

Sakana AI has released Fugu Ultra, a multi-agent orchestration system that routes tasks across pools of underlying models rather than operating as a single monolithic model. The system supports a 1M token context window and is priced at $5 per million input tokens and $30 per million output tokens.

2 min read
0

Fugu Ultra — Quick Specs

Context window1000K tokens
Input$5/1M tokens
Output$30/1M tokens

Sakana AI Releases Fugu Ultra: Multi-Agent Orchestration System with 1M Context Window

Sakana AI has released Fugu Ultra, the higher-performance model in its Fugu family, now available through OpenRouter. Unlike traditional language models, Fugu Ultra is a learned multi-agent orchestration system trained to route tasks across a swappable pool of underlying models and recursively call instances of itself.

Architecture and Capabilities

According to Sakana AI, Fugu Ultra prioritizes answer quality on complex, multi-step reasoning, coding, and agentic workflows. The system supports:

  • 1 million token context window
  • Configurable reasoning effort
  • Tool calling
  • Built-in web search capabilities

Orchestration tokens consumed by the system are billed as standard input/output tokens, with no separate pricing tier for the routing logic.

Pricing

Fugu Ultra is priced at:

  • Input: $5 per million tokens
  • Output: $30 per million tokens

The effective price can be 60-80% lower when prompt caching is applied for repeated context, according to OpenRouter's monitoring data.

Technical Approach

Rather than training a single large model, Sakana AI's approach involves training a language model to act as an orchestrator, deciding which underlying models to route tasks to and when to recursively invoke additional instances of itself. This architecture represents a departure from the monolithic model paradigm adopted by most major AI labs.

The model was released June 24, 2026, according to OpenRouter's listing. Sakana AI has not disclosed benchmark scores, parameter count, or details about the underlying model pool at this time.

Availability

Fugu Ultra is currently available exclusively through OpenRouter, which forwards requests directly to Sakana AI's infrastructure with no intermediate routing layer.

What This Means

Sakana AI's orchestration approach represents a significant architectural departure from the single-model paradigm. By routing tasks to specialized models rather than attempting to embed all capabilities in one system, Fugu Ultra could potentially offer better cost-performance ratios for complex workflows. However, without published benchmarks, it's unclear how the system compares to frontier models like Claude 3.5 Sonnet or GPT-4 on standardized reasoning and coding tasks. The success of this approach will depend on whether the orchestration overhead is offset by more efficient task routing.

Related Articles

model release

NVIDIA Releases Nemotron VoiceChat 11B, an Open Full-Duplex Speech Model with Live Tool Calling

NVIDIA has released NemotronLabs VoiceChat 11B, an 11-billion-parameter end-to-end full-duplex speech model that unifies streaming speech understanding and generation in one architecture. The model claims to be the first open full-duplex system to support live tool calling during natural conversation, with ~450ms turn-taking latency.

model release

Liquid AI Releases LFM2.5-2.6B, a 2.6B-Parameter Agent Model for On-Device Deployment

Liquid AI has released LFM2.5-2.6B, a 2.6B-parameter model designed to run capable tool-calling agents locally on laptops and phones. The company claims it matches or beats models up to 4x its size on instruction-following and tool-use benchmarks while running under 2.5GB of memory.

model release

Alibaba Releases Qwen3.8 Max, a Multimodal Reasoning Model with 1M Token Context

Alibaba has moved Qwen3.8 Max out of preview into general availability, positioning it as the flagship of the Qwen3.8 series with a 1 million token context window and multimodal input support. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens via OpenRouter.

model release

OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold

OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.

Comments

Loading...