InclusionAI Releases Ling-3.0-flash, a 124B MoE Model with 5.1B Active Parameters
InclusionAI has released Ling-3.0-flash, a 124-billion-parameter Mixture-of-Experts model that activates roughly 5.1 billion parameters per token. The model targets production-scale agentic workloads with a 262K context window and an emphasis on token efficiency.
InclusionAI has released Ling-3.0-flash, a 124-billion-parameter Mixture-of-Experts (MoE) model that activates approximately 5.1 billion parameters per token during inference. The model is listed on OpenRouter with a 262,000-token context window.
What's new
Ling-3.0-flash follows the sparse MoE architecture pattern now common among large-scale open and semi-open model releases: a large total parameter count paired with a much smaller active-parameter footprint per forward pass. At 124B total parameters with 5.1B active, the model activates roughly 4% of its parameters per token, a ratio designed to keep inference costs and latency down relative to a dense model of comparable total size.
According to InclusionAI, the model was built with two priorities: token efficiency and production-scale agentic inference. The company states the goal is to let developers complete more useful work within constrained token, latency, and serving-cost budgets — language that points to agentic and multi-step tool-use workloads as a primary target use case, rather than single-shot chat or completion tasks.
Specifications
- Total parameters: 124B
- Active parameters per token: ~5.1B
- Context window: 262,144 tokens
- Architecture: Mixture-of-Experts (MoE)
- Pricing: Not yet disclosed
- Benchmark scores: Not yet disclosed
OpenRouter's listing shows a release date of July 23, 2026, and does not yet include third-party benchmark results, provider-specific pricing, or an official model card detailing training data cutoff, license terms, or supported modalities beyond text. Independent verification of the parameter counts and efficiency claims has not been published.
Context
InclusionAI, a relatively new entrant among Chinese AI labs shipping openly-listed models on aggregators like OpenRouter, joins a growing field of companies — including Alibaba Qwen, DeepSeek, Moonshot AI, and Zhipu AI — releasing large MoE models optimized for cost-efficient inference. The "flash" naming convention, also used by Google for its lower-latency Gemini variants, signals positioning as a faster, cheaper option relative to a presumed larger "Ling-3.0" flagship, though InclusionAI has not detailed a broader model family lineup in the available listing.
The 262K context window places Ling-3.0-flash in the same range as long-context models from Anthropic and Google, useful for large-document analysis, extended agentic sessions, or multi-turn tool-calling chains where conversation history accumulates quickly.
What this means
Ling-3.0-flash's core claim — a 124B/5.1B split — is standard practice for cost-conscious MoE deployment, and the emphasis on agentic inference suggests InclusionAI is targeting developers building autonomous or tool-using systems rather than general chatbot use cases. However, with no published pricing, no independent benchmark scores, and no detailed model card yet available, the practical value of this release can't be assessed until those numbers surface. Developers evaluating the model should treat efficiency and capability claims as unverified until pricing from serving providers and third-party benchmark comparisons against similarly-sized MoE models (such as Qwen's and DeepSeek's offerings) become available.
Related Articles
Anthropic Releases Claude Opus 5 with 1M-Token Context and $5/$25 Per-Million-Token Pricing
Anthropic has released Claude Opus 5, its new flagship model built for complex reasoning, coding, and multi-agent coordination. The model ships with a 1 million token context window and pricing of $5 per million input tokens and $25 per million output tokens.
Poolside Releases Laguna S 2.1, an 8B-Active-Parameter Open Coding Model That Rivals Systems 20x Its Size
Poolside has released Laguna S 2.1, a mixture-of-experts coding model with 8 billion active parameters out of 118 billion total, its third coding model release in three months. The company claims it outperforms open-weight models 10 to 20 times its size on agentic coding benchmarks like Terminal-Bench 2.1 and DeepSWE.
Anthropic Launches Claude Opus 5, Claims Parity With Rival Fable 5 at Half the Cost
Anthropic has released Claude Opus 5, its new flagship model, claiming performance comparable to rival model Fable 5 at half the cost. The company says Opus 5 leads on several coding and knowledge-work benchmarks while requiring far less manual intervention.
Anthropic Releases Claude Opus 5, Claims Near-Fable-5 Performance at Opus Pricing
Anthropic released Claude Opus 5 on July 24, 2026, pricing it identically to Opus 4.8 at $5 per million input tokens and $25 per million output tokens while claiming performance approaching its higher-tier Fable 5 model. The release includes a faster processing mode, automatic safety fallbacks, and mid-conversation tool switching that preserves prompt caching.
Comments
Loading...