InclusionAI Releases Ling-3.0-flash, a 124B MoE Model with 5.1B Active Parameters
InclusionAI has released Ling-3.0-flash, a 124-billion-parameter Mixture-of-Experts model that activates roughly 5.1 billion parameters per token. The model targets production-scale agentic workloads with a 262K context window and an emphasis on token efficiency.
InclusionAI has released Ling-3.0-flash, a 124-billion-parameter Mixture-of-Experts (MoE) model that activates approximately 5.1 billion parameters per token during inference. The model is listed on OpenRouter with a 262,000-token context window.
What's new
Ling-3.0-flash follows the sparse MoE architecture pattern now common among large-scale open and semi-open model releases: a large total parameter count paired with a much smaller active-parameter footprint per forward pass. At 124B total parameters with 5.1B active, the model activates roughly 4% of its parameters per token, a ratio designed to keep inference costs and latency down relative to a dense model of comparable total size.
According to InclusionAI, the model was built with two priorities: token efficiency and production-scale agentic inference. The company states the goal is to let developers complete more useful work within constrained token, latency, and serving-cost budgets — language that points to agentic and multi-step tool-use workloads as a primary target use case, rather than single-shot chat or completion tasks.
Specifications
- Total parameters: 124B
- Active parameters per token: ~5.1B
- Context window: 262,144 tokens
- Architecture: Mixture-of-Experts (MoE)
- Pricing: Not yet disclosed
- Benchmark scores: Not yet disclosed
OpenRouter's listing shows a release date of July 23, 2026, and does not yet include third-party benchmark results, provider-specific pricing, or an official model card detailing training data cutoff, license terms, or supported modalities beyond text. Independent verification of the parameter counts and efficiency claims has not been published.
Context
InclusionAI, a relatively new entrant among Chinese AI labs shipping openly-listed models on aggregators like OpenRouter, joins a growing field of companies — including Alibaba Qwen, DeepSeek, Moonshot AI, and Zhipu AI — releasing large MoE models optimized for cost-efficient inference. The "flash" naming convention, also used by Google for its lower-latency Gemini variants, signals positioning as a faster, cheaper option relative to a presumed larger "Ling-3.0" flagship, though InclusionAI has not detailed a broader model family lineup in the available listing.
The 262K context window places Ling-3.0-flash in the same range as long-context models from Anthropic and Google, useful for large-document analysis, extended agentic sessions, or multi-turn tool-calling chains where conversation history accumulates quickly.
What this means
Ling-3.0-flash's core claim — a 124B/5.1B split — is standard practice for cost-conscious MoE deployment, and the emphasis on agentic inference suggests InclusionAI is targeting developers building autonomous or tool-using systems rather than general chatbot use cases. However, with no published pricing, no independent benchmark scores, and no detailed model card yet available, the practical value of this release can't be assessed until those numbers surface. Developers evaluating the model should treat efficiency and capability claims as unverified until pricing from serving providers and third-party benchmark comparisons against similarly-sized MoE models (such as Qwen's and DeepSeek's offerings) become available.
Related Articles
Alibaba Releases Qwen3.8 Max (0902), a 2.4-Trillion-Parameter MoE Model With 1M-Token Context
Alibaba's Qwen team released Qwen3.8 Max (0902), a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window that accepts text, image, and video input. The snapshot is post-trained for coding, agentic workflows, and long-horizon task execution, priced at $2/$6 per 1M input/output tokens.
InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters
InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.
OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing
A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.
Alibaba Releases Qwen-Drive 1.0, an Open Driving Model That Explains Its Own Decisions
Alibaba has released Qwen-Drive 1.0, a driving model built on Qwen3.5-4B that handles spatial perception, route planning, and cockpit dialogue in a single system. Reinforcement learning cut the rate of off-road driving errors in simulation from 24 percent to 12 percent, though the model's stated reasoning doesn't always match its actual maneuvers.
Comments
Loading...