inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter
inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.
inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts (MoE) model with 560B total parameters and 25B active per token. It has a 262K-token context window. According to its OpenRouter listing, the model was released on October 2, 2026 and is currently priced at $0 per 1M input tokens and $0 per 1M output tokens.
Key specifications
| Spec | Detail |
|---|---|
| Developer | inclusionAI |
| Architecture | Mixture-of-experts, hybrid reasoning |
| Total parameters | 560B |
| Active parameters | 25B |
| Context window | 262K tokens |
| Release date | Oct 2, 2026 (per OpenRouter listing) |
| Pricing | Free (input and output) |
| Training cutoff | Not disclosed |
| Benchmark scores | Not published on the listing |
Availability and performance
The model is available through OpenRouter, with NovitaAI as the sole listed provider. OpenRouter's listing shows these figures:
- Latency: 1.12s P50
- Throughput: 66 tokens per second P50
- Provider uptime: 100.00% (NovitaAI)
- OpenRouter availability: 98.93% over the last 24 hours
The listing does not say whether the free pricing is promotional or how long it will last. Paid rates have not been disclosed. The listing also does not itemize supported modalities. Its description refers to a language model, and inclusionAI sells vision capability separately in the Ling 3.0 Flash VL variant.
Position within the Ling family
Ling 3.1 Flash is much larger than earlier "Flash" models in the lineup, according to OpenRouter's catalog:
- Ling 3.0 Flash: 124B total, about 5.1B active
- Ling-2.6-flash: 104B total, 7.4B active
- Ling 3.1 Flash: 560B total, 25B active
That is roughly 4.5x the total parameters and about 5x the active parameters of Ling 3.0 Flash. It sits below the trillion-parameter Ling-2.6-1T and Ring-2.6-1T (63B active) models. Unlike the Ling-2.6 instruct models, it combines instant and reasoning behavior in one hybrid model, as Ling 3.0 Tiny and Ling 3.0 Flash VL also do.
inclusionAI has not published benchmark results for 3.1 Flash on the listing, so its capability claims cannot yet be verified against peers.
What this means
The jump from 5B to 25B active parameters under the "Flash" name suggests inclusionAI is moving that tier toward higher capability, at the expense of the very low serving cost that defined Ling 3.0 Flash ($0.021 input / $0.063 output per 1M tokens). Active parameter count drives inference cost, so paid pricing for 3.1 Flash will probably be well above the 3.0 tier once it is set.
The free listing makes it easy to test now. A 66 tokens-per-second throughput on a single provider is moderate for a model with 25B active parameters. Builders should treat the free tier as an evaluation window rather than a production commitment until pricing and rate limits are published.
The main gap is evidence. Without benchmark scores, a technical report, or a stated training cutoff, the model's standing against other open MoE models at this scale is unknown. Independent evaluations on coding, tool use and long-context tasks will determine whether the larger active footprint is justified.
Related Articles
Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter
Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.
Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta
Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.
Microsoft's MAI-Transcribe-2-Streaming returns first results in ~100 ms across 60 languages
Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model covering 60 languages with first partial results in just over 100 milliseconds. It also launched two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, aimed at voice agents.
Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text
Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.
Comments
Loading...