model releaseInclusionai

inclusionAI releases Ling 3.1 Flash: 560B MoE, 25B active, 262K context, free on OpenRouter

TL;DR

inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total and 25B active parameters and a 262K-token context window. It is listed as free on OpenRouter through NovitaAI. No benchmark scores have been published on the listing.

3 min read
0

inclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts (MoE) model with 560B total parameters and 25B active per token. It has a 262K-token context window. According to its OpenRouter listing, the model was released on October 2, 2026 and is currently priced at $0 per 1M input tokens and $0 per 1M output tokens.

Key specifications

Spec Detail
Developer inclusionAI
Architecture Mixture-of-experts, hybrid reasoning
Total parameters 560B
Active parameters 25B
Context window 262K tokens
Release date Oct 2, 2026 (per OpenRouter listing)
Pricing Free (input and output)
Training cutoff Not disclosed
Benchmark scores Not published on the listing

Availability and performance

The model is available through OpenRouter, with NovitaAI as the sole listed provider. OpenRouter's listing shows these figures:

  • Latency: 1.12s P50
  • Throughput: 66 tokens per second P50
  • Provider uptime: 100.00% (NovitaAI)
  • OpenRouter availability: 98.93% over the last 24 hours

The listing does not say whether the free pricing is promotional or how long it will last. Paid rates have not been disclosed. The listing also does not itemize supported modalities. Its description refers to a language model, and inclusionAI sells vision capability separately in the Ling 3.0 Flash VL variant.

Position within the Ling family

Ling 3.1 Flash is much larger than earlier "Flash" models in the lineup, according to OpenRouter's catalog:

  • Ling 3.0 Flash: 124B total, about 5.1B active
  • Ling-2.6-flash: 104B total, 7.4B active
  • Ling 3.1 Flash: 560B total, 25B active

That is roughly 4.5x the total parameters and about 5x the active parameters of Ling 3.0 Flash. It sits below the trillion-parameter Ling-2.6-1T and Ring-2.6-1T (63B active) models. Unlike the Ling-2.6 instruct models, it combines instant and reasoning behavior in one hybrid model, as Ling 3.0 Tiny and Ling 3.0 Flash VL also do.

inclusionAI has not published benchmark results for 3.1 Flash on the listing, so its capability claims cannot yet be verified against peers.

What this means

The jump from 5B to 25B active parameters under the "Flash" name suggests inclusionAI is moving that tier toward higher capability, at the expense of the very low serving cost that defined Ling 3.0 Flash ($0.021 input / $0.063 output per 1M tokens). Active parameter count drives inference cost, so paid pricing for 3.1 Flash will probably be well above the 3.0 tier once it is set.

The free listing makes it easy to test now. A 66 tokens-per-second throughput on a single provider is moderate for a model with 25B active parameters. Builders should treat the free tier as an evaluation window rather than a production commitment until pricing and rate limits are published.

The main gap is evidence. Without benchmark scores, a technical report, or a stated training cutoff, the model's standing against other open MoE models at this scale is unknown. Independent evaluations on coding, tool use and long-context tasks will determine whether the larger active footprint is justified.

Related Articles

model release

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

model release

Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta

Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.

model release

Microsoft's MAI-Transcribe-2-Streaming returns first results in ~100 ms across 60 languages

Microsoft AI released MAI-Transcribe-2-Streaming, a real-time transcription model covering 60 languages with first partial results in just over 100 milliseconds. It also launched two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, aimed at voice agents.

model release

Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text

Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.

Comments

Loading...