model release

InclusionAI releases Ling-2.6-flash: 104B parameter model with 7.4B active parameters, free on OpenRouter

TL;DR

InclusionAI has released Ling-2.6-flash, an instruction-tuned model with 104 billion total parameters and 7.4 billion active parameters, available free through OpenRouter. The model features a 262,144-token context window and is designed for agent workflows requiring fast responses and high token efficiency.

2 min read
0

InclusionAI releases Ling-2.6-flash: 104B parameter model with 7.4B active parameters, free on OpenRouter

InclusionAI has released Ling-2.6-flash, an instruction-tuned model with 104 billion total parameters and 7.4 billion active parameters. The model is available at no cost through OpenRouter as of April 21, 2026.

Model Specifications

Ling-2.6-flash features a 262,144-token context window (approximately 262K tokens) and is offered with $0 per million tokens for both input and output. The model uses a sparse architecture, activating only 7.4B of its 104B total parameters during inference.

According to inclusionAI, the model is designed for "real-world agents that require fast responses, strong execution, and high token efficiency." The company claims it delivers performance comparable to state-of-the-art models at similar scale while reducing token usage across coding, document processing, and lightweight agent workflows.

Technical Architecture

The model's sparse activation approach—using only 7.1% of its total parameters per inference—enables faster response times compared to fully-activated models of similar total parameter count. This design pattern follows recent trends in mixture-of-experts and sparse architectures.

The model is accessible through OpenRouter's unified API, which provides OpenAI-compatible endpoints. OpenRouter routes requests to available providers with automatic fallbacks for uptime optimization.

Availability

Ling-2.6-flash is currently available exclusively through OpenRouter's platform. The company has not disclosed whether the model will be released through other providers or made available for self-hosting. No benchmark scores have been published at this time.

InclusionAI is not among the previously established AI model providers tracked in industry databases, suggesting this is either a new entrant or an independent research team making their first public model release.

What This Means

The release of a free, high-parameter-count model with sparse activation represents competitive pressure on existing model providers. If performance claims are verified through independent benchmarks, the 262K context window at zero cost could make this attractive for agent applications and document processing tasks. However, without published benchmark scores or information about training data and capabilities, adoption will likely depend on real-world testing by developers. The sparse activation design (7.4B active from 104B total) suggests this is optimized for cost-efficient inference rather than maximum capability.

Related Articles

model release

Meituan launches LongCat 2.0: 1.6T parameter MoE model with 1M+ context window at $0.30 per 1M input tokens

Meituan has released LongCat 2.0, a sparse mixture-of-experts language model with 48 billion active parameters out of 1.6 trillion total. The model features a 1,049,000 token context window and costs $0.30 per 1M input tokens and $1.20 per 1M output tokens.

model release

Alibaba previews Qwen3.8 with 2.4 trillion parameters, claims second place without benchmark data

Alibaba unveiled Qwen3.8 at the World Artificial Intelligence Conference in Shanghai, claiming the 2.4 trillion parameter model ranks second only to Anthropic's Fable 5. The company provided no benchmark scores, model card, or independent verification to support the claim.

model release

Moonshot AI releases Kimi K3, largest open-weight model at 2.8 trillion parameters

Moonshot AI released Kimi K3 on July 16, 2025, an open-weight model with 2.8 trillion parameters. The model represents the largest openly available model by parameter count, entering what the industry categorizes as the 3T class.

model release

Alibaba releases Qwen 3.8, a 2.4 trillion parameter open-weight model claiming second place behind Fable 5

Alibaba has released Qwen 3.8, a 2.4 trillion parameter open-weight model that the company claims trails only Fable 5. The multimodal model processes images, videos, and documents, with a preview available through Alibaba's platforms at 10 percent of standard pricing.

Comments

Loading...