InclusionAI releases Ling-2.6-flash: 104B parameter model with 7.4B active parameters, free on OpenRouter
InclusionAI has released Ling-2.6-flash, an instruction-tuned model with 104 billion total parameters and 7.4 billion active parameters, available free through OpenRouter. The model features a 262,144-token context window and is designed for agent workflows requiring fast responses and high token efficiency.
InclusionAI releases Ling-2.6-flash: 104B parameter model with 7.4B active parameters, free on OpenRouter
InclusionAI has released Ling-2.6-flash, an instruction-tuned model with 104 billion total parameters and 7.4 billion active parameters. The model is available at no cost through OpenRouter as of April 21, 2026.
Model Specifications
Ling-2.6-flash features a 262,144-token context window (approximately 262K tokens) and is offered with $0 per million tokens for both input and output. The model uses a sparse architecture, activating only 7.4B of its 104B total parameters during inference.
According to inclusionAI, the model is designed for "real-world agents that require fast responses, strong execution, and high token efficiency." The company claims it delivers performance comparable to state-of-the-art models at similar scale while reducing token usage across coding, document processing, and lightweight agent workflows.
Technical Architecture
The model's sparse activation approach—using only 7.1% of its total parameters per inference—enables faster response times compared to fully-activated models of similar total parameter count. This design pattern follows recent trends in mixture-of-experts and sparse architectures.
The model is accessible through OpenRouter's unified API, which provides OpenAI-compatible endpoints. OpenRouter routes requests to available providers with automatic fallbacks for uptime optimization.
Availability
Ling-2.6-flash is currently available exclusively through OpenRouter's platform. The company has not disclosed whether the model will be released through other providers or made available for self-hosting. No benchmark scores have been published at this time.
InclusionAI is not among the previously established AI model providers tracked in industry databases, suggesting this is either a new entrant or an independent research team making their first public model release.
What This Means
The release of a free, high-parameter-count model with sparse activation represents competitive pressure on existing model providers. If performance claims are verified through independent benchmarks, the 262K context window at zero cost could make this attractive for agent applications and document processing tasks. However, without published benchmark scores or information about training data and capabilities, adoption will likely depend on real-world testing by developers. The sparse activation design (7.4B active from 104B total) suggests this is optimized for cost-efficient inference rather than maximum capability.
Related Articles
Meituan launches LongCat 2.0: 1.6T parameter MoE model with 1M+ context window at $0.30 per 1M input tokens
Meituan has released LongCat 2.0, a sparse mixture-of-experts language model with 48 billion active parameters out of 1.6 trillion total. The model features a 1,049,000 token context window and costs $0.30 per 1M input tokens and $1.20 per 1M output tokens.
Alibaba previews Qwen3.8 with 2.4 trillion parameters, claims second place without benchmark data
Alibaba unveiled Qwen3.8 at the World Artificial Intelligence Conference in Shanghai, claiming the 2.4 trillion parameter model ranks second only to Anthropic's Fable 5. The company provided no benchmark scores, model card, or independent verification to support the claim.
Moonshot AI releases Kimi K3, largest open-weight model at 2.8 trillion parameters
Moonshot AI released Kimi K3 on July 16, 2025, an open-weight model with 2.8 trillion parameters. The model represents the largest openly available model by parameter count, entering what the industry categorizes as the 3T class.
Alibaba releases Qwen 3.8, a 2.4 trillion parameter open-weight model claiming second place behind Fable 5
Alibaba has released Qwen 3.8, a 2.4 trillion parameter open-weight model that the company claims trails only Fable 5. The multimodal model processes images, videos, and documents, with a preview available through Alibaba's platforms at 10 percent of standard pricing.
Comments
Loading...