model releaseInclusionai

InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters

TL;DR

InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.

2 min read
0

Ling 3.0 Flash Fin — Quick Specs

Context window262K tokens
Input$0.06/1M tokens
Output$0.18/1M tokens

InclusionAI Releases Ling 3.0 Flash Fin

InclusionAI has released Ling 3.0 Flash Fin, a finance-focused mixture-of-experts (MoE) model built on its Ling 3.0 Flash architecture. The model activates 5.1 billion parameters out of a 124 billion total parameter pool, using sparse MoE routing to keep inference costs low while maintaining a large overall capacity.

Specs and Pricing

The model ships with a 262,000-token context window, positioning it for long documents such as financial filings, research reports, and multi-step transaction records. Pricing on OpenRouter, via DeepInfra, is set at $0.06 per 1M input tokens and $0.18 per 1M output tokens, with cached input reads priced at $0.012 per 1M tokens. DeepInfra reports a median (P50) latency of 0.52 seconds and throughput of 99 tokens per second for the endpoint, though OpenRouter's own availability tracking shows the endpoint at 26.20% uptime over the last 24 hours and 0.00% over the trailing three days at the time of writing.

According to InclusionAI, Ling 3.0 Flash Fin is designed for "real-world investment workflows" that require complex, multi-step task execution and long-horizon planning. The company states the model retains general-purpose capabilities in reasoning, coding, and mathematics inherited from the base Ling 3.0 Flash model, rather than sacrificing them for domain specialization.

No independent benchmark scores have been published for Ling 3.0 Flash Fin at this time. InclusionAI has not disclosed a training data cutoff date for the model.

Availability

The model is listed on OpenRouter with a listed release date of August 27, 2026, and is currently served through DeepInfra as its primary inference provider. OpenRouter's routing infrastructure allows fallback to alternative providers if the primary endpoint experiences errors, though as of publication only one provider serves this model.

What This Means

Ling 3.0 Flash Fin represents a narrower bet than a general-purpose frontier release: rather than competing on raw benchmark leaderboards, InclusionAI is targeting a vertical — investment and financial workflows — where long context windows and multi-step planning matter more than peak reasoning scores. The 5.1B active-parameter MoE design keeps per-token costs low ($0.06/$0.18 per 1M) relative to dense models of similar total size, which could make it attractive for high-volume financial document processing if reliability improves.

That said, the current uptime figures reported by OpenRouter — near-zero over the past three days — suggest the endpoint is not yet production-stable. Teams evaluating this model for finance applications should treat it as an early-access release and verify benchmark performance independently, since InclusionAI has not published third-party-verifiable scores for reasoning, coding, or domain-specific financial tasks.

Related Articles

model release

Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window

Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.

model release

Inception Launches Mercury 2.5 Preview, a Diffusion LLM Claiming 1,107 Tokens/Sec

Inception released Mercury 2.5 Preview, a diffusion-based language model that generates tokens in parallel rather than sequentially, claiming throughput of 1,107 tokens per second on standard GPUs. The model is available on OpenRouter with a 260K context window and an 80% launch discount through September 8, 2026.

model release

OpenAI Releases Astra, Claims New Flagship Model Beats Rivals on Coding and Cybersecurity Benchmarks

OpenAI released Astra on Thursday, calling it its most capable and most aligned model yet. The model uses a reasoning technique called 'opaque recurrence' that critics say reduces visibility into its chain of thought.

model release

Meta's Muse Spark 1.3 Claims #3 Global Ranking, Matches OpenAI's GPT-5.6-Sol on Coding Benchmarks

Meta Superintelligence Labs shipped Muse Spark 1.3, which the company claims ranks #3 globally on the Artificial Analysis Intelligence Index and matches OpenAI's GPT-5.6-Sol on coding and agentic benchmarks. The model is available now via Muse Code and Meta's API, with open weights and a follow-up model promised soon.

Comments

Loading...