InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters
InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.
Ling 3.0 Flash Fin — Quick Specs
InclusionAI Releases Ling 3.0 Flash Fin
InclusionAI has released Ling 3.0 Flash Fin, a finance-focused mixture-of-experts (MoE) model built on its Ling 3.0 Flash architecture. The model activates 5.1 billion parameters out of a 124 billion total parameter pool, using sparse MoE routing to keep inference costs low while maintaining a large overall capacity.
Specs and Pricing
The model ships with a 262,000-token context window, positioning it for long documents such as financial filings, research reports, and multi-step transaction records. Pricing on OpenRouter, via DeepInfra, is set at $0.06 per 1M input tokens and $0.18 per 1M output tokens, with cached input reads priced at $0.012 per 1M tokens. DeepInfra reports a median (P50) latency of 0.52 seconds and throughput of 99 tokens per second for the endpoint, though OpenRouter's own availability tracking shows the endpoint at 26.20% uptime over the last 24 hours and 0.00% over the trailing three days at the time of writing.
According to InclusionAI, Ling 3.0 Flash Fin is designed for "real-world investment workflows" that require complex, multi-step task execution and long-horizon planning. The company states the model retains general-purpose capabilities in reasoning, coding, and mathematics inherited from the base Ling 3.0 Flash model, rather than sacrificing them for domain specialization.
No independent benchmark scores have been published for Ling 3.0 Flash Fin at this time. InclusionAI has not disclosed a training data cutoff date for the model.
Availability
The model is listed on OpenRouter with a listed release date of August 27, 2026, and is currently served through DeepInfra as its primary inference provider. OpenRouter's routing infrastructure allows fallback to alternative providers if the primary endpoint experiences errors, though as of publication only one provider serves this model.
What This Means
Ling 3.0 Flash Fin represents a narrower bet than a general-purpose frontier release: rather than competing on raw benchmark leaderboards, InclusionAI is targeting a vertical — investment and financial workflows — where long context windows and multi-step planning matter more than peak reasoning scores. The 5.1B active-parameter MoE design keeps per-token costs low ($0.06/$0.18 per 1M) relative to dense models of similar total size, which could make it attractive for high-volume financial document processing if reliability improves.
That said, the current uptime figures reported by OpenRouter — near-zero over the past three days — suggest the endpoint is not yet production-stable. Teams evaluating this model for finance applications should treat it as an early-access release and verify benchmark performance independently, since InclusionAI has not published third-party-verifiable scores for reasoning, coding, or domain-specific financial tasks.
Related Articles
Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window
Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.
Inception Launches Mercury 2.5 Preview, a Diffusion LLM Claiming 1,107 Tokens/Sec
Inception released Mercury 2.5 Preview, a diffusion-based language model that generates tokens in parallel rather than sequentially, claiming throughput of 1,107 tokens per second on standard GPUs. The model is available on OpenRouter with a 260K context window and an 80% launch discount through September 8, 2026.
OpenAI Releases Astra, Claims New Flagship Model Beats Rivals on Coding and Cybersecurity Benchmarks
OpenAI released Astra on Thursday, calling it its most capable and most aligned model yet. The model uses a reasoning technique called 'opaque recurrence' that critics say reduces visibility into its chain of thought.
Meta's Muse Spark 1.3 Claims #3 Global Ranking, Matches OpenAI's GPT-5.6-Sol on Coding Benchmarks
Meta Superintelligence Labs shipped Muse Spark 1.3, which the company claims ranks #3 globally on the Artificial Analysis Intelligence Index and matches OpenAI's GPT-5.6-Sol on coding and agentic benchmarks. The model is available now via Muse Code and Meta's API, with open weights and a follow-up model promised soon.
Comments
Loading...