Ling 3.0 Flash Tops Artificial Analysis Rankings for Open Models Under 124B Parameters
Ant Group's inclusionAI released Ling 3.0 Flash, which scores 38 points on the Artificial Analysis Intelligence Index — the highest of any open model under 124 billion total parameters. The model cuts its hallucination rate from 97 to 44 percent versus its predecessor and ships under an MIT license.
Ling 3.0 Flash claims top spot among small open models
Ant Group's inclusionAI has released Ling 3.0 Flash, an open-weight model that now ranks as the most capable open model under 124 billion total parameters, according to Artificial Analysis benchmark data.
On the Artificial Analysis Intelligence Index, Ling 3.0 Flash scores 38 points — a substantial jump over its predecessor. That score puts it roughly on par with Qwen3.6 27B, despite Ling 3.0 Flash using far fewer active parameters. The current leader among open models, DeepSeek V4 Flash, still scores higher at 52 points.
According to Artificial Analysis, no smaller model currently matches Ling 3.0 Flash's score, making it the top performer specifically within its size class rather than an outright leader across all open models.
Hallucination rate drops sharply
One of the more notable changes is in reliability. On the AA Omniscience test, which measures how often a model produces incorrect answers versus appropriately declining to answer, Ling 3.0 Flash's hallucination rate fell from 97 percent to 44 percent compared to the previous version. The model now refuses to answer questions it lacks reliable information for far more frequently — a shift that trades some raw answer volume for accuracy.
The model also shows gains in agentic task performance over its predecessor, including improved results on the t3-Bench Banking benchmark, which tests multi-step task execution in simulated financial workflows.
Pricing and efficiency tradeoffs
On a per-token basis, Ling 3.0 Flash undercuts every comparably capable model on price, according to Artificial Analysis, though inclusionAI has not disclosed exact per-million-token rates. The model does consume more tokens on complex tasks than some similarly strong alternatives, meaning its efficiency advantage narrows on harder problems. Even so, on a per-task cost basis it reportedly remains cheaper than Qwen3.6 27B.
Availability
inclusionAI is releasing Ling 3.0 Flash under an MIT license, making it usable commercially with minimal restriction. The model is available through the inclusionAI API and via DeepInfra, with model weights published directly on Hugging Face for self-hosting.
No training cutoff date, exact active parameter count, or context window size has been disclosed in the released benchmark materials.
What this means
Ling 3.0 Flash's positioning is narrow but real: it's not challenging DeepSeek V4 Flash for the top spot among open models overall, but it does appear to set a new bar for models under 124 billion total parameters, based on Artificial Analysis's independent testing. The bigger story may be the hallucination fix — cutting the error rate from 97 to 44 percent on AA Omniscience suggests inclusionAI prioritized calibration and refusal behavior over raw benchmark chasing in this release. For teams evaluating smaller open models for cost-sensitive or self-hosted deployments, Ling 3.0 Flash's MIT license and multi-platform availability (Hugging Face, DeepInfra, native API) lower the barrier to testing it directly rather than relying on vendor claims alone.
Related Articles
GLM-5.3 Ties Kimi K3 for Top Open-Model Ranking, Undercuts Rivals on Price — But Open Weights Delayed
Z.ai's GLM-5.3 ties Kimi K3 for the top spot among open models on the Artificial Analysis Intelligence Index, driven by a major leap in agentic task performance. The company is delaying the open-weight release by about two weeks, citing the model's unusually strong vulnerability-detection capabilities.
Artificial Analysis Launches Search Index Benchmark for AI Agent Search APIs
Artificial Analysis has released the Search Index, a benchmark measuring how search API providers perform for AI agents across quality, cost, and speed. Parallel, Exa, and Firecrawl lead the initial rankings, with search access boosting model scores from 33 to as high as 75 points.
Artificial Analysis Launches Optima, a Platform to Build Custom AI Benchmarks on Your Own Data
Artificial Analysis has launched Optima, a platform that lets users build custom AI benchmarks using their own data, workflows, or use-case descriptions. Unlike public benchmarks, Optima compares models on cost per task and time per task in addition to quality.
Qwen3.8 Max Matches Claude Opus 4.8 on Intelligence Index, But Costs 2x More Per Task Than Predecessor
Alibaba's Qwen3.8 Max jumps 10 points to 56 on the Artificial Analysis Intelligence Index, putting it on par with Claude Opus 4.8. But Kimi K3 still edges it out at a lower per-task cost, and Qwen3.8 Max shows a sharp rise in hallucination rate.
Comments
Loading...