Ling 3.0 Flash Tops Artificial Analysis Rankings for Open Models Under 124B Parameters
Ant Group's inclusionAI released Ling 3.0 Flash, which scores 38 points on the Artificial Analysis Intelligence Index — the highest of any open model under 124 billion total parameters. The model cuts its hallucination rate from 97 to 44 percent versus its predecessor and ships under an MIT license.
Ling 3.0 Flash claims top spot among small open models
Ant Group's inclusionAI has released Ling 3.0 Flash, an open-weight model that now ranks as the most capable open model under 124 billion total parameters, according to Artificial Analysis benchmark data.
On the Artificial Analysis Intelligence Index, Ling 3.0 Flash scores 38 points — a substantial jump over its predecessor. That score puts it roughly on par with Qwen3.6 27B, despite Ling 3.0 Flash using far fewer active parameters. The current leader among open models, DeepSeek V4 Flash, still scores higher at 52 points.
According to Artificial Analysis, no smaller model currently matches Ling 3.0 Flash's score, making it the top performer specifically within its size class rather than an outright leader across all open models.
Hallucination rate drops sharply
One of the more notable changes is in reliability. On the AA Omniscience test, which measures how often a model produces incorrect answers versus appropriately declining to answer, Ling 3.0 Flash's hallucination rate fell from 97 percent to 44 percent compared to the previous version. The model now refuses to answer questions it lacks reliable information for far more frequently — a shift that trades some raw answer volume for accuracy.
The model also shows gains in agentic task performance over its predecessor, including improved results on the t3-Bench Banking benchmark, which tests multi-step task execution in simulated financial workflows.
Pricing and efficiency tradeoffs
On a per-token basis, Ling 3.0 Flash undercuts every comparably capable model on price, according to Artificial Analysis, though inclusionAI has not disclosed exact per-million-token rates. The model does consume more tokens on complex tasks than some similarly strong alternatives, meaning its efficiency advantage narrows on harder problems. Even so, on a per-task cost basis it reportedly remains cheaper than Qwen3.6 27B.
Availability
inclusionAI is releasing Ling 3.0 Flash under an MIT license, making it usable commercially with minimal restriction. The model is available through the inclusionAI API and via DeepInfra, with model weights published directly on Hugging Face for self-hosting.
No training cutoff date, exact active parameter count, or context window size has been disclosed in the released benchmark materials.
What this means
Ling 3.0 Flash's positioning is narrow but real: it's not challenging DeepSeek V4 Flash for the top spot among open models overall, but it does appear to set a new bar for models under 124 billion total parameters, based on Artificial Analysis's independent testing. The bigger story may be the hallucination fix — cutting the error rate from 97 to 44 percent on AA Omniscience suggests inclusionAI prioritized calibration and refusal behavior over raw benchmark chasing in this release. For teams evaluating smaller open models for cost-sensitive or self-hosted deployments, Ling 3.0 Flash's MIT license and multi-platform availability (Hugging Face, DeepInfra, native API) lower the barrier to testing it directly rather than relying on vendor claims alone.
Related Articles
Qwen3.8 Max Matches Claude Opus 4.8 on Intelligence Index, But Costs 2x More Per Task Than Predecessor
Alibaba's Qwen3.8 Max jumps 10 points to 56 on the Artificial Analysis Intelligence Index, putting it on par with Claude Opus 4.8. But Kimi K3 still edges it out at a lower per-task cost, and Qwen3.8 Max shows a sharp rise in hallucination rate.
Claude Opus 5 Scores 61 on Intelligence Index, Beats Fable 5 on Cost Across Most Benchmarks
Anthropic's Claude Opus 5 posts a 61 on the Artificial Analysis Intelligence Index, narrowly beating Claude Fable 5 (60) and GPT-5.6 Sol (59) while costing less per task. The model leads in coding and knowledge-work benchmarks but shows a rising hallucination rate of 50 percent.
Microsoft's MAI Code 1.1 Flash Loses on Price and Performance to DeepSeek-V4-Flash
Microsoft's MAI Code 1.1 Flash beats its predecessor and mini-models from Anthropic and OpenAI on SWE-bench Verified, but DeepSeek-V4-Flash-0731 outperforms it on Terminal Bench 2.1 (82.7% vs 62.9%) while costing roughly a third as much per token.
xAI's Imagine Image 2.0 Scores Second Place Behind GPT-Image-2 in Arena Benchmarks
xAI's new Imagine Image 2.0 lands in second place globally on both the Image Edit and Text-to-Image Arena leaderboards, trailing OpenAI's GPT-Image-2. The model ships with editing tools including Magic Wand, Multi-Ref Editing, and Smart Resize.
Comments
Loading...