AI Model Benchmarks
778 scores across 9 benchmarks — quality, coding, reasoning, and speed.
Every score we hold, including models older than 12 months. For current picks see the rankings, which use the last 12 months only.
Output throughput in tokens per second measured on standard prompts. Higher = faster responses. Source: artificialanalysis.ai and vendor-published measurements.
Model
tok/sReleased
Quality scores from official model cards, published technical reports, and independent leaderboards (vals.ai, Vectara). Speed benchmarks from artificialanalysis.ai (approximate medians). Higher is better for all metrics except Hallucination Rate, where lower is better.