AI Model Benchmarks

778 scores across 9 benchmarks — quality, coding, reasoning, and speed.

Every score we hold, including models older than 12 months. For current picks see the rankings, which use the last 12 months only.

Output throughput in tokens per second measured on standard prompts. Higher = faster responses. Source: artificialanalysis.ai and vendor-published measurements.

Model
tok/s
1
1007
17
215
20
201
23
192
26
184
27
175
29
163
33
158
34
155
38
145
41
138
42
138
43
136
44
135
45
125
46
125
51
115
52
114
53
114
54
114
55
113
56
112
58
108
59
108
60
106
63
98
65
92
68
87
74
79
76
78
78
76
79
75
84
70
95
62
96
62
99
60
100
60
103
58
104
58
105
57
107
57
110
55
118
43.167
119
42
123
39
124
39
127
32
129
28
130
28
133
22

Quality scores from official model cards, published technical reports, and independent leaderboards (vals.ai, Vectara). Speed benchmarks from artificialanalysis.ai (approximate medians). Higher is better for all metrics except Hallucination Rate, where lower is better.