Best LLM for Coding in 2026

Ranked by composite coding score averaging SWE-bench Verified and LiveCodeBench. All scores from official model cards and technical reports.

Updated automatically as new models release. Full benchmark leaderboard →

1
96.2%
avg
2
95.0%
avg
3
93.9%
avg
4
93.0%
avg
5
88.6%
avg
6
87.6%
avg
7
87.0%
avg
8
86.6%
avg
9
86.0%
avg
10
84.9%
avg
11
84.6%
avg
12
84.5%
avg
13
83.3%
avg
14
81.6%
avg
15
80.7%
avg
16
80.5%
avg
17
80.2%
avg
18
80.2%
avg
19
79.5%
avg
20
79.2%
avg
21
78.8%
avg
22
78.8%
avg
23
78.0%
avg
24
78.0%
avg
25
77.8%
avg
26
77.7%
avg
27
77.6%
avg
28
77.4%
avg
29
77.3%
avg
30
77.2%
avg
31
77.2%
avg
32
76.8%
avg
33
76.4%
avg
34
76.0%
avg
35
76.0%
avg
36
76.0%
avg
37
75.2%
avg
38
74.8%
avg
39
74.6%
avg
40
74.4%
avg
41
74.4%
avg
42
74.4%
avg
43
73.5%
avg
44
73.5%
avg
45
73.4%
avg
46
73.4%
avg
47
73.2%
avg
48
73.2%
avg
49
72.7%
avg
50
72.6%
avg
51
72.4%
avg
52
71.9%
avg
53
70.9%
avg
54
69.2%
avg
55
68.5%
avg
56
68.2%
avg
57
68.2%
avg
58
67.6%
avg
59
64.8%
avg
60
63.2%
avg
61
62.1%
avg
62
62.0%
avg
63
61.6%
avg
64
60.5%
avg
65
59.3%
avg
66
58.0%
avg
67
57.0%
avg
68
56.1%
avg
69
52.8%
avg
70
51.6%
avg
71
47.7%
avg

How we rank coding models

We average three industry-standard benchmarks:

  • SWE-bench Verified — Real GitHub issues from popular open-source repos. Tests autonomous software engineering: read the issue, write a fix, pass the test suite. The most practical real-world coding benchmark available.
  • LiveCodeBench — Competitive programming problems from LeetCode, Codeforces, and AtCoder, collected after model training cutoffs to prevent contamination. Harder than HumanEval.

All scores are from official model cards, technical reports, or the HuggingFace Open LLM Leaderboard. Rankings update automatically as new models are released.

Also see: Best Reasoning LLM, Best Cheap LLM, Compare any two models.