Best LLM for Coding in 2026

Ranked by composite coding score averaging SWE-bench Verified and LiveCodeBench. All scores from official model cards and technical reports.

Updated automatically as new models release. Full benchmark leaderboard →

1
96.2%
avg
2
95.0%
avg
3
93.9%
avg
4
93.0%
avg
5
88.6%
avg
6
87.6%
avg
7
87.0%
avg
8
86.6%
avg
9
86.0%
avg
10
84.9%
avg
11
84.6%
avg
12
84.5%
avg
13
83.3%
avg
14
81.6%
avg
15
80.7%
avg
16
80.5%
avg
17
80.2%
avg
18
79.5%
avg
19
78.8%
avg
20
78.8%
avg
21
78.0%
avg
22
78.0%
avg
23
77.8%
avg
24
77.7%
avg
25
77.6%
avg
26
77.4%
avg
27
77.3%
avg
28
77.2%
avg
29
77.2%
avg
30
76.8%
avg
31
76.4%
avg
32
76.0%
avg
33
76.0%
avg
34
75.2%
avg
35
74.8%
avg
36
74.6%
avg
37
74.4%
avg
38
74.4%
avg
39
74.4%
avg
40
73.5%
avg
41
73.5%
avg
42
73.4%
avg
43
73.4%
avg
44
73.2%
avg
45
73.2%
avg
46
72.7%
avg
47
72.4%
avg
48
71.9%
avg
49
70.9%
avg
50
69.2%
avg
51
68.5%
avg
52
68.2%
avg
53
68.2%
avg
54
67.6%
avg
55
64.8%
avg
56
63.2%
avg
57
62.1%
avg
58
62.0%
avg
59
61.6%
avg
60
60.5%
avg
61
59.3%
avg
62
58.0%
avg
63
56.1%
avg

How we rank coding models

We average three industry-standard benchmarks:

  • SWE-bench Verified — Real GitHub issues from popular open-source repos. Tests autonomous software engineering: read the issue, write a fix, pass the test suite. The most practical real-world coding benchmark available.
  • LiveCodeBench — Competitive programming problems from LeetCode, Codeforces, and AtCoder, collected after model training cutoffs to prevent contamination. Harder than HumanEval.

All scores are from official model cards, technical reports, or the HuggingFace Open LLM Leaderboard. Rankings update automatically as new models are released.

Also see: Best Reasoning LLM, Best Cheap LLM, Compare any two models.