DataLearner logo
Back to Main Leaderboard

LLM Coding Benchmark Leaderboard

This page provides the LLM coding benchmark leaderboard, covering SWE-Bench Verified, SWE-Bench Pro, LiveCodeBench, and SWE-bench Multilingual datasets, comparing GPT, Claude, Qwen, and DeepSeek models.

Updated on 2026-09-09 21:31:23

As of 2026-09, this page covers SWE-bench Verified, LiveCodeBench, SWE-Bench Pro - Public, SWE-bench Multilingual and related benchmarks for LLM Coding Benchmark Leaderboard, making it straightforward to compare within the same task family.

Click any model name to check context length, licensing, and pricing on its detail page. See Data Methodology for scoring details.

Reference: Composite Coding Rankings

There is no single, universally accepted coding leaderboard. Static benchmarks like SWE-bench and HumanEval measure specific skills but can be gamed through targeted fine-tuning. We selected two complementary human-preference leaderboards: LMArena Coding Arena ranks models on general programming tasks (debugging, algorithms, code generation) via anonymous crowd-sourced voting; DesignArena Code Category focuses specifically on visual, front-end code generation (websites, UI components, games) using the same blind-voting methodology. Reading both together gives a fuller picture of coding capability.

LMArena Coding Arena

Full ranking

Elo ratings from anonymous A/B voting on real general coding tasks (debugging, algorithms, code generation) submitted by developers.

Updated 2026-09-02

#ModelElo
1
1552
3
1551
4
Anthropic
Opus 4.7
1547
5
1546
6
Moonshot
kimi-k3-max
1542
7
Google
gemini-3.8-flash-high
1537
9
Z.ai
glm-5.3-flash
1534
10
F
Muse Spark 1.2 (xhigh)
1533
Source: LMArena

DesignArena Code Category

Full ranking

Elo ratings from anonymous voting on visual front-end code tasks (websites, UI components, games, data viz) by Arcada Labs.

Updated 2026-09-08

#ModelElo
1
Moonshot AI
Kimi K3
1394
4
1346
5
Anthropic
Claude Opus 5
1341
7
GLM-5.3
1334
8
F
Muse Spark 1.2
1331
9
1328
10
Google Deep Mind
Gemini 3.7 Flash
1322
Source: DesignArena

LLM Performance Results

Data source: DataLearnerAI
No chart data available

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side.

SWE-bench Verified60.10
LiveCodeBench
SWE-Bench Pro - Public
SWE-bench Multilingual
Free commercial
SWE-bench Verified
LiveCodeBench24.60
SWE-Bench Pro - Public
SWE-bench Multilingual
Free commercial
Sort by: