DataLearner logo

LMArena Coding Arena Leaderboard

The latest AI coding model leaderboard based on LMArena Coding Arena anonymous user voting. Covers Elo scores, confidence intervals, and vote counts for Claude, GPT, Gemini, DeepSeek, Qwen, and more.

Top Model

kimi-k3-max

Top Score

1542.00

Model Count

380

Data version

2026年08月06日

Data source: LM Arena

About This Leaderboard

This leaderboard ranks AI models by coding ability. Data comes from LMArena (formerly LMSYS Chatbot Arena)'s Coding sub-track, evaluated through anonymous blind testing by real users on programming tasks.

Methodology Overview

Blind testing: Users submit coding questions, two anonymous models generate code answers, and users vote for the better response — eliminating brand bias.

Elo scoring: Uses the Bradley-Terry model to calculate Elo scores. Higher scores mean users more frequently prefer that model's code solutions.

Broad scenario coverage: Testing spans code generation, bug fixing, algorithm implementation, code explanation, and more real-world programming scenarios.

DataLearner provides in-depth analysis on top of the raw data, linking leaderboard models to the DataLearner model database so you can quickly access model details, API pricing, benchmark scores, and more.

Origin:AllChina
Leaderboard snapshot month:

Ranking Table

RankModelScore95% CIVotesOrganizationLicense
6Moonshotkimi-k3-maxMoonshot1542.00+/-122,611MoonshotKimi K3 license
8Alibabaqwen3.8-maxAlibaba1532.00+/-161,411AlibabaProprietary
32Moonshot AIKimi K2.6Moonshot AI1515.00+/-710,310Moonshot AIModified MIT
52Moonshot AIKimi K2.5 InstantMoonshot AI1504.00+/-141,797Moonshot AIModified MIT
55Moonshot AIKimi K2 ThinkingMoonshot AI1502.00+/-518,836Moonshot AIModified MIT
56DeepSeek-AIDeepSeek-V4-ProDeepSeek-AI1501.00+/-615,292DeepSeek-AIMIT
58MiniMaxAIMiniMax M3MiniMaxAI1499.00+/-710,050MiniMaxAIMiniMax Community License
75DeepSeekdeepseek-v4-pro-high-previewDeepSeek1489.00+/-614,268DeepSeekMIT
79Moonshot AIKimi K2 Thinking (thinking-turbo)Moonshot AI1486.00+/-614,758Moonshot AIModified MIT
82DeepSeek-AIDeepSeek-V4-FlashDeepSeek-AI1483.00+/-614,249DeepSeek-AIMIT
86DeepSeekdeepseek-v4-flash-high-previewDeepSeek1479.00+/-614,203DeepSeekMIT
90MiniMaxAIMiniMax-M2.7MiniMaxAI1479.00+/-616,121MiniMaxAIModified MIT
91DeepSeek-AIDeepSeek V3.2-Exp (thinking)DeepSeek-AI1475.00+/-78,493DeepSeek-AIMIT
92DeepSeek-AIDeepSeek V3.2-Exp (thinking)DeepSeek-AI1475.00+/-131,914DeepSeek-AIMIT
95Alibabaqwen3-max-2025-09-23Alibaba1474.00+/-132,039AlibabaProprietary
100DeepSeek-AIDeepSeek V3.2DeepSeek-AI1470.00+/-610,556DeepSeek-AIMIT
104Moonshot AIKimi K2 0905Moonshot AI1468.00+/-132,240Moonshot AIModified MIT
106DeepSeek-AIDeepSeek V3.2-ExpDeepSeek-AI1465.00+/-122,491DeepSeek-AIMIT
108DeepSeek-AIDeepSeek-R1-0528DeepSeek-AI1464.00+/-112,725DeepSeek-AIMIT
111DeepSeek-AIDeepSeek-V3.1 Terminus (thinking)DeepSeek-AI1463.00+/-24635DeepSeek-AIMIT
114Moonshot AIKimi K2Moonshot AI1460.00+/-85,235Moonshot AIModified MIT
115Tencenthunyuan-hy3-previewTencent1460.00+/-141,942Tencenttencent-hunyuan-community
122DeepSeek-AIDeepSeek-V3.1 (thinking)DeepSeek-AI1457.00+/-131,902DeepSeek-AIMIT
129StepFunAIStep 3.5 FlashStepFunAI1451.00+/-614,275StepFunAIApache 2.0
132DeepSeek-AIDeepSeek-V3.1DeepSeek-AI1448.00+/-122,624DeepSeek-AIMIT
134Alibabaqwen3-235b-a22b-no-thinkingAlibaba1446.00+/-86,967AlibabaApache 2.0
136DeepSeek-AIDeepSeek-R1DeepSeek-AI1445.00+/-122,317DeepSeek-AIMIT
137MiniMaxAIMiniMax M2.5MiniMaxAI1444.00+/-710,772MiniMaxAIModified MIT
140Alibabaqwen3-235b-a22b-thinking-2507Alibaba1442.00+/-151,612AlibabaApache 2.0
141MiniMaxAIM2.1MiniMaxAI1439.00+/-103,409MiniMaxAIMIT
143DeepSeek-AIDeepSeek-V3.1 TerminusDeepSeek-AI1439.00+/-21778DeepSeek-AIMIT
145Tencenthunyuan-vision-1.5-thinkingTencent1437.00+/-27436TencentProprietary
146StepFunAIStep 3.5 FlashStepFunAI1437.00+/-616,519StepFunAIProprietary
161DeepSeek-AIDeepSeek-V3-0324DeepSeek-AI1429.00+/-78,358DeepSeek-AIMIT
169MiniMaxminimax-m1MiniMax1416.00+/-86,471MiniMaxApache 2.0
178StepFunAIStep3StepFunAI1408.00+/-171,231StepFunAIApache 2.0
183Tencenthunyuan-turbos-20250226Tencent1400.00+/-31275TencentProprietary
189Tencenthunyuan-turbos-20250416Tencent1394.00+/-141,776TencentProprietary
196DeepSeek-AIDeepSeek-V3DeepSeek-AI1388.00+/-103,280DeepSeek-AIDeepSeek
202MiniMaxAIMiniMax M2MiniMaxAI1385.00+/-151,545MiniMaxAIApache 2.0
207Alibabaqwen-plus-0125Alibaba1380.00+/-18893AlibabaProprietary
209DeepSeekdeepseek-v2.5-1210DeepSeek1375.00+/-171,079DeepSeekDeepSeek
212Tencenthunyuan-turbo-0110Tencent1371.00+/-30299TencentProprietary
213StepFunstep-2-16k-exp-202412StepFun1371.00+/-20737StepFunProprietary
219DeepSeek-AIDeepSeek V2.5DeepSeek-AI1368.00+/-94,252DeepSeek-AIDeepSeek
221Tencenthunyuan-large-2025-02-10Tencent1367.00+/-25519TencentProprietary
231Alibabaqwen2.5-plus-1127Alibaba1357.00+/-141,553AlibabaProprietary
233Tencenthunyuan-large-visionTencent1356.00+/-19963TencentProprietary
237StepFunstep-1o-turbo-202506StepFun1353.00+/-151,505StepFunProprietary
238Alibabaqwen-max-0919Alibaba1353.00+/-112,756AlibabaQwen
240glm-4-plusZhipu AI1352.00+/-94,449Zhipu AIProprietary
249DeepSeekdeepseek-coder-v2DeepSeek1342.00+/-122,671DeepSeekDeepSeek License
256Tencenthunyuan-standard-2025-02-10Tencent1332.00+/-24549TencentProprietary
258glm-4-plus-0111Zhipu1330.00+/-18894ZhipuProprietary
280Tencenthunyuan-standard-256kTencent1301.00+/-24497TencentProprietary
305Alibabaqwen1.5-32b-chatAlibaba1262.00+/-113,930AlibabaQianwen LICENSE
325DeepSeek-AIDeepSeek LLM 67B ChatDeepSeek-AI1218.00+/-24649DeepSeek-AIDeepSeek License

Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.

FAQ

01

What is LMArena Coding Arena?

LMArena Coding Arena is an anonymous evaluation track focused on coding ability. Users submit real programming tasks such as debugging, code generation, and algorithm implementation; two hidden model answers are shown side by side, and user votes are aggregated into an Elo leaderboard.

02

How is Coding Arena different from SWE-bench or HumanEval?

Static benchmarks use fixed test sets and automated scoring, which makes them reproducible but easier to over-optimize for. Coding Arena uses open-ended user tasks and human preference votes, so it better reflects practical coding experience. The two approaches are complementary.

03

How do China-developed models perform on coding tasks?

Models such as DeepSeek and Qwen rank competitively on coding leaderboards. They are especially relevant when open deployment, Chinese-language developer workflows, or cost control matter.

04

How can AI help with day-to-day programming?

Common workflows include code completion and generation, debugging, code review, unit test generation, and cross-language translation.