DataLearner logo

LMArena Math Arena Leaderboard

The latest AI math reasoning leaderboard based on LMArena Math Arena anonymous user voting. Covers Elo scores, confidence intervals, and vote counts for Claude, GPT, Gemini, DeepSeek, Qwen, and more.

Top Model

qwen3.8-max

Top Score

1517.00

Model Count

373

Data version

2026年08月06日

Data source: LM Arena

About This Leaderboard

This leaderboard ranks AI models by mathematical reasoning ability. Data comes from LMArena's Math sub-track, evaluated through anonymous blind testing by real users on math problem-solving tasks.

Methodology Overview

Blind testing: Users submit math problems, two anonymous models provide solutions, and users vote for the better answer — eliminating brand bias.

Elo scoring: Uses the Bradley-Terry model to calculate Elo scores. Higher scores mean users more frequently prefer that model's math solutions.

Broad scenario coverage: Testing spans algebra, geometry, calculus, competition math, and more diverse real-world math tasks.

DataLearner provides in-depth analysis on top of the raw data, linking leaderboard models to the DataLearner model database so you can quickly access model details, API pricing, benchmark scores, and more.

Origin:AllChina
Leaderboard snapshot month:

Ranking Table

RankModelScore95% CIVotesOrganizationLicense
6Alibabaqwen3.8-maxAlibaba1517.00+/-40218AlibabaProprietary
11Moonshotkimi-k3-maxMoonshot1496.00+/-28422MoonshotKimi K3 license
24Moonshot AIKimi K2.6Moonshot AI1479.00+/-141,921Moonshot AIModified MIT
38Moonshot AIKimi K2 ThinkingMoonshot AI1470.00+/-103,781Moonshot AIModified MIT
43DeepSeekdeepseek-v4-pro-high-previewDeepSeek1467.00+/-132,373DeepSeekMIT
63DeepSeek-AIDeepSeek-V4-ProDeepSeek-AI1444.00+/-122,732DeepSeek-AIMIT
68DeepSeekdeepseek-v4-flash-high-previewDeepSeek1441.00+/-122,442DeepSeekMIT
70Moonshot AIKimi K2.5 InstantMoonshot AI1441.00+/-25508Moonshot AIModified MIT
71MiniMaxAIMiniMax M3MiniMaxAI1440.00+/-151,639MiniMaxAIMiniMax Community License
75Moonshot AIKimi K2 Thinking (thinking-turbo)Moonshot AI1437.00+/-103,743Moonshot AIModified MIT
84DeepSeek-AIDeepSeek V3.2-Exp (thinking)DeepSeek-AI1429.00+/-26481DeepSeek-AIMIT
85DeepSeek-AIDeepSeek V3.2DeepSeek-AI1428.00+/-112,978DeepSeek-AIMIT
87Tencenthunyuan-hy3-previewTencent1428.00+/-28399Tencenttencent-hunyuan-community
92Alibabaqwen3-max-2025-09-23Alibaba1427.00+/-24581AlibabaProprietary
93DeepSeek-AIDeepSeek-V4-FlashDeepSeek-AI1426.00+/-122,498DeepSeek-AIMIT
95DeepSeek-AIDeepSeek V3.2-Exp (thinking)DeepSeek-AI1425.00+/-122,477DeepSeek-AIMIT
98MiniMaxAIMiniMax-M2.7MiniMaxAI1424.00+/-112,900MiniMaxAIModified MIT
107DeepSeek-AIDeepSeek V3.2-ExpDeepSeek-AI1418.00+/-21774DeepSeek-AIMIT
109Moonshot AIKimi K2 0905Moonshot AI1417.00+/-21757Moonshot AIModified MIT
112DeepSeek-AIDeepSeek-V3.1DeepSeek-AI1415.00+/-18992DeepSeek-AIMIT
113DeepSeek-AIDeepSeek-V3.1 (thinking)DeepSeek-AI1414.00+/-22664DeepSeek-AIMIT
116DeepSeek-AIDeepSeek-R1DeepSeek-AI1412.00+/-141,606DeepSeek-AIMIT
124StepFunAIStep 3.5 FlashStepFunAI1407.00+/-113,216StepFunAIApache 2.0
128StepFunAIStep 3.5 FlashStepFunAI1404.00+/-113,192StepFunAIProprietary
142Alibabaqwen3-235b-a22b-thinking-2507Alibaba1397.00+/-25486AlibabaApache 2.0
143DeepSeek-AIDeepSeek-R1-0528DeepSeek-AI1396.00+/-20863DeepSeek-AIMIT
144MiniMaxAIMiniMax M2.5MiniMaxAI1396.00+/-122,422MiniMaxAIModified MIT
145DeepSeek-AIDeepSeek-V3.1 TerminusDeepSeek-AI1396.00+/-39217DeepSeek-AIMIT
147Alibabaqwen3-235b-a22b-no-thinkingAlibaba1393.00+/-122,381AlibabaApache 2.0
149MiniMaxAIM2.1MiniMaxAI1390.00+/-19988MiniMaxAIMIT
152Moonshot AIKimi K2Moonshot AI1389.00+/-141,689Moonshot AIModified MIT
170MiniMaxminimax-m1MiniMax1372.00+/-131,789MiniMaxApache 2.0
171DeepSeek-AIDeepSeek-V3-0324DeepSeek-AI1370.00+/-103,183DeepSeek-AIMIT
177StepFunAIStep3StepFunAI1363.00+/-31353StepFunAIApache 2.0
185MiniMaxAIMiniMax M2MiniMaxAI1354.00+/-33320MiniMaxAIApache 2.0
191Tencenthunyuan-turbos-20250416Tencent1347.00+/-20845TencentProprietary
200Alibabaqwen-plus-0125Alibaba1323.00+/-19732AlibabaProprietary
207StepFunstep-2-16k-exp-202412StepFun1312.00+/-20642StepFunProprietary
211DeepSeek-AIDeepSeek-V3DeepSeek-AI1311.00+/-112,721DeepSeek-AIDeepSeek
219Alibabaqwen2.5-plus-1127Alibaba1304.00+/-141,404AlibabaProprietary
221Tencenthunyuan-turbos-20250226Tencent1301.00+/-31238TencentProprietary
224glm-4-plus-0111Zhipu1298.00+/-19721ZhipuProprietary
225StepFunstep-1o-turbo-202506StepFun1298.00+/-24564StepFunProprietary
231Tencenthunyuan-large-2025-02-10Tencent1294.00+/-24497TencentProprietary
232DeepSeekdeepseek-v2.5-1210DeepSeek1292.00+/-171,031DeepSeekDeepSeek
233Alibabaqwen-max-0919Alibaba1291.00+/-122,249AlibabaQwen
234Tencenthunyuan-standard-2025-02-10Tencent1290.00+/-24499TencentProprietary
237DeepSeek-AIDeepSeek V2.5DeepSeek-AI1288.00+/-103,649DeepSeek-AIDeepSeek
238glm-4-plusZhipu AI1287.00+/-103,599Zhipu AIProprietary
243Tencenthunyuan-large-visionTencent1281.00+/-30350TencentProprietary
244Tencenthunyuan-turbo-0110Tencent1279.00+/-31243TencentProprietary
253DeepSeekdeepseek-coder-v2DeepSeek1272.00+/-141,858DeepSeekDeepSeek License
268Tencenthunyuan-standard-256kTencent1250.00+/-29361TencentProprietary
300Alibabaqwen1.5-32b-chatAlibaba1201.00+/-122,649AlibabaQianwen LICENSE
325DeepSeek-AIDeepSeek LLM 67B ChatDeepSeek-AI1155.00+/-24576DeepSeek-AIDeepSeek License

Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.

FAQ

01

What is LMArena Math Arena?

LMArena Math Arena is an anonymous evaluation track focused on mathematical reasoning. Users submit real math questions, compare hidden model solutions side by side, and vote for the better answer; the leaderboard is then calculated with Elo-style scoring.

02

How is Math Arena different from MATH-500 or AIME?

Static benchmarks such as MATH-500 and AIME use fixed problem sets and automated grading. Math Arena uses open-ended user questions and human preference voting, making it a useful complement for measuring how models handle varied real-world math tasks.

03

Do thinking models perform better in Math Arena?

Models with extended reasoning or chain-of-thought style capabilities often rank higher on math tasks because they spend more time decomposing and checking solutions. That benefit can come with higher latency and cost.

04

How do China-developed models perform in math?

DeepSeek, Qwen, GLM, and related models have become competitive in math reasoning leaderboards. Open licenses and Chinese-language support can make them especially useful for local deployment and education scenarios.