LMArena Math Arena Leaderboard
The latest AI math reasoning leaderboard based on LMArena Math Arena anonymous user voting. Covers Elo scores, confidence intervals, and vote counts for Claude, GPT, Gemini, DeepSeek, Qwen, and more.
Top Model
claude-opus-5-max
Top Score
1555.00
Model Count
373
Data version
2026年08月06日
Data source: LM Arena
About This Leaderboard
This leaderboard ranks AI models by mathematical reasoning ability. Data comes from LMArena's Math sub-track, evaluated through anonymous blind testing by real users on math problem-solving tasks.
Methodology Overview
Blind testing: Users submit math problems, two anonymous models provide solutions, and users vote for the better answer — eliminating brand bias.
Elo scoring: Uses the Bradley-Terry model to calculate Elo scores. Higher scores mean users more frequently prefer that model's math solutions.
Broad scenario coverage: Testing spans algebra, geometry, calculus, competition math, and more diverse real-world math tasks.
DataLearner provides in-depth analysis on top of the raw data, linking leaderboard models to the DataLearner model database so you can quickly access model details, API pricing, benchmark scores, and more.
Ranking Table
| Rank | Model | Score | 95% CI | Votes | Organization | License |
|---|---|---|---|---|---|---|
claude-opus-5-maxAnthropic | 1555.00 | +/-32 | 329 | Anthropic | Proprietary | |
Claude Fable 5Anthropic | 1537.00 | +/-19 | 948 | Anthropic | Proprietary | |
claude-opus-5-highAnthropic | 1529.00 | +/-23 | 700 | Anthropic | Proprietary | |
| 4 | gemini-3.6-flashGoogle | 1521.00 | +/-26 | 529 | Proprietary | |
| 5 | Claude Opus 4.6 (thinking)Anthropic | 1517.00 | +/-11 | 3,574 | Anthropic | Proprietary |
| 6 | qwen3.8-maxAlibaba | 1517.00 | +/-40 | 218 | Alibaba | Proprietary |
| 7 | gemini-3.5-flash-highGoogle | 1506.00 | +/-18 | 1,161 | Proprietary | |
| 8 | Claude Opus 4.6Anthropic | 1505.00 | +/-10 | 3,965 | Anthropic | Proprietary |
| 9 | Opus 4.7 (thinking)Anthropic | 1504.00 | +/-12 | 2,876 | Anthropic | Proprietary |
| 10 | GPT-5.5OpenAI | 1499.00 | +/-12 | 2,722 | OpenAI | Proprietary |
| 11 | kimi-k3-maxMoonshot | 1496.00 | +/-28 | 422 | Moonshot | Kimi K3 license |
| 12 | GPT-5.4 (high)OpenAI | 1494.00 | +/-11 | 3,166 | OpenAI | Proprietary |
| 13 | GPT-5.5 (high)OpenAI | 1492.00 | +/-12 | 2,676 | OpenAI | Proprietary |
| 14 | Gemini 3.1 Pro PreviewGoogle Deep Mind | 1492.00 | +/-9 | 4,903 | Google Deep Mind | Proprietary |
| 15 | Claude Opus 4.8 (thinking)Anthropic | 1491.00 | +/-14 | 1,816 | Anthropic | Proprietary |
| 16 | Opus 4.7Anthropic | 1490.00 | +/-11 | 3,003 | Anthropic | Proprietary |
| 17 | Qwen3.7-Max-Preview阿里巴巴 | 1490.00 | +/-40 | 218 | 阿里巴巴 | Proprietary |
| 18 | Muse Spark 1.1Facebook AI研究实验室 | 1487.00 | +/-23 | 639 | Facebook AI研究实验室 | Proprietary |
| 19 | gpt-5.6-sol-xhighOpenAI | 1484.00 | +/-24 | 597 | OpenAI | Proprietary |
| 20 | GLM 5.1智谱AI | 1483.00 | +/-14 | 1,875 | 智谱AI | MIT |
| 21 | InklingThinking Machines Lab | 1482.00 | +/-25 | 546 | Thinking Machines Lab | Apache 2.0 |
| 22 | gemini-3.5-flash-mediumGoogle | 1482.00 | +/-19 | 983 | Proprietary | |
| 23 | gpt-5.6-terra-xhighOpenAI | 1480.00 | +/-23 | 609 | OpenAI | Proprietary |
| 24 | Kimi K2.6Moonshot AI | 1479.00 | +/-14 | 1,921 | Moonshot AI | Modified MIT |
| 25 | Hy3腾讯AI实验室 | 1479.00 | +/-42 | 201 | 腾讯AI实验室 | Apache 2.0 |
| 26 | 1479.00 | +/-23 | 633 | xAI | Proprietary | |
| 27 | Gemini 3 ProGoogle Deep Mind | 1478.00 | +/-11 | 2,640 | Google Deep Mind | Proprietary |
| 28 | ERNIE-5.1-Preview百度 | 1476.00 | +/-14 | 1,877 | 百度 | Proprietary |
| 29 | mimo-v2.5-proXiaomi | 1476.00 | +/-13 | 2,316 | Xiaomi | MIT |
| 30 | Gemini 3.0 FlashGoogle Deep Mind | 1476.00 | +/-13 | 1,982 | Google Deep Mind | Proprietary |
| 31 | Claude Opus 4.8Anthropic | 1475.00 | +/-14 | 1,850 | Anthropic | Proprietary |
| 32 | glm-5.2-maxZ.ai | 1475.00 | +/-17 | 1,160 | Z.ai | MIT |
| 33 | Qwen3.6-Max-Preview阿里巴巴 | 1473.00 | +/-30 | 356 | 阿里巴巴 | Proprietary |
| 34 | Claude Sonnet 4.5 (high)Anthropic | 1472.00 | +/-19 | 924 | Anthropic | Proprietary |
| 35 | gpt-5.6-luna-xhighOpenAI | 1472.00 | +/-23 | 634 | OpenAI | Proprietary |
| 36 | Qwen3.7-Plus阿里巴巴 | 1471.00 | +/-17 | 1,292 | 阿里巴巴 | Proprietary |
| 37 | Gemma 4 31BDeepMind | 1471.00 | +/-27 | 398 | DeepMind | Apache 2.0 |
| 38 | Kimi K2 ThinkingMoonshot AI | 1470.00 | +/-10 | 3,781 | Moonshot AI | Modified MIT |
| 39 | Claude Opus 4 (thinking-32k)Anthropic | 1469.00 | +/-12 | 2,246 | Anthropic | Proprietary |
| 40 | Qwen3.5 Max Preview阿里巴巴 | 1469.00 | +/-16 | 1,335 | 阿里巴巴 | Proprietary |
| 41 | Gemma 4 26B A4BDeepMind | 1468.00 | +/-28 | 370 | DeepMind | Apache 2.0 |
| 42 | 1467.00 | +/-11 | 3,304 | xAI | Proprietary | |
| 43 | deepseek-v4-pro-high-previewDeepSeek | 1467.00 | +/-13 | 2,373 | DeepSeek | MIT |
| 44 | Claude Opus 4Anthropic | 1465.00 | +/-9 | 4,292 | Anthropic | Proprietary |
| 45 | Claude Sonnet 4.6Anthropic | 1464.00 | +/-11 | 3,439 | Anthropic | Proprietary |
| 46 | GPT-5.5 InstantOpenAI | 1463.00 | +/-16 | 1,451 | OpenAI | Proprietary |
| 47 | Muse SparkFacebook AI研究实验室 | 1462.00 | +/-20 | 856 | Facebook AI研究实验室 | Proprietary |
| 48 | GPT-5.4OpenAI | 1461.00 | +/-11 | 3,334 | OpenAI | Proprietary |
| 49 | GPT-5.2 Pro (high)OpenAI | 1458.00 | +/-11 | 2,957 | OpenAI | Proprietary |
| 50 | GPT-5.1 Pro (high)OpenAI | 1455.00 | +/-12 | 2,479 | OpenAI | Proprietary |
| 51 | Qwen 3.6 Plus Preview阿里巴巴 | 1455.00 | +/-12 | 2,513 | 阿里巴巴 | Proprietary |
| 52 | Claude Sonnet 4.5 (thinking-32k)Anthropic | 1455.00 | +/-9 | 4,870 | Anthropic | Proprietary |
| 53 | GPT-5.2 Chat (0210)OpenAI | 1454.00 | +/-13 | 2,066 | OpenAI | Proprietary |
| 54 | Gemini 3.0 Flash (minimal)Google Deep Mind | 1453.00 | +/-9 | 4,666 | Google Deep Mind | Proprietary |
| 55 | 1453.00 | +/-11 | 3,207 | xAI | Proprietary | |
| 56 | mimo-v2-proXiaomi | 1452.00 | +/-15 | 1,606 | Xiaomi | Proprietary |
| 57 | 1452.00 | +/-15 | 1,591 | xAI | Proprietary | |
| 58 | DOLA Seed 2.0 Pro字节跳动Seed团队 | 1451.00 | +/-10 | 4,054 | 字节跳动Seed团队 | Proprietary |
| 59 | Qwen3.5-397B-A17B阿里巴巴 | 1449.00 | +/-10 | 3,622 | 阿里巴巴 | Apache 2.0 |
| 60 | OpenAI o3OpenAI | 1447.00 | +/-10 | 3,721 | OpenAI | Proprietary |
| 61 | Nemotron 3 UltraNVIDIA | 1445.00 | +/-25 | 541 | NVIDIA | OpenMDW-1.1 |
| 62 | GLM-5V-Turbo智谱AI | 1444.00 | +/-26 | 447 | 智谱AI | Proprietary |
| 63 | DeepSeek-V4-ProDeepSeek-AI | 1444.00 | +/-12 | 2,732 | DeepSeek-AI | MIT |
| 64 | Opus 4.1 (thinking-16k)Anthropic | 1444.00 | +/-11 | 3,025 | Anthropic | Proprietary |
| 65 | 1443.00 | +/-10 | 3,795 | xAI | Proprietary | |
| 66 | GLM-5智谱AI | 1443.00 | +/-15 | 1,599 | 智谱AI | MIT |
| 67 | mimo-v2.5Xiaomi | 1441.00 | +/-13 | 2,318 | Xiaomi | MIT |
| 68 | deepseek-v4-flash-high-previewDeepSeek | 1441.00 | +/-12 | 2,442 | DeepSeek | MIT |
| 69 | Gemini 2.5 ProGoogle Deep Mind | 1441.00 | +/-7 | 7,613 | Google Deep Mind | Proprietary |
| 70 | Kimi K2.5 InstantMoonshot AI | 1441.00 | +/-25 | 508 | Moonshot AI | Modified MIT |
| 71 | 1440.00 | +/-15 | 1,639 | MiniMaxAI | MiniMax Community License | |
| 72 | GPT-5.4 mini (high)OpenAI | 1439.00 | +/-11 | 3,155 | OpenAI | Proprietary |
| 73 | Gemini 3.1 Flash-LiteGoogle Deep Mind | 1439.00 | +/-10 | 3,386 | Google Deep Mind | Proprietary |
| 74 | Qwen3 Max (Preview)阿里巴巴 | 1438.00 | +/-15 | 1,516 | 阿里巴巴 | Proprietary |
| 75 | Kimi K2 Thinking (thinking-turbo)Moonshot AI | 1437.00 | +/-10 | 3,743 | Moonshot AI | Modified MIT |
| 76 | ERNIE 5.0 (0110)百度 | 1437.00 | +/-13 | 2,120 | 百度 | Proprietary |
| 77 | mimo-v2-omniXiaomi | 1436.00 | +/-20 | 904 | Xiaomi | Proprietary |
| 78 | longcat-flash-chat-2602-expMeituan | 1435.00 | +/-14 | 1,732 | Meituan | Proprietary |
| 79 | GPT-5-Pro (high)OpenAI | 1434.00 | +/-14 | 1,883 | OpenAI | Proprietary |
| 80 | Opus 4.1Anthropic | 1433.00 | +/-9 | 4,705 | Anthropic | Proprietary |
| 81 | GPT-5.2OpenAI | 1433.00 | +/-9 | 4,346 | OpenAI | Proprietary |
| 82 | Mistral Medium 3.5MistralAI | 1431.00 | +/-25 | 537 | MistralAI | Modified MIT |
| 83 | gemini-3.5-flash-liteGoogle | 1430.00 | +/-27 | 453 | Proprietary | |
| 84 | DeepSeek V3.2-Exp (thinking)DeepSeek-AI | 1429.00 | +/-26 | 481 | DeepSeek-AI | MIT |
| 85 | DeepSeek V3.2DeepSeek-AI | 1428.00 | +/-11 | 2,978 | DeepSeek-AI | MIT |
| 86 | 1428.00 | +/-9 | 4,195 | xAI | Proprietary | |
| 87 | hunyuan-hy3-previewTencent | 1428.00 | +/-28 | 399 | Tencent | tencent-hunyuan-community |
| 88 | Claude Sonnet 4.5Anthropic | 1428.00 | +/-9 | 4,877 | Anthropic | Proprietary |
| 89 | Qwen3.5-27B阿里巴巴 | 1428.00 | +/-15 | 1,634 | 阿里巴巴 | Apache 2.0 |
| 90 | 1427.00 | +/-12 | 2,256 | xAI | Proprietary | |
| 91 | GLM-4.7智谱AI | 1427.00 | +/-21 | 695 | 智谱AI | MIT |
| 92 | qwen3-max-2025-09-23Alibaba | 1427.00 | +/-24 | 581 | Alibaba | Proprietary |
| 93 | DeepSeek-V4-FlashDeepSeek-AI | 1426.00 | +/-12 | 2,498 | DeepSeek-AI | MIT |
| 94 | GPT-5.3 ChatOpenAI | 1426.00 | +/-14 | 2,018 | OpenAI | Proprietary |
| 95 | DeepSeek V3.2-Exp (thinking)DeepSeek-AI | 1425.00 | +/-12 | 2,477 | DeepSeek-AI | MIT |
| 96 | amazon-nova-experimental-chat-26-02-10Amazon | 1424.00 | +/-39 | 207 | Amazon | Proprietary |
| 97 | GPT-5.4 nano (high)OpenAI | 1424.00 | +/-11 | 3,067 | OpenAI | Proprietary |
| 98 | 1424.00 | +/-11 | 2,900 | MiniMaxAI | Modified MIT | |
| 99 | 1423.00 | +/-29 | 397 | xAI | Proprietary | |
| 100 | Qwen3.5-122B-A10B阿里巴巴 | 1423.00 | +/-14 | 1,763 | 阿里巴巴 | Apache 2.0 |
| 101 | GPT-5.1OpenAI | 1422.00 | +/-11 | 2,846 | OpenAI | Proprietary |
| 102 | GLM-4.6智谱AI | 1421.00 | +/-13 | 2,104 | 智谱AI | MIT |
| 103 | Claude Opus 4 (thinking-16k)Anthropic | 1420.00 | +/-12 | 2,238 | Anthropic | Proprietary |
| 104 | 1420.00 | +/-12 | 2,589 | xAI | Proprietary | |
| 105 | Qwen3-235B-A22B-2507阿里巴巴 | 1418.00 | +/-8 | 5,887 | 阿里巴巴 | Apache 2.0 |
| 106 | 1418.00 | +/-10 | 3,462 | xAI | Proprietary | |
| 107 | DeepSeek V3.2-ExpDeepSeek-AI | 1418.00 | +/-21 | 774 | DeepSeek-AI | MIT |
| 108 | Qwen3-Next阿里巴巴 | 1417.00 | +/-17 | 1,209 | 阿里巴巴 | Apache 2.0 |
| 109 | Kimi K2 0905Moonshot AI | 1417.00 | +/-21 | 757 | Moonshot AI | Modified MIT |
| 110 | OpenAI o4 - miniOpenAI | 1415.00 | +/-11 | 2,933 | OpenAI | Proprietary |
| 111 | longcat-flash-chatMeituan | 1415.00 | +/-22 | 687 | Meituan | MIT |
| 112 | DeepSeek-V3.1DeepSeek-AI | 1415.00 | +/-18 | 992 | DeepSeek-AI | MIT |
| 113 | DeepSeek-V3.1 (thinking)DeepSeek-AI | 1414.00 | +/-22 | 664 | DeepSeek-AI | MIT |
| 114 | GPT-5.2 Chat (0210)OpenAI | 1413.00 | +/-14 | 1,783 | OpenAI | Proprietary |
| 115 | GLM-4.5智谱AI | 1413.00 | +/-16 | 1,420 | 智谱AI | MIT |
| 116 | DeepSeek-R1DeepSeek-AI | 1412.00 | +/-14 | 1,606 | DeepSeek-AI | MIT |
| 117 | Gemini 2.5 Flash-Preview-09-2025Google Deep Mind | 1411.00 | +/-13 | 1,942 | Google Deep Mind | Proprietary |
| 118 | 1411.00 | +/-18 | 1,076 | xAI | Proprietary | |
| 119 | Qwen3-VL-235B-A22B-Instruct阿里巴巴 | 1410.00 | +/-23 | 702 | 阿里巴巴 | Apache 2.0 |
| 120 | OpenAI o1OpenAI | 1409.00 | +/-11 | 2,986 | OpenAI | Proprietary |
| 121 | GPT-4.5OpenAI | 1408.00 | +/-15 | 1,393 | OpenAI | Proprietary |
| 122 | ERNIE 5.0 Preview (1203)百度 | 1408.00 | +/-23 | 617 | 百度 | Proprietary |
| 123 | amazon-nova-experimental-chat-26-01-10Amazon | 1408.00 | +/-34 | 260 | Amazon | Proprietary |
| 124 | Step 3.5 FlashStepFunAI | 1407.00 | +/-11 | 3,216 | StepFunAI | Apache 2.0 |
| 125 | Gemini 2.5 FlashGoogle Deep Mind | 1405.00 | +/-7 | 7,831 | Google Deep Mind | Proprietary |
| 126 | OpenAI o3-mini (high)OpenAI | 1405.00 | +/-13 | 1,909 | OpenAI | Proprietary |
| 127 | GPT-5-mini (high)OpenAI | 1405.00 | +/-15 | 1,458 | OpenAI | Proprietary |
| 128 | Step 3.5 FlashStepFunAI | 1404.00 | +/-11 | 3,192 | StepFunAI | Proprietary |
| 129 | Qwen3-VL-235B-A22B-Instruct (thinking)阿里巴巴 | 1404.00 | +/-28 | 424 | 阿里巴巴 | Apache 2.0 |
| 130 | Claude Opus 4Anthropic | 1404.00 | +/-11 | 2,759 | Anthropic | Proprietary |
| 131 | GPT-4o(2025-03-27)OpenAI | 1404.00 | +/-8 | 5,708 | OpenAI | Proprietary |
| 132 | Claude Sonnet 4 (thinking-32k)Anthropic | 1403.00 | +/-13 | 2,019 | Anthropic | Proprietary |
| 133 | Mistral Large 3MistralAI | 1403.00 | +/-10 | 3,400 | MistralAI | Apache 2.0 |
| 134 | Hunyuan-T1腾讯AI实验室 | 1401.00 | +/-38 | 236 | 腾讯AI实验室 | Proprietary |
| 135 | Qwen3.5-35B-A3B阿里巴巴 | 1400.00 | +/-14 | 1,746 | 阿里巴巴 | Apache 2.0 |
| 136 | ERNIE 5.0百度 | 1399.00 | +/-34 | 267 | 百度 | Proprietary |
| 137 | Qwen3-32B阿里巴巴 | 1399.00 | +/-30 | 316 | 阿里巴巴 | Apache 2.0 |
| 138 | amazon-nova-experimental-chat-12-10Amazon | 1399.00 | +/-37 | 233 | Amazon | Proprietary |
| 139 | Haiku 4.5Anthropic | 1398.00 | +/-8 | 6,534 | Anthropic | Proprietary |
| 140 | Magistral-Medium-2506MistralAI | 1398.00 | +/-8 | 5,770 | MistralAI | Proprietary |
| 141 | amazon-nova-experimental-chat-11-10Amazon | 1397.00 | +/-15 | 1,567 | Amazon | Proprietary |
| 142 | qwen3-235b-a22b-thinking-2507Alibaba | 1397.00 | +/-25 | 486 | Alibaba | Apache 2.0 |
| 143 | DeepSeek-R1-0528DeepSeek-AI | 1396.00 | +/-20 | 863 | DeepSeek-AI | MIT |
| 144 | 1396.00 | +/-12 | 2,422 | MiniMaxAI | Modified MIT | |
| 145 | DeepSeek-V3.1 TerminusDeepSeek-AI | 1396.00 | +/-39 | 217 | DeepSeek-AI | MIT |
| 146 | amazon-nova-experimental-chat-10-20Amazon | 1394.00 | +/-20 | 806 | Amazon | Proprietary |
| 147 | qwen3-235b-a22b-no-thinkingAlibaba | 1393.00 | +/-12 | 2,381 | Alibaba | Apache 2.0 |
| 148 | Qwen3-235B-A22B阿里巴巴 | 1393.00 | +/-14 | 1,602 | 阿里巴巴 | Apache 2.0 |
| 149 | 1390.00 | +/-19 | 988 | MiniMaxAI | MIT | |
| 150 | GLM-4.5-Air智谱AI | 1389.00 | +/-15 | 1,535 | 智谱AI | MIT |
| 151 | Qwen3-Next (thinking)阿里巴巴 | 1389.00 | +/-20 | 826 | 阿里巴巴 | Apache 2.0 |
| 152 | Kimi K2Moonshot AI | 1389.00 | +/-14 | 1,689 | Moonshot AI | Modified MIT |
| 153 | Claude Sonnet 4Anthropic | 1388.00 | +/-12 | 2,462 | Anthropic | Proprietary |
| 154 | nvidia-llama-3.3-nemotron-super-49b-v1.5Nvidia | 1388.00 | +/-39 | 193 | Nvidia | Nvidia Open |
| 155 | OpenAI o3-mini (high)OpenAI | 1388.00 | +/-18 | 976 | OpenAI | Proprietary |
| 156 | OpenAI o1OpenAI | 1386.00 | +/-10 | 4,569 | OpenAI | Proprietary |
| 157 | Claude3-Sonnet (thinking-32k)Anthropic | 1385.00 | +/-11 | 2,785 | Anthropic | Proprietary |
| 158 | trinity-large-thinking Apache 2.0 | 1384.00 | +/-15 | 1,607 | — | — |
| 159 | OpenAI o3-miniOpenAI | 1382.00 | +/-9 | 4,715 | OpenAI | Proprietary |
| 160 | intellect-3 MIT | 1381.00 | +/-31 | 332 | — | — |
| 161 | GPT OSS 120BOpenAI | 1381.00 | +/-14 | 1,794 | OpenAI | Apache 2.0 |
| 162 | llama-3.1-nemotron-ultra-253b-v1Nvidia | 1380.00 | +/-37 | 209 | Nvidia | Nvidia Open Model |
| 163 | Qwen3-30B-A3B-2507阿里巴巴 | 1380.00 | +/-15 | 1,424 | 阿里巴巴 | Apache 2.0 |
| 164 | mimo-v2-flash (non-thinking)Xiaomi | 1378.00 | +/-11 | 2,809 | Xiaomi | MIT |
| 165 | Qwen3-Coder-480B-A35B阿里巴巴 | 1376.00 | +/-15 | 1,621 | 阿里巴巴 | Apache 2.0 |
| 166 | nvidia-nemotron-3-super-120b-a12bNvidia | 1376.00 | +/-25 | 517 | Nvidia | NVIDIA Open Model |
| 167 | mimo-v2-flash (thinking)Xiaomi | 1375.00 | +/-23 | 619 | Xiaomi | MIT |
| 168 | GPT-4.1OpenAI | 1373.00 | +/-10 | 3,221 | OpenAI | Proprietary |
| 169 | 1373.00 | +/-11 | 2,676 | xAI | Proprietary | |
| 170 | minimax-m1MiniMax | 1372.00 | +/-13 | 1,789 | MiniMax | Apache 2.0 |
| 171 | DeepSeek-V3-0324DeepSeek-AI | 1370.00 | +/-10 | 3,183 | DeepSeek-AI | MIT |
| 172 | 1369.00 | +/-14 | 1,520 | xAI | Proprietary | |
| 173 | Gemini 2.5 Flash-Lite (thinking)Google Deep Mind | 1365.00 | +/-12 | 2,084 | Google Deep Mind | Proprietary |
| 174 | GLM-4.7-Flash智谱AI | 1365.00 | +/-21 | 708 | 智谱AI | MIT |
| 175 | Gemini 2.5 Flash-Lite-Preview-09-2025 (no-thinking)Google Deep Mind | 1364.00 | +/-11 | 2,857 | Google Deep Mind | Proprietary |
| 176 | QwQ-32B阿里巴巴 | 1364.00 | +/-14 | 1,713 | 阿里巴巴 | Apache 2.0 |
| 177 | Step3StepFunAI | 1363.00 | +/-31 | 353 | StepFunAI | Apache 2.0 |
| 178 | Qwen2.5-Max阿里巴巴 | 1363.00 | +/-10 | 3,305 | 阿里巴巴 | Proprietary |
| 179 | Claude3-SonnetAnthropic | 1363.00 | +/-10 | 3,351 | Anthropic | Proprietary |
| 180 | OpenAI o1-miniOpenAI | 1362.00 | +/-8 | 7,499 | OpenAI | Proprietary |
| 181 | trinity-large-preview Apache 2.0 | 1361.00 | +/-14 | 1,865 | — | — |
| 182 | GLM-4.5V智谱AI | 1361.00 | +/-34 | 274 | 智谱AI | MIT |
| 183 | Gemini 2.0 Flash ExperimentalDeepMind | 1356.00 | +/-9 | 4,058 | DeepMind | Proprietary |
| 184 | GPT-4.1 miniOpenAI | 1354.00 | +/-11 | 2,689 | OpenAI | Proprietary |
| 185 | 1354.00 | +/-33 | 320 | MiniMaxAI | Apache 2.0 | |
| 186 | Qwen3-30B-A3B阿里巴巴 | 1353.00 | +/-14 | 1,701 | 阿里巴巴 | Apache 2.0 |
| 187 | ling-flash-2.0 AntGroup | 1353.00 | +/-27 | 461 | Group | MIT |
| 188 | Claude 3.5 SonnetAnthropic | 1351.00 | +/-7 | 10,014 | Anthropic | Proprietary |
| 189 | nvidia-nemotron-3-nano-30b-a3b-bf16Nvidia | 1351.00 | +/-19 | 975 | Nvidia | NVIDIA Open Model |
| 190 | mistral-medium-2505Mistral | 1348.00 | +/-12 | 2,222 | Mistral | Proprietary |
| 191 | hunyuan-turbos-20250416Tencent | 1347.00 | +/-20 | 845 | Tencent | Proprietary |
| 192 | GPT-5-Nano (high)OpenAI | 1345.00 | +/-27 | 491 | OpenAI | Proprietary |
| 193 | Claude 3.5 SonnetAnthropic | 1342.00 | +/-8 | 11,359 | Anthropic | Proprietary |
| 194 | ring-flash-2.0 AntGroup | 1339.00 | +/-27 | 450 | Group | MIT |
| 195 | Gemini 1.5 ProGoogle Deep Mind | 1339.00 | +/-7 | 7,610 | Google Deep Mind | Proprietary |
| 196 | Mistral-Small-3.2MistralAI | 1339.00 | +/-18 | 1,041 | MistralAI | Apache 2.0 |
| 197 | GPT OSS 20BOpenAI | 1336.00 | +/-22 | 678 | OpenAI | Apache 2.0 |
| 198 | Nova 2 Lite亚马逊 | 1333.00 | +/-20 | 825 | 亚马逊 | Proprietary |
| 199 | Gemini 2.0 Flash-LiteDeepMind | 1326.00 | +/-10 | 2,814 | DeepMind | Proprietary |
| 200 | qwen-plus-0125Alibaba | 1323.00 | +/-19 | 732 | Alibaba | Proprietary |
| 201 | Gemma 3 - 27B (IT)Google Deep Mind | 1322.00 | +/-10 | 3,576 | Google Deep Mind | Gemma |
| 202 | granite-4.1-8bIBM | 1319.00 | +/-39 | 234 | IBM | Apache 2.0 |
| 203 | llama-3.1-405b-instruct-fp8Meta | 1319.00 | +/-8 | 8,482 | Meta | Llama 3.1 Community |
| 204 | Gemma 3 - 12B (IT)Google Deep Mind | 1318.00 | +/-27 | 389 | Google Deep Mind | Gemma |
| 205 | Llama 4 Maverick InstructFacebook AI研究实验室 | 1317.00 | +/-11 | 2,834 | Facebook AI研究实验室 | Llama 4 |
| 206 | llama-3.1-405b-instruct-bf16Meta | 1315.00 | +/-8 | 5,215 | Meta | Llama 3.1 Community |
| 207 | step-2-16k-exp-202412StepFun | 1312.00 | +/-20 | 642 | StepFun | Proprietary |
| 208 | Claude3-OpusAnthropic | 1312.00 | +/-7 | 25,769 | Anthropic | Proprietary |
| 209 | athene-v2-chat NexusFlow | 1312.00 | +/-10 | 3,412 | — | — |
| 210 | olmo-3-32b-thinkAi2 | 1311.00 | +/-32 | 315 | Ai2 | Apache 2.0 |
| 211 | DeepSeek-V3DeepSeek-AI | 1311.00 | +/-11 | 2,721 | DeepSeek-AI | DeepSeek |
| 212 | C4AI Command A (202503)CohereAI | 1309.00 | +/-9 | 3,987 | CohereAI | CC-BY-NC-4.0 |
| 213 | Llama 4 Scout InstructFacebook AI研究实验室 | 1309.00 | +/-13 | 1,938 | Facebook AI研究实验室 | Llama |
| 214 | GPT-4oOpenAI | 1309.00 | +/-8 | 6,826 | OpenAI | Proprietary |
| 215 | gemini-advanced-0514Google | 1306.00 | +/-10 | 6,395 | Proprietary | |
| 216 | yi-lightning Proprietary | 1305.00 | +/-10 | 3,921 | — | — |
| 217 | GPT-4oOpenAI | 1305.00 | +/-7 | 15,103 | OpenAI | Proprietary |
| 218 | olmo-3.1-32b-instructAi2 | 1304.00 | +/-23 | 682 | Ai2 | Apache 2.0 |
| 219 | qwen2.5-plus-1127Alibaba | 1304.00 | +/-14 | 1,404 | Alibaba | Proprietary |
| 220 | GPT-4OpenAI | 1303.00 | +/-8 | 13,306 | OpenAI | Proprietary |
| 221 | hunyuan-turbos-20250226Tencent | 1301.00 | +/-31 | 238 | Tencent | Proprietary |
| 222 | GPT-4OpenAI | 1299.00 | +/-8 | 12,374 | OpenAI | Proprietary |
| 223 | Gemini 1.5 ProGoogle Deep Mind | 1299.00 | +/-8 | 10,492 | Google Deep Mind | Proprietary |
| 224 | glm-4-plus-0111Zhipu | 1298.00 | +/-19 | 721 | Zhipu | Proprietary |
| 225 | step-1o-turbo-202506StepFun | 1298.00 | +/-24 | 564 | StepFun | Proprietary |
| 226 | olmo-3.1-32b-thinkAi2 | 1297.00 | +/-26 | 472 | Ai2 | Apache 2.0 |
| 227 | Qwen2.5-VL-72B-Instruct阿里巴巴 | 1296.00 | +/-8 | 5,415 | 阿里巴巴 | Qwen |
| 228 | gpt-4-turbo-2024-04-09OpenAI | 1296.00 | +/-8 | 13,217 | OpenAI | Proprietary |
| 229 | Llama3.3-70B-InstructFacebook AI研究实验室 | 1296.00 | +/-8 | 5,772 | Facebook AI研究实验室 | Llama-3.3 |
| 230 | 1294.00 | +/-7 | 8,950 | xAI | Proprietary | |
| 231 | hunyuan-large-2025-02-10Tencent | 1294.00 | +/-24 | 497 | Tencent | Proprietary |
| 232 | deepseek-v2.5-1210DeepSeek | 1292.00 | +/-17 | 1,031 | DeepSeek | DeepSeek |
| 233 | qwen-max-0919Alibaba | 1291.00 | +/-12 | 2,249 | Alibaba | Qwen |
| 234 | hunyuan-standard-2025-02-10Tencent | 1290.00 | +/-24 | 499 | Tencent | Proprietary |
| 235 | gemini-1.5-flash-002Google | 1289.00 | +/-9 | 4,789 | Proprietary | |
| 236 | mistral-large-2407Mistral | 1288.00 | +/-8 | 6,664 | Mistral | Mistral Research |
| 237 | DeepSeek V2.5DeepSeek-AI | 1288.00 | +/-10 | 3,649 | DeepSeek-AI | DeepSeek |
| 238 | glm-4-plusZhipu AI | 1287.00 | +/-10 | 3,599 | Zhipu AI | Proprietary |
| 239 | Magistral-Medium-2506MistralAI | 1287.00 | +/-26 | 551 | MistralAI | Proprietary |
| 240 | Claude 3.5 HaikuAnthropic | 1286.00 | +/-8 | 6,353 | Anthropic | Proprietary |
| 241 | GPT-4OpenAI | 1284.00 | +/-10 | 7,052 | OpenAI | Proprietary |
| 242 | mistral-large-2411Mistral | 1282.00 | +/-9 | 3,574 | Mistral | MRL |
| 243 | hunyuan-large-visionTencent | 1281.00 | +/-30 | 350 | Tencent | Proprietary |
| 244 | hunyuan-turbo-0110Tencent | 1279.00 | +/-31 | 243 | Tencent | Proprietary |
| 245 | Llama3.1-70B-InstructFacebook AI研究实验室 | 1279.00 | +/-17 | 1,041 | Facebook AI研究实验室 | Llama 3.1 |
| 246 | ibm-granite-h-smallIBM | 1278.00 | +/-32 | 357 | IBM | Apache 2.0 |
| 247 | Mistral-Small-3.1-24B-Instruct-2503MistralAI | 1278.00 | +/-13 | 2,128 | MistralAI | Apache 2.0 |
| 248 | GPT-4OpenAI | 1276.00 | +/-8 | 11,181 | OpenAI | Proprietary |
| 249 | GPT-4o miniOpenAI | 1276.00 | +/-7 | 9,322 | OpenAI | Proprietary |
| 250 | GPT-4.1 nanoOpenAI | 1274.00 | +/-23 | 582 | OpenAI | Proprietary |
| 251 | Qwen2-72B-Instruct阿里巴巴 | 1273.00 | +/-9 | 4,835 | 阿里巴巴 | Qianwen LICENSE |
| 252 | 1273.00 | +/-8 | 7,261 | xAI | Proprietary | |
| 253 | deepseek-coder-v2DeepSeek | 1272.00 | +/-14 | 1,858 | DeepSeek | DeepSeek License |
| 254 | llama-3.1-nemotron-51b-instructNvidia | 1271.00 | +/-22 | 507 | Nvidia | Llama 3.1 |
| 255 | Qwen2.5-Coder-32B-Instruct阿里巴巴 | 1270.00 | +/-19 | 725 | 阿里巴巴 | Apache 2.0 |
| 256 | Llama3.1-70B-InstructFacebook AI研究实验室 | 1269.00 | +/-8 | 7,677 | Facebook AI研究实验室 | Llama 3.1 Community |
| 257 | amazon-nova-pro-v1.0Amazon | 1269.00 | +/-10 | 2,978 | Amazon | Proprietary |
| 258 | Phi-4-reasoningMicrosoft Azure | 1265.00 | +/-10 | 2,764 | Microsoft Azure | MIT |
| 259 | llama-3.1-tulu-3-70bAi2 | 1263.00 | +/-25 | 397 | Ai2 | Llama 3.1 |
| 260 | athene-70b-0725 CC-BY-NC-4.0 | 1262.00 | +/-10 | 2,921 | — | — |
| 261 | Mistral Small 24B Instruct 2501MistralAI | 1261.00 | +/-13 | 1,683 | MistralAI | Apache 2.0 |
| 262 | Gemma-3n-E4BGoogle Deep Mind | 1260.00 | +/-15 | 1,571 | Google Deep Mind | Gemma |
| 263 | gemini-1.5-flash-001Google | 1258.00 | +/-8 | 8,392 | Proprietary | |
| 264 | Llama3-70B-InstructFacebook AI研究实验室 | 1258.00 | +/-7 | 20,941 | Facebook AI研究实验室 | Llama 3 Community |
| 265 | Gemma 3 - 4B (IT)Google Deep Mind | 1254.00 | +/-28 | 423 | Google Deep Mind | Gemma |
| 266 | Claude3-SonnetAnthropic | 1253.00 | +/-8 | 13,766 | Anthropic | Proprietary |
| 267 | nemotron-4-340b-instructNvidia | 1252.00 | +/-12 | 2,352 | Nvidia | NVIDIA Open Model |
| 268 | hunyuan-standard-256kTencent | 1250.00 | +/-29 | 361 | Tencent | Proprietary |
| 269 | gemma-2-27b-itGoogle | 1247.00 | +/-7 | 10,170 | Gemma license | |
| 270 | GLM4智谱AI | 1246.00 | +/-16 | 1,191 | 智谱AI | Proprietary |
| 271 | reka-core-20240904 Proprietary | 1246.00 | +/-14 | 1,207 | — | — |
| 272 | jamba-1.5-large Jamba Open | 1245.00 | +/-15 | 1,147 | — | — |
| 273 | mistral-large-2402Mistral | 1244.00 | +/-9 | 7,987 | Mistral | Proprietary |
| 274 | amazon-nova-lite-v1.0Amazon | 1244.00 | +/-11 | 2,511 | Amazon | Proprietary |
| 275 | C4AI Aya Vision 32BCohereAI | 1233.00 | +/-10 | 3,854 | CohereAI | CC-BY-NC-4.0 |
| 276 | reka-flash-20240904 Proprietary | 1232.00 | +/-14 | 1,284 | — | — |
| 277 | Claude3-HaikuAnthropic | 1231.00 | +/-7 | 14,983 | Anthropic | Proprietary |
| 278 | command-r-plus-08-2024Cohere | 1231.00 | +/-14 | 1,467 | Cohere | CC-BY-NC-4.0 |
| 279 | gemini-1.5-flash-8b-001Google | 1229.00 | +/-9 | 5,036 | Proprietary | |
| 280 | Mixtral-8x22B-Instruct-v0.1MistralAI | 1228.00 | +/-9 | 6,778 | MistralAI | Apache 2.0 |
| 281 | olmo-2-0325-32b-instructAi2 | 1227.00 | +/-28 | 375 | Ai2 | Apache-2.0 |
| 282 | amazon-nova-micro-v1.0Amazon | 1224.00 | +/-11 | 2,455 | Amazon | Proprietary |
| 283 | Qwen1.5-110B-Chat阿里巴巴 | 1221.00 | +/-11 | 3,188 | 阿里巴巴 | Qianwen LICENSE |
| 284 | mistral-mediumMistral | 1220.00 | +/-11 | 4,406 | Mistral | Proprietary |
| 285 | gemma-2-9b-itGoogle | 1219.00 | +/-8 | 7,110 | Gemma license | |
| 286 | Phi-3-medium 14B-previewMicrosoft Azure | 1215.00 | +/-11 | 3,238 | Microsoft Azure | MIT |
| 287 | ministral-8b-2410Mistral | 1214.00 | +/-20 | 683 | Mistral | MRL |
| 288 | C4AI Command R+CohereAI | 1213.00 | +/-8 | 9,769 | CohereAI | CC-BY-NC-4.0 |
| 289 | Yi-1.5-34B零一万物 | 1213.00 | +/-11 | 2,985 | 零一万物 | — |
| 290 | reka-flash-21b-20240226-online Proprietary | 1212.00 | +/-14 | 2,028 | — | — |
| 291 | QwQ-32B-Preview阿里巴巴 | 1210.00 | +/-24 | 480 | 阿里巴巴 | Apache 2.0 |
| 292 | Qwen1.5-72B-Chat阿里巴巴 | 1209.00 | +/-10 | 5,327 | 阿里巴巴 | Qianwen LICENSE |
| 293 | gemma-2-9b-it-simpo MIT | 1207.00 | +/-15 | 1,285 | — | — |
| 294 | InternLM2-Base-20B上海人工智能实验室 | 1207.00 | +/-15 | 1,387 | 上海人工智能实验室 | — |
| 295 | command-r-08-2024Cohere | 1207.00 | +/-14 | 1,601 | Cohere | CC-BY-NC-4.0 |
| 296 | llama-3.1-tulu-3-8bAi2 | 1206.00 | +/-26 | 363 | Ai2 | Llama 3.1 |
| 297 | gpt-3.5-turbo-1106OpenAI | 1204.00 | +/-15 | 2,134 | OpenAI | Proprietary |
| 298 | Gemini-proDeepMind | 1201.00 | +/-19 | 993 | DeepMind | Proprietary |
| 299 | C4AI Aya Vision 8BCohereAI | 1201.00 | +/-15 | 1,307 | CohereAI | CC-BY-NC-4.0 |
| 300 | qwen1.5-32b-chatAlibaba | 1201.00 | +/-12 | 2,649 | Alibaba | Qianwen LICENSE |
| 301 | gpt-3.5-turbo-0125OpenAI | 1201.00 | +/-9 | 8,626 | OpenAI | Proprietary |
| 302 | reka-flash-21b-20240226 Proprietary | 1199.00 | +/-11 | 3,363 | — | — |
| 303 | granite-3.1-2b-instructIBM | 1198.00 | +/-26 | 391 | IBM | Apache 2.0 |
| 304 | granite-3.0-8b-instructIBM | 1197.00 | +/-19 | 873 | IBM | Apache 2.0 |
| 305 | gemini-pro-dev-apiGoogle | 1197.00 | +/-14 | 2,274 | Proprietary | |
| 306 | zephyr-orpo-141b-A35b-v0.1 Apache 2.0 | 1196.00 | +/-22 | 589 | — | — |
| 307 | dbrx-instruct-preview DBRX LICENSE | 1196.00 | +/-11 | 4,001 | — | — |
| 308 | Phi-3-mini 3.8BMicrosoft Azure | 1193.00 | +/-14 | 1,568 | Microsoft Azure | MIT |
| 309 | Phi-3-small 7BMicrosoft Azure | 1193.00 | +/-13 | 2,092 | Microsoft Azure | MIT |
| 310 | Llama3-8B-InstructFacebook AI研究实验室 | 1193.00 | +/-8 | 14,252 | Facebook AI研究实验室 | Llama 3 Community |
| 311 | mixtral-8x7b-instruct-v0.1Mistral | 1191.00 | +/-9 | 9,663 | Mistral | Apache 2.0 |
| 312 | Llama3.1-8B-InstructFacebook AI研究实验室 | 1190.00 | +/-28 | 382 | Facebook AI研究实验室 | Apache 2.0 |
| 313 | Llama3.1-8B-InstructFacebook AI研究实验室 | 1189.00 | +/-8 | 7,135 | Facebook AI研究实验室 | Llama 3.1 Community |
| 314 | jamba-1.5-mini Jamba Open | 1186.00 | +/-16 | 1,094 | — | — |
| 315 | command-rCohere | 1176.00 | +/-10 | 6,682 | Cohere | CC-BY-NC-4.0 |
| 316 | Qwen3-VL-2B阿里巴巴 | 1169.00 | +/-19 | 908 | 阿里巴巴 | Apache 2.0 |
| 317 | Qwen1.5-14B-Chat阿里巴巴 | 1167.00 | +/-14 | 2,184 | 阿里巴巴 | Qianwen LICENSE |
| 318 | llama-3.2-3b-instructMeta | 1165.00 | +/-16 | 1,136 | Meta | Llama 3.2 |
| 319 | gemma-2-2b-itGoogle | 1164.00 | +/-8 | 6,599 | Gemma license | |
| 320 | snowflake-arctic-instruct Apache 2.0 | 1163.00 | +/-11 | 4,793 | — | — |
| 321 | Gemma 1.1-7B-ITGoogle Research | 1161.00 | +/-11 | 3,039 | Google Research | Gemma license |
| 322 | openchat-3.5-0106 Apache-2.0 | 1158.00 | +/-14 | 1,726 | — | — |
| 323 | starling-lm-7b-beta Apache-2.0 | 1157.00 | +/-14 | 1,973 | — | — |
| 324 | WizardLM-70B-V1.0WizardLM Team | 1157.00 | +/-19 | 903 | WizardLM Team | Llama 2 Community |
| 325 | DeepSeek LLM 67B ChatDeepSeek-AI | 1155.00 | +/-24 | 576 | DeepSeek-AI | DeepSeek License |
| 326 | smollm2-1.7b-instruct Apache 2.0 | 1152.00 | +/-33 | 271 | — | — |
| 327 | openhermes-2.5-mistral-7b Apache-2.0 | 1152.00 | +/-20 | 697 | — | — |
| 328 | Yi-34B零一万物 | 1150.00 | +/-13 | 2,043 | 零一万物 | — |
| 329 | Phi-3-mini 3.8BMicrosoft Azure | 1150.00 | +/-12 | 2,564 | Microsoft Azure | MIT |
| 330 | LLaMA2 70BFacebook AI研究实验室 | 1144.00 | +/-19 | 888 | Facebook AI研究实验室 | — |
| 331 | Phi-3-mini 3.8BMicrosoft Azure | 1139.00 | +/-13 | 2,813 | Microsoft Azure | MIT |
| 332 | llama-2-70b-chatMeta | 1136.00 | +/-10 | 4,740 | Meta | Llama 2 Community |
| 333 | Mistral-7B-Instruct-v0.2MistralAI | 1127.00 | +/-12 | 2,605 | MistralAI | Apache-2.0 |
| 334 | starling-lm-7b-alpha CC-BY-NC-4.0 | 1126.00 | +/-16 | 1,300 | — | — |
| 335 | Qwen-14B-Chat阿里巴巴 | 1126.00 | +/-24 | 534 | 阿里巴巴 | Qianwen LICENSE |
| 336 | openchat-3.5 Apache-2.0 | 1125.00 | +/-18 | 945 | — | — |
| 337 | dolphin-2.2.1-mistral-7b Apache-2.0 | 1125.00 | +/-32 | 219 | — | — |
| 338 | llama-3.2-1b-instructMeta | 1123.00 | +/-16 | 1,162 | Meta | Llama 3.2 |
| 339 | Qwen1.5-7B-Chat阿里巴巴 | 1121.00 | +/-21 | 690 | 阿里巴巴 | Qianwen LICENSE |
| 340 | Gemma 7B - ItGoogle Research | 1119.00 | +/-16 | 1,120 | Google Research | Gemma license |
| 341 | PaLM 2Google Research | 1116.00 | +/-19 | 901 | Google Research | Proprietary |
| 342 | Vicuna 33BLM-SYS | 1116.00 | +/-13 | 2,663 | LM-SYS | — |
| 343 | llama2-70b-steerlm-chatNvidia | 1114.00 | +/-27 | 440 | Nvidia | Llama 2 Community |
| 344 | Baichuan2-13B-Chat百川智能 | 1110.00 | +/-13 | 2,218 | 百川智能 | Llama 2 Community |
| 345 | Gemma 1.1-2B-ITGoogle Research | 1109.00 | +/-16 | 1,355 | Google Research | Gemma license |
| 346 | CodeLLaMA-34BFacebook AI研究实验室 | 1109.00 | +/-19 | 770 | Facebook AI研究实验室 | Llama 2 Community |
| 347 | solar-10.7b-instruct-v1.0 CC-BY-NC-4.0 | 1109.00 | +/-22 | 604 | — | — |
| 348 | mpt-30b-chat CC-BY-NC-SA-4.0 | 1095.00 | +/-34 | 242 | — | — |
| 349 | nous-hermes-2-mixtral-8x7b-dpo Apache-2.0 | 1093.00 | +/-21 | 628 | — | — |
| 350 | Qwen1.5-4B-Chat阿里巴巴 | 1087.00 | +/-18 | 988 | 阿里巴巴 | Qianwen LICENSE |
| 351 | Baichuan2-7B-Chat百川智能 | 1086.00 | +/-14 | 1,656 | 百川智能 | Llama 2 Community |
| 352 | stripedhyena-nous-7b Apache 2.0 | 1085.00 | +/-20 | 676 | — | — |
| 353 | vicuna-13b Llama 2 Community | 1083.00 | +/-14 | 2,146 | — | — |
| 354 | Mistral 7B InstructMistralAI | 1082.00 | +/-19 | 974 | MistralAI | Apache 2.0 |
| 355 | zephyr-7b-beta MIT | 1082.00 | +/-17 | 1,250 | — | — |
| 356 | guanaco-33b Non-commercial | 1080.00 | +/-32 | 280 | — | — |
| 357 | Gemma 2B - ItGoogle Research | 1072.00 | +/-22 | 597 | Google Research | Gemma license |
| 358 | wizardlm-13bMicrosoft | 1064.00 | +/-21 | 669 | Microsoft | Llama 2 Community |
| 359 | olmo-7b-instructAi2 | 1054.00 | +/-19 | 848 | Ai2 | Apache-2.0 |
| 360 | vicuna-7b Llama 2 Community | 1048.00 | +/-22 | 658 | — | — |
| 361 | ChatGLM3-6B智谱AI | 1042.00 | +/-23 | 576 | 智谱AI | — |
| 362 | GPT4All 13BNomic AI | 999.00 | +/-37 | 211 | Nomic AI | — |
| 363 | alpaca-13b Non-commercial | 994.00 | +/-23 | 652 | — | — |
| 364 | mpt-7b-chat CC-BY-NC-SA-4.0 | 986.00 | +/-25 | 471 | — | — |
| 365 | RWKV-4-Raven-14B Apache 2.0 | 984.00 | +/-24 | 544 | — | — |
| 366 | Koala达摩院 | 980.00 | +/-21 | 751 | 达摩院 | — |
| 367 | ChatGLM-6B智谱AI | 977.00 | +/-26 | 525 | 智谱AI | — |
| 368 | ChatGLM2-6B智谱AI | 972.00 | +/-35 | 227 | 智谱AI | — |
| 369 | oasst-pythia-12b Apache 2.0 | 961.00 | +/-22 | 687 | — | — |
| 370 | dolly-v2-12b MIT | 952.00 | +/-29 | 370 | — | — |
| 371 | LLaMA 13BFacebook AI研究实验室 | 921.00 | +/-33 | 252 | Facebook AI研究实验室 | Non-commercial |
| 372 | fastchat-t5-3b Apache 2.0 | 920.00 | +/-26 | 462 | — | — |
| 373 | stablelm-tuned-alpha-7b CC-BY-NC-SA-4.0 | 890.00 | +/-29 | 353 | — | — |
Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.
FAQ
What is LMArena Math Arena?
LMArena Math Arena is an anonymous evaluation track focused on mathematical reasoning. Users submit real math questions, compare hidden model solutions side by side, and vote for the better answer; the leaderboard is then calculated with Elo-style scoring.
How is Math Arena different from MATH-500 or AIME?
Static benchmarks such as MATH-500 and AIME use fixed problem sets and automated grading. Math Arena uses open-ended user questions and human preference voting, making it a useful complement for measuring how models handle varied real-world math tasks.
Do thinking models perform better in Math Arena?
Models with extended reasoning or chain-of-thought style capabilities often rank higher on math tasks because they spend more time decomposing and checking solutions. That benefit can come with higher latency and cost.
How do China-developed models perform in math?
DeepSeek, Qwen, GLM, and related models have become competitive in math reasoning leaderboards. Open licenses and Chinese-language support can make them especially useful for local deployment and education scenarios.
















