LMArena Math Arena 数学推理能力排行榜
基于 LMArena Math Arena 用户匿名投票的最新AI大模型数学推理能力排行榜,涵盖各模型的 Elo 得分、95% 置信区间、投票量、机构与许可证。
榜首模型
Claude Fable 5
最高得分
1529.00
模型数量
383
数据版本
2026年09月02日
数据来源: LM Arena
关于本排行榜
本排行榜展示了当前 AI 大模型在数学推理任务中的实力排名。数据来源于 LMArena 的 Math 子赛道,通过真实用户匿名盲测投票评估各模型在数学解题任务中的表现。
评测方法概要
匿名盲测:用户提出数学题目后,由两个"隐藏身份"的模型分别作答,用户投票选出解题更优的一方,排除品牌偏见。
Elo 评分:采用 Bradley-Terry 模型计算 Elo 分数,分数越高说明该模型在数学场景中被用户更频繁地选择。
覆盖多种数学场景:包括代数、几何、计算推理、竞赛数学等多元化的真实数学任务。
DataLearner 在原始数据基础上提供中文解读与深度分析,并将排行榜模型关联至 DataLearner 模型库,方便您一键查看模型详情、API 定价、评测得分等完整信息。
排名总表
| 排名 | 模型名称 | 得分 | 95% CI | 投票数 | 机构 | 许可证 |
|---|---|---|---|---|---|---|
Claude Fable 5Anthropic | 1529.00 | +/-16 | 1,331 | Anthropic | Proprietary | |
claude-opus-5-maxAnthropic | 1529.00 | +/-22 | 740 | Anthropic | Proprietary | |
claude-opus-5-highAnthropic | 1526.00 | +/-16 | 1,523 | Anthropic | Proprietary | |
| 4 | gemini-3.7-flash-highGoogle | 1524.00 | +/-33 | 313 | Proprietary | |
| 5 | Claude Opus 4.6 (high)Anthropic | 1517.00 | +/-10 | 3,706 | Anthropic | Proprietary |
| 6 | glm-5.3-maxZ.ai | 1514.00 | +/-32 | 300 | Z.ai | MIT |
| 7 | Claude Opus 4.6Anthropic | 1506.00 | +/-10 | 4,122 | Anthropic | Proprietary |
| 8 | kimi-k3-maxMoonshot | 1506.00 | +/-21 | 760 | Moonshot | Kimi K3 license |
| 9 | gemini-3.6-flash-highGoogle | 1504.00 | +/-18 | 1,040 | Proprietary | |
| 10 | gemini-3.5-flash-highGoogle | 1504.00 | +/-15 | 1,652 | Proprietary | |
| 11 | Opus 4.7 (high)Anthropic | 1503.00 | +/-11 | 3,031 | Anthropic | Proprietary |
| 12 | GPT-5.5OpenAI | 1502.00 | +/-11 | 3,239 | OpenAI | Proprietary |
| 13 | gemini-3.8-flash-highGoogle | 1501.00 | +/-42 | 191 | Proprietary | |
| 14 | qwen3.8-maxAlibaba | 1497.00 | +/-24 | 595 | Alibaba | Proprietary |
| 15 | Claude Opus 4.8 (high)Anthropic | 1496.00 | +/-13 | 2,238 | Anthropic | Proprietary |
| 16 | GPT-5.4 (high)OpenAI | 1494.00 | +/-11 | 3,183 | OpenAI | Proprietary |
| 17 | Opus 4.7Anthropic | 1492.00 | +/-11 | 3,153 | Anthropic | Proprietary |
| 18 | GPT-5.5 (high)OpenAI | 1491.00 | +/-11 | 3,170 | OpenAI | Proprietary |
| 19 | Muse Spark 1.1Facebook AI研究实验室 | 1490.00 | +/-18 | 1,025 | Facebook AI研究实验室 | Proprietary |
| 20 | Qwen3.7-Max-Preview阿里巴巴 | 1490.00 | +/-40 | 218 | 阿里巴巴 | Proprietary |
| 21 | Gemini 3.1 Pro PreviewGoogle Deep Mind | 1489.00 | +/-9 | 5,406 | Google Deep Mind | Proprietary |
| 22 | GPT-5.6 Sol (xhigh)OpenAI | 1485.00 | +/-18 | 1,033 | OpenAI | Proprietary |
| 23 | GPT-5.6 Terra (xhigh)OpenAI | 1482.00 | +/-18 | 1,021 | OpenAI | Proprietary |
| 24 | 1479.00 | +/-18 | 1,087 | xAI | Proprietary | |
| 25 | Kimi K2.6Moonshot AI | 1478.00 | +/-13 | 1,936 | Moonshot AI | Modified MIT |
| 26 | GPT-5.6 Luna (xhigh)OpenAI | 1478.00 | +/-18 | 1,094 | OpenAI | Proprietary |
| 27 | gemini-3.5-flash-mediumGoogle | 1478.00 | +/-16 | 1,474 | Proprietary | |
| 28 | GLM 5.1智谱AI | 1477.00 | +/-13 | 2,244 | 智谱AI | MIT |
| 29 | InklingThinking Machines Lab | 1477.00 | +/-19 | 954 | Thinking Machines Lab | Apache 2.0 |
| 30 | Gemini 3 ProGoogle Deep Mind | 1477.00 | +/-11 | 2,622 | Google Deep Mind | Proprietary |
| 31 | mimo-v2.5-proXiaomi | 1477.00 | +/-12 | 2,720 | Xiaomi | MIT |
| 32 | ERNIE-5.1-Preview百度 | 1477.00 | +/-14 | 1,891 | 百度 | Proprietary |
| 33 | glm-5.2-maxZ.ai | 1476.00 | +/-15 | 1,575 | Z.ai | MIT |
| 34 | Gemini 3.0 FlashGoogle Deep Mind | 1475.00 | +/-13 | 1,963 | Google Deep Mind | Proprietary |
| 35 | Hy3腾讯AI实验室 | 1475.00 | +/-31 | 359 | 腾讯AI实验室 | Apache 2.0 |
| 36 | Qwen3.6-Max-Preview阿里巴巴 | 1475.00 | +/-30 | 360 | 阿里巴巴 | Proprietary |
| 37 | glm-5.3-flashZ.ai | 1475.00 | +/-43 | 182 | Z.ai | MIT |
| 38 | Claude Opus 4.8Anthropic | 1474.00 | +/-13 | 2,332 | Anthropic | Proprietary |
| 39 | Claude Sonnet 4.5 (high)Anthropic | 1474.00 | +/-16 | 1,363 | Anthropic | Proprietary |
| 40 | Gemma 4 31BDeepMind | 1472.00 | +/-27 | 402 | DeepMind | Apache 2.0 |
| 41 | Kimi K2 ThinkingMoonshot AI | 1471.00 | +/-10 | 3,925 | Moonshot AI | Modified MIT |
| 42 | Opus 4.5Anthropic | 1470.00 | +/-12 | 2,223 | Anthropic | Proprietary |
| 43 | Qwen3.5 Max Preview阿里巴巴 | 1470.00 | +/-16 | 1,346 | 阿里巴巴 | Proprietary |
| 44 | deepseek-v4-pro-high-previewDeepSeek | 1469.00 | +/-12 | 2,505 | DeepSeek | MIT |
| 45 | Gemma 4 26B A4BDeepMind | 1468.00 | +/-28 | 374 | DeepMind | Apache 2.0 |
| 46 | 1467.00 | +/-11 | 3,319 | xAI | Proprietary | |
| 47 | Qwen3.7-Plus阿里巴巴 | 1467.00 | +/-15 | 1,691 | 阿里巴巴 | Proprietary |
| 48 | Claude Opus 4Anthropic | 1465.00 | +/-9 | 4,282 | Anthropic | Proprietary |
| 49 | GPT-5.5 InstantOpenAI | 1463.00 | +/-16 | 1,455 | OpenAI | Proprietary |
| 50 | Claude Sonnet 4.6Anthropic | 1463.00 | +/-10 | 3,558 | Anthropic | Proprietary |
| 51 | deepseek-v4-pro-high-20260813DeepSeek | 1463.00 | +/-40 | 192 | DeepSeek | MIT |
| 52 | Muse SparkFacebook AI研究实验室 | 1461.00 | +/-20 | 857 | Facebook AI研究实验室 | Proprietary |
| 53 | GPT-5.4OpenAI | 1461.00 | +/-11 | 3,352 | OpenAI | Proprietary |
| 54 | qwen3.8-27bAlibaba | 1460.00 | +/-34 | 289 | Alibaba | Apache 2.0 |
| 55 | GPT-5.2 Pro (high)OpenAI | 1458.00 | +/-11 | 2,961 | OpenAI | Proprietary |
| 56 | GPT-5.1 Pro (high)OpenAI | 1456.00 | +/-12 | 2,460 | OpenAI | Proprietary |
| 57 | Claude Sonnet 4.5Anthropic | 1455.00 | +/-9 | 4,841 | Anthropic | Proprietary |
| 58 | Qwen 3.6 Plus Preview阿里巴巴 | 1455.00 | +/-12 | 2,531 | 阿里巴巴 | Proprietary |
| 59 | Gemini 3.0 Flash (minimal)Google Deep Mind | 1455.00 | +/-9 | 4,679 | Google Deep Mind | Proprietary |
| 60 | Inkling SmallThinky | 1453.00 | +/-24 | 589 | Thinky | Apache 2.0 |
| 61 | GPT-5.2 Chat (0210)OpenAI | 1453.00 | +/-13 | 2,072 | OpenAI | Proprietary |
| 62 | 1453.00 | +/-11 | 3,230 | xAI | Proprietary | |
| 63 | Muse Glimmer-30BFacebook AI研究实验室 | 1453.00 | +/-41 | 207 | Facebook AI研究实验室 | Apache-2.0 |
| 64 | mimo-v2-proXiaomi | 1452.00 | +/-15 | 1,608 | Xiaomi | Proprietary |
| 65 | 1452.00 | +/-15 | 1,595 | xAI | Proprietary | |
| 66 | DOLA Seed 2.0 Pro字节跳动Seed团队 | 1451.00 | +/-10 | 4,083 | 字节跳动Seed团队 | Proprietary |
| 67 | Qwen3.5-397B-A17B阿里巴巴 | 1448.00 | +/-10 | 4,095 | 阿里巴巴 | Apache 2.0 |
| 68 | OpenAI o3OpenAI | 1448.00 | +/-10 | 3,679 | OpenAI | Proprietary |
| 69 | DeepSeek-V4-ProDeepSeek-AI | 1445.00 | +/-12 | 2,872 | DeepSeek-AI | MIT |
| 70 | Nemotron 3 UltraNVIDIA | 1445.00 | +/-25 | 548 | NVIDIA | OpenMDW-1.1 |
| 71 | 1444.00 | +/-10 | 3,765 | xAI | Proprietary | |
| 72 | Opus 4.1 (thinking-16k)Anthropic | 1444.00 | +/-11 | 2,996 | Anthropic | Proprietary |
| 73 | GLM-5智谱AI | 1443.00 | +/-14 | 1,607 | 智谱AI | MIT |
| 74 | GLM-5V-Turbo智谱AI | 1443.00 | +/-26 | 443 | 智谱AI | Proprietary |
| 75 | Gemini 2.5 ProGoogle Deep Mind | 1442.00 | +/-7 | 7,565 | Google Deep Mind | Proprietary |
| 76 | deepseek-v4-flash-high-previewDeepSeek | 1441.00 | +/-12 | 2,459 | DeepSeek | MIT |
| 77 | gemini-3.5-flash-liteGoogle | 1441.00 | +/-19 | 956 | Proprietary | |
| 78 | GPT-5.4 mini (high)OpenAI | 1441.00 | +/-11 | 3,165 | OpenAI | Proprietary |
| 79 | mimo-v2.5Xiaomi | 1441.00 | +/-12 | 2,327 | Xiaomi | MIT |
| 80 | Kimi K2.5 InstantMoonshot AI | 1439.00 | +/-25 | 509 | Moonshot AI | Modified MIT |
| 81 | MiniMax M3MiniMaxAI | 1438.00 | +/-13 | 2,113 | MiniMaxAI | MiniMax Community License |
| 82 | Gemini 3.1 Flash-LiteGoogle Deep Mind | 1438.00 | +/-10 | 3,408 | Google Deep Mind | Proprietary |
| 83 | Kimi K2 Thinking (thinking-turbo)Moonshot AI | 1437.00 | +/-10 | 3,720 | Moonshot AI | Modified MIT |
| 84 | ERNIE 5.0 (0110)百度 | 1437.00 | +/-13 | 2,123 | 百度 | Proprietary |
| 85 | Qwen3 Max (Preview)阿里巴巴 | 1436.00 | +/-15 | 1,484 | 阿里巴巴 | Proprietary |
| 86 | longcat-flash-chat-2602-expMeituan | 1436.00 | +/-14 | 1,748 | Meituan | Proprietary |
| 87 | mimo-v2-omniXiaomi | 1436.00 | +/-20 | 914 | Xiaomi | Proprietary |
| 88 | GPT-5-Pro (high)OpenAI | 1433.00 | +/-14 | 1,858 | OpenAI | Proprietary |
| 89 | GPT-5.2OpenAI | 1433.00 | +/-9 | 4,366 | OpenAI | Proprietary |
| 90 | Opus 4.1Anthropic | 1432.00 | +/-9 | 4,663 | Anthropic | Proprietary |
| 91 | Mistral Medium 3.5MistralAI | 1430.00 | +/-25 | 539 | MistralAI | Modified MIT |
| 92 | DeepSeek V3.2DeepSeek-AI | 1429.00 | +/-11 | 2,973 | DeepSeek-AI | MIT |
| 93 | Qwen3.5-27B阿里巴巴 | 1429.00 | +/-15 | 1,642 | 阿里巴巴 | Apache 2.0 |
| 94 | GLM-4.7智谱AI | 1428.00 | +/-21 | 677 | 智谱AI | MIT |
| 95 | GPT-5.3 ChatOpenAI | 1427.00 | +/-13 | 2,031 | OpenAI | Proprietary |
| 96 | 1427.00 | +/-9 | 4,176 | xAI | Proprietary | |
| 97 | Claude Sonnet 4.5Anthropic | 1427.00 | +/-9 | 4,881 | Anthropic | Proprietary |
| 98 | 1426.00 | +/-12 | 2,220 | xAI | Proprietary | |
| 99 | 1426.00 | +/-40 | 216 | SpaceXAI | Proprietary | |
| 100 | hunyuan-hy3-previewTencent | 1426.00 | +/-28 | 404 | Tencent | tencent-hunyuan-community |
| 101 | DeepSeek-V4-FlashDeepSeek-AI | 1426.00 | +/-12 | 2,515 | DeepSeek-AI | MIT |
| 102 | qwen3-max-2025-09-23Alibaba | 1425.00 | +/-24 | 565 | Alibaba | Proprietary |
| 103 | DeepSeek V3.2-Exp (thinking)DeepSeek-AI | 1425.00 | +/-12 | 2,474 | DeepSeek-AI | MIT |
| 104 | 1425.00 | +/-30 | 382 | xAI | Proprietary | |
| 105 | DeepSeek V3.2-Exp (thinking)DeepSeek-AI | 1424.00 | +/-27 | 470 | DeepSeek-AI | MIT |
| 106 | amazon-nova-experimental-chat-26-02-10Amazon | 1424.00 | +/-39 | 208 | Amazon | Proprietary |
| 107 | Qwen3.5-122B-A10B阿里巴巴 | 1424.00 | +/-14 | 1,762 | 阿里巴巴 | Apache 2.0 |
| 108 | GPT-5.4 nano (high)OpenAI | 1424.00 | +/-11 | 3,081 | OpenAI | Proprietary |
| 109 | GPT-5.1OpenAI | 1422.00 | +/-11 | 2,823 | OpenAI | Proprietary |
| 110 | Claude Opus 4 (thinking-16k)Anthropic | 1422.00 | +/-12 | 2,203 | Anthropic | Proprietary |
| 111 | MiniMax-M2.7MiniMaxAI | 1422.00 | +/-11 | 3,305 | MiniMaxAI | Modified MIT |
| 112 | 1420.00 | +/-11 | 3,103 | xAI | Proprietary | |
| 113 | Qwen3-235B-A22B-2507阿里巴巴 | 1418.00 | +/-8 | 5,851 | 阿里巴巴 | Apache 2.0 |
| 114 | GLM-4.6智谱AI | 1417.00 | +/-13 | 2,058 | 智谱AI | MIT |
| 115 | Qwen3-Next阿里巴巴 | 1417.00 | +/-17 | 1,191 | 阿里巴巴 | Apache 2.0 |
| 116 | OpenAI o4 - miniOpenAI | 1417.00 | +/-11 | 2,899 | OpenAI | Proprietary |
| 117 | 1417.00 | +/-10 | 3,445 | xAI | Proprietary | |
| 118 | DeepSeek-V3.1 (thinking)DeepSeek-AI | 1416.00 | +/-22 | 658 | DeepSeek-AI | MIT |
| 119 | longcat-flash-chatMeituan | 1415.00 | +/-22 | 686 | Meituan | MIT |
| 120 | Kimi K2 0905Moonshot AI | 1415.00 | +/-21 | 753 | Moonshot AI | Modified MIT |
| 121 | DeepSeek-V3.1DeepSeek-AI | 1414.00 | +/-18 | 981 | DeepSeek-AI | MIT |
| 122 | DeepSeek V3.2-ExpDeepSeek-AI | 1414.00 | +/-21 | 771 | DeepSeek-AI | MIT |
| 123 | GLM-4.5智谱AI | 1413.00 | +/-16 | 1,408 | 智谱AI | MIT |
| 124 | DeepSeek-R1DeepSeek-AI | 1412.00 | +/-14 | 1,606 | DeepSeek-AI | MIT |
| 125 | GPT-5.2 Chat (0210)OpenAI | 1412.00 | +/-14 | 1,763 | OpenAI | Proprietary |
| 126 | 1411.00 | +/-18 | 1,053 | xAI | Proprietary | |
| 127 | Gemini 2.5 Flash-Preview-09-2025Google Deep Mind | 1410.00 | +/-13 | 1,926 | Google Deep Mind | Proprietary |
| 128 | OpenAI o1OpenAI | 1409.00 | +/-11 | 2,986 | OpenAI | Proprietary |
| 129 | ERNIE 5.0 Preview (1203)百度 | 1409.00 | +/-23 | 604 | 百度 | Proprietary |
| 130 | GPT-4.5OpenAI | 1408.00 | +/-15 | 1,393 | OpenAI | Proprietary |
| 131 | amazon-nova-experimental-chat-26-01-10Amazon | 1408.00 | +/-34 | 260 | Amazon | Proprietary |
| 132 | Qwen3-VL-235B-A22B-Instruct阿里巴巴 | 1408.00 | +/-23 | 694 | 阿里巴巴 | Apache 2.0 |
| 133 | Step 3.5 FlashStepFunAI | 1406.00 | +/-11 | 3,235 | StepFunAI | Apache 2.0 |
| 134 | Gemini 2.5 FlashGoogle Deep Mind | 1406.00 | +/-7 | 7,781 | Google Deep Mind | Proprietary |
| 135 | OpenAI o3-mini (high)OpenAI | 1405.00 | +/-13 | 1,909 | OpenAI | Proprietary |
| 136 | GPT-5-mini (high)OpenAI | 1405.00 | +/-16 | 1,439 | OpenAI | Proprietary |
| 137 | Hunyuan-T1腾讯AI实验室 | 1404.00 | +/-38 | 233 | 腾讯AI实验室 | Proprietary |
| 138 | Claude Opus 4Anthropic | 1404.00 | +/-11 | 2,722 | Anthropic | Proprietary |
| 139 | GPT-4o(2025-03-27)OpenAI | 1404.00 | +/-8 | 5,654 | OpenAI | Proprietary |
| 140 | Claude Sonnet 4 (thinking-32k)Anthropic | 1403.00 | +/-13 | 1,993 | Anthropic | Proprietary |
| 141 | Step 3.5 FlashStepFunAI | 1403.00 | +/-11 | 3,195 | StepFunAI | Proprietary |
| 142 | Mistral Large 3MistralAI | 1403.00 | +/-10 | 3,760 | MistralAI | Apache 2.0 |
| 143 | Qwen3-VL-235B-A22B-Instruct (thinking)阿里巴巴 | 1401.00 | +/-29 | 416 | 阿里巴巴 | Apache 2.0 |
| 144 | amazon-nova-experimental-chat-12-10Amazon | 1400.00 | +/-37 | 233 | Amazon | Proprietary |
| 145 | Qwen3.5-35B-A3B阿里巴巴 | 1400.00 | +/-14 | 1,754 | 阿里巴巴 | Apache 2.0 |
| 146 | Qwen3-32B阿里巴巴 | 1399.00 | +/-30 | 316 | 阿里巴巴 | Apache 2.0 |
| 147 | Haiku 4.5Anthropic | 1399.00 | +/-8 | 6,985 | Anthropic | Proprietary |
| 148 | qwen3-235b-a22b-thinking-2507Alibaba | 1398.00 | +/-25 | 479 | Alibaba | Apache 2.0 |
| 149 | Magistral-Medium-2506MistralAI | 1398.00 | +/-8 | 5,762 | MistralAI | Proprietary |
| 150 | amazon-nova-experimental-chat-11-10Amazon | 1397.00 | +/-15 | 1,556 | Amazon | Proprietary |
| 151 | ERNIE 5.0百度 | 1396.00 | +/-34 | 265 | 百度 | Proprietary |
| 152 | DeepSeek-V3.1 TerminusDeepSeek-AI | 1396.00 | +/-39 | 217 | DeepSeek-AI | MIT |
| 153 | MiniMax M2.5MiniMaxAI | 1396.00 | +/-12 | 2,426 | MiniMaxAI | Modified MIT |
| 154 | amazon-nova-experimental-chat-10-20Amazon | 1393.00 | +/-20 | 803 | Amazon | Proprietary |
| 155 | DeepSeek-R1-0528DeepSeek-AI | 1393.00 | +/-20 | 854 | DeepSeek-AI | MIT |
| 156 | Qwen3-Next (thinking)阿里巴巴 | 1393.00 | +/-20 | 816 | 阿里巴巴 | Apache 2.0 |
| 157 | qwen3-235b-a22b-no-thinkingAlibaba | 1392.00 | +/-12 | 2,363 | Alibaba | Apache 2.0 |
| 158 | Qwen3-235B-A22B阿里巴巴 | 1392.00 | +/-14 | 1,597 | 阿里巴巴 | Apache 2.0 |
| 159 | M2.1MiniMaxAI | 1390.00 | +/-19 | 973 | MiniMaxAI | MIT |
| 160 | Claude Sonnet 4Anthropic | 1390.00 | +/-12 | 2,428 | Anthropic | Proprietary |
| 161 | GLM-4.5-Air智谱AI | 1389.00 | +/-15 | 1,513 | 智谱AI | MIT |
| 162 | OpenAI o3-mini (high)OpenAI | 1388.00 | +/-18 | 953 | OpenAI | Proprietary |
| 163 | Kimi K2Moonshot AI | 1388.00 | +/-14 | 1,680 | Moonshot AI | Modified MIT |
| 164 | OpenAI o1OpenAI | 1386.00 | +/-10 | 4,569 | OpenAI | Proprietary |
| 165 | nvidia-llama-3.3-nemotron-super-49b-v1.5Nvidia | 1386.00 | +/-39 | 194 | Nvidia | Nvidia Open |
| 166 | trinity-large-thinking Apache 2.0 | 1385.00 | +/-15 | 1,611 | — | — |
| 167 | Claude3-Sonnet (thinking-32k)Anthropic | 1385.00 | +/-11 | 2,782 | Anthropic | Proprietary |
| 168 | OpenAI o3-miniOpenAI | 1381.00 | +/-9 | 4,693 | OpenAI | Proprietary |
| 169 | intellect-3 MIT | 1381.00 | +/-31 | 332 | — | — |
| 170 | llama-3.1-nemotron-ultra-253b-v1Nvidia | 1381.00 | +/-37 | 209 | Nvidia | Nvidia Open Model |
| 171 | GPT OSS 120BOpenAI | 1380.00 | +/-14 | 1,763 | OpenAI | Apache 2.0 |
| 172 | Qwen3-30B-A3B-2507阿里巴巴 | 1378.00 | +/-15 | 1,396 | 阿里巴巴 | Apache 2.0 |
| 173 | mimo-v2-flash (non-thinking)Xiaomi | 1378.00 | +/-11 | 2,819 | Xiaomi | MIT |
| 174 | mimo-v2-flash (thinking)Xiaomi | 1375.00 | +/-23 | 617 | Xiaomi | MIT |
| 175 | nvidia-nemotron-3-super-120b-a12bNvidia | 1374.00 | +/-25 | 518 | Nvidia | NVIDIA Open Model |
| 176 | GPT-4.1OpenAI | 1374.00 | +/-10 | 3,194 | OpenAI | Proprietary |
| 177 | Qwen3-Coder-480B-A35B阿里巴巴 | 1374.00 | +/-15 | 1,607 | 阿里巴巴 | Apache 2.0 |
| 178 | 1373.00 | +/-11 | 2,667 | xAI | Proprietary | |
| 179 | 1371.00 | +/-14 | 1,490 | xAI | Proprietary | |
| 180 | minimax-m1MiniMax | 1369.00 | +/-14 | 1,763 | MiniMax | Apache 2.0 |
| 181 | DeepSeek-V3-0324DeepSeek-AI | 1368.00 | +/-10 | 3,150 | DeepSeek-AI | MIT |
| 182 | GLM-4.7-Flash智谱AI | 1366.00 | +/-21 | 712 | 智谱AI | MIT |
| 183 | Gemini 2.5 Flash-Lite-Preview-09-2025 (no-thinking)Google Deep Mind | 1364.00 | +/-11 | 2,823 | Google Deep Mind | Proprietary |
| 184 | Gemini 2.5 Flash-Lite (thinking)Google Deep Mind | 1364.00 | +/-13 | 2,058 | Google Deep Mind | Proprietary |
| 185 | QwQ-32B阿里巴巴 | 1363.00 | +/-14 | 1,704 | 阿里巴巴 | Apache 2.0 |
| 186 | Claude3-SonnetAnthropic | 1363.00 | +/-10 | 3,339 | Anthropic | Proprietary |
| 187 | Qwen2.5-Max阿里巴巴 | 1363.00 | +/-10 | 3,298 | 阿里巴巴 | Proprietary |
| 188 | GLM-4.5V智谱AI | 1362.00 | +/-34 | 271 | 智谱AI | MIT |
| 189 | trinity-large-preview Apache 2.0 | 1362.00 | +/-14 | 1,870 | — | — |
| 190 | OpenAI o1-miniOpenAI | 1362.00 | +/-8 | 7,499 | OpenAI | Proprietary |
| 191 | Step3StepFunAI | 1361.00 | +/-31 | 348 | StepFunAI | Apache 2.0 |
| 192 | Gemini 2.0 Flash ExperimentalDeepMind | 1356.00 | +/-9 | 4,044 | DeepMind | Proprietary |
| 193 | MiniMax M2MiniMaxAI | 1355.00 | +/-33 | 318 | MiniMaxAI | Apache 2.0 |
| 194 | GPT-4.1 miniOpenAI | 1353.00 | +/-11 | 2,666 | OpenAI | Proprietary |
| 195 | Qwen3-30B-A3B阿里巴巴 | 1352.00 | +/-14 | 1,692 | 阿里巴巴 | Apache 2.0 |
| 196 | Claude 3.5 SonnetAnthropic | 1351.00 | +/-7 | 10,000 | Anthropic | Proprietary |
| 197 | ling-flash-2.0 AntGroup | 1349.00 | +/-27 | 446 | Group | MIT |
| 198 | nvidia-nemotron-3-nano-30b-a3b-bf16Nvidia | 1349.00 | +/-19 | 955 | Nvidia | NVIDIA Open Model |
| 199 | mistral-medium-2505Mistral | 1348.00 | +/-12 | 2,209 | Mistral | Proprietary |
| 200 | hunyuan-turbos-20250416Tencent | 1347.00 | +/-20 | 846 | Tencent | Proprietary |
| 201 | GPT-5-Nano (high)OpenAI | 1346.00 | +/-27 | 492 | OpenAI | Proprietary |
| 202 | Claude 3.5 SonnetAnthropic | 1342.00 | +/-8 | 11,359 | Anthropic | Proprietary |
| 203 | Gemini 1.5 ProGoogle Deep Mind | 1339.00 | +/-7 | 7,610 | Google Deep Mind | Proprietary |
| 204 | Mistral-Small-3.2MistralAI | 1337.00 | +/-18 | 1,040 | MistralAI | Apache 2.0 |
| 205 | nvidia-nemotron-3.5-lightning-30b-a3b-nvfp4Nvidia | 1337.00 | +/-43 | 200 | Nvidia | OpenMDW-1.1 |
| 206 | ring-flash-2.0 AntGroup | 1335.00 | +/-27 | 446 | Group | MIT |
| 207 | GPT OSS 20BOpenAI | 1335.00 | +/-22 | 672 | OpenAI | Apache 2.0 |
| 208 | Nova 2 Lite亚马逊 | 1333.00 | +/-20 | 820 | 亚马逊 | Proprietary |
| 209 | Gemini 2.0 Flash-LiteDeepMind | 1326.00 | +/-10 | 2,814 | DeepMind | Proprietary |
| 210 | qwen-plus-0125Alibaba | 1323.00 | +/-19 | 732 | Alibaba | Proprietary |
| 211 | Gemma 3 - 27B (IT)Google Deep Mind | 1322.00 | +/-10 | 3,551 | Google Deep Mind | Gemma |
| 212 | granite-4.1-8bIBM | 1320.00 | +/-39 | 236 | IBM | Apache 2.0 |
| 213 | llama-3.1-405b-instruct-fp8Meta | 1319.00 | +/-8 | 8,482 | Meta | Llama 3.1 Community |
| 214 | Gemma 3 - 12B (IT)Google Deep Mind | 1318.00 | +/-27 | 389 | Google Deep Mind | Gemma |
| 215 | Llama 4 Maverick InstructFacebook AI研究实验室 | 1317.00 | +/-11 | 2,810 | Facebook AI研究实验室 | Llama 4 |
| 216 | llama-3.1-405b-instruct-bf16Meta | 1315.00 | +/-8 | 5,215 | Meta | Llama 3.1 Community |
| 217 | step-2-16k-exp-202412StepFun | 1312.00 | +/-20 | 642 | StepFun | Proprietary |
| 218 | Claude3-OpusAnthropic | 1312.00 | +/-7 | 25,769 | Anthropic | Proprietary |
| 219 | athene-v2-chat NexusFlow | 1312.00 | +/-10 | 3,412 | — | — |
| 220 | olmo-3-32b-thinkAi2 | 1311.00 | +/-33 | 305 | Ai2 | Apache 2.0 |
| 221 | DeepSeek-V3DeepSeek-AI | 1310.00 | +/-11 | 2,721 | DeepSeek-AI | DeepSeek |
| 222 | C4AI Command A (202503)CohereAI | 1309.00 | +/-9 | 3,961 | CohereAI | CC-BY-NC-4.0 |
| 223 | GPT-4oOpenAI | 1309.00 | +/-8 | 6,826 | OpenAI | Proprietary |
| 224 | Llama 4 Scout InstructFacebook AI研究实验室 | 1308.00 | +/-13 | 1,928 | Facebook AI研究实验室 | Llama |
| 225 | gemini-advanced-0514Google | 1307.00 | +/-10 | 6,395 | Proprietary | |
| 226 | yi-lightning Proprietary | 1305.00 | +/-10 | 3,921 | — | — |
| 227 | GPT-4oOpenAI | 1305.00 | +/-7 | 15,103 | OpenAI | Proprietary |
| 228 | qwen2.5-plus-1127Alibaba | 1304.00 | +/-14 | 1,404 | Alibaba | Proprietary |
| 229 | GPT-4OpenAI | 1303.00 | +/-8 | 13,306 | OpenAI | Proprietary |
| 230 | olmo-3.1-32b-instructAi2 | 1302.00 | +/-23 | 659 | Ai2 | Apache 2.0 |
| 231 | hunyuan-turbos-20250226Tencent | 1301.00 | +/-31 | 238 | Tencent | Proprietary |
| 232 | GPT-4OpenAI | 1299.00 | +/-8 | 12,374 | OpenAI | Proprietary |
| 233 | Gemini 1.5 ProGoogle Deep Mind | 1299.00 | +/-8 | 10,492 | Google Deep Mind | Proprietary |
| 234 | glm-4-plus-0111Zhipu | 1298.00 | +/-19 | 721 | Zhipu | Proprietary |
| 235 | step-1o-turbo-202506StepFun | 1297.00 | +/-24 | 560 | StepFun | Proprietary |
| 236 | gpt-4-turbo-2024-04-09OpenAI | 1296.00 | +/-8 | 13,217 | OpenAI | Proprietary |
| 237 | Qwen2.5-VL-72B-Instruct阿里巴巴 | 1296.00 | +/-8 | 5,415 | 阿里巴巴 | Qwen |
| 238 | Llama3.3-70B-InstructFacebook AI研究实验室 | 1296.00 | +/-8 | 5,768 | Facebook AI研究实验室 | Llama-3.3 |
| 239 | olmo-3.1-32b-thinkAi2 | 1296.00 | +/-27 | 453 | Ai2 | Apache 2.0 |
| 240 | 1294.00 | +/-7 | 8,950 | xAI | Proprietary | |
| 241 | hunyuan-large-2025-02-10Tencent | 1293.00 | +/-24 | 497 | Tencent | Proprietary |
| 242 | deepseek-v2.5-1210DeepSeek | 1292.00 | +/-17 | 1,031 | DeepSeek | DeepSeek |
| 243 | qwen-max-0919Alibaba | 1291.00 | +/-12 | 2,249 | Alibaba | Qwen |
| 244 | hunyuan-standard-2025-02-10Tencent | 1290.00 | +/-24 | 499 | Tencent | Proprietary |
| 245 | gemini-1.5-flash-002Google | 1289.00 | +/-9 | 4,789 | Proprietary | |
| 246 | mistral-large-2407Mistral | 1288.00 | +/-8 | 6,664 | Mistral | Mistral Research |
| 247 | DeepSeek V2.5DeepSeek-AI | 1287.00 | +/-10 | 3,649 | DeepSeek-AI | DeepSeek |
| 248 | glm-4-plusZhipu AI | 1287.00 | +/-10 | 3,599 | Zhipu AI | Proprietary |
| 249 | Claude 3.5 HaikuAnthropic | 1286.00 | +/-8 | 6,319 | Anthropic | Proprietary |
| 250 | Magistral-Medium-2506MistralAI | 1285.00 | +/-26 | 550 | MistralAI | Proprietary |
| 251 | GPT-4OpenAI | 1284.00 | +/-10 | 7,052 | OpenAI | Proprietary |
| 252 | mistral-large-2411Mistral | 1282.00 | +/-9 | 3,574 | Mistral | MRL |
| 253 | ibm-granite-h-smallIBM | 1281.00 | +/-33 | 337 | IBM | Apache 2.0 |
| 254 | hunyuan-large-visionTencent | 1281.00 | +/-30 | 344 | Tencent | Proprietary |
| 255 | hunyuan-turbo-0110Tencent | 1279.00 | +/-31 | 243 | Tencent | Proprietary |
| 256 | Llama3.1-70B-InstructFacebook AI研究实验室 | 1279.00 | +/-17 | 1,041 | Facebook AI研究实验室 | Llama 3.1 |
| 257 | Mistral-Small-3.1-24B-Instruct-2503MistralAI | 1277.00 | +/-13 | 2,094 | MistralAI | Apache 2.0 |
| 258 | GPT-4OpenAI | 1276.00 | +/-8 | 11,181 | OpenAI | Proprietary |
| 259 | GPT-4o miniOpenAI | 1275.00 | +/-7 | 9,319 | OpenAI | Proprietary |
| 260 | GPT-4.1 nanoOpenAI | 1274.00 | +/-23 | 582 | OpenAI | Proprietary |
| 261 | Qwen2-72B-Instruct阿里巴巴 | 1273.00 | +/-9 | 4,835 | 阿里巴巴 | Qianwen LICENSE |
| 262 | 1273.00 | +/-8 | 7,261 | xAI | Proprietary | |
| 263 | deepseek-coder-v2DeepSeek | 1272.00 | +/-14 | 1,858 | DeepSeek | DeepSeek License |
| 264 | llama-3.1-nemotron-51b-instructNvidia | 1271.00 | +/-22 | 507 | Nvidia | Llama 3.1 |
| 265 | Qwen2.5-Coder-32B-Instruct阿里巴巴 | 1270.00 | +/-19 | 725 | 阿里巴巴 | Apache 2.0 |
| 266 | Llama3.1-70B-InstructFacebook AI研究实验室 | 1269.00 | +/-8 | 7,677 | Facebook AI研究实验室 | Llama 3.1 Community |
| 267 | amazon-nova-pro-v1.0Amazon | 1269.00 | +/-10 | 2,978 | Amazon | Proprietary |
| 268 | Phi-4-reasoningMicrosoft Azure | 1265.00 | +/-10 | 2,764 | Microsoft Azure | MIT |
| 269 | llama-3.1-tulu-3-70bAi2 | 1263.00 | +/-25 | 397 | Ai2 | Llama 3.1 |
| 270 | athene-70b-0725 CC-BY-NC-4.0 | 1262.00 | +/-10 | 2,921 | — | — |
| 271 | Mistral Small 24B Instruct 2501MistralAI | 1261.00 | +/-13 | 1,683 | MistralAI | Apache 2.0 |
| 272 | Gemma-3n-E4BGoogle Deep Mind | 1260.00 | +/-15 | 1,553 | Google Deep Mind | Gemma |
| 273 | gemini-1.5-flash-001Google | 1259.00 | +/-8 | 8,392 | Proprietary | |
| 274 | Llama3-70B-InstructFacebook AI研究实验室 | 1258.00 | +/-7 | 20,941 | Facebook AI研究实验室 | Llama 3 Community |
| 275 | Gemma 3 - 4B (IT)Google Deep Mind | 1254.00 | +/-28 | 423 | Google Deep Mind | Gemma |
| 276 | Claude3-SonnetAnthropic | 1253.00 | +/-8 | 13,766 | Anthropic | Proprietary |
| 277 | nemotron-4-340b-instructNvidia | 1252.00 | +/-12 | 2,352 | Nvidia | NVIDIA Open Model |
| 278 | hunyuan-standard-256kTencent | 1250.00 | +/-29 | 361 | Tencent | Proprietary |
| 279 | gemma-2-27b-itGoogle | 1247.00 | +/-7 | 10,170 | Gemma license | |
| 280 | GLM4智谱AI | 1246.00 | +/-16 | 1,191 | 智谱AI | Proprietary |
| 281 | reka-core-20240904 Proprietary | 1246.00 | +/-14 | 1,207 | — | — |
| 282 | jamba-1.5-large Jamba Open | 1245.00 | +/-15 | 1,147 | — | — |
| 283 | mistral-large-2402Mistral | 1244.00 | +/-9 | 7,987 | Mistral | Proprietary |
| 284 | amazon-nova-lite-v1.0Amazon | 1244.00 | +/-11 | 2,511 | Amazon | Proprietary |
| 285 | C4AI Aya Vision 32BCohereAI | 1233.00 | +/-10 | 3,854 | CohereAI | CC-BY-NC-4.0 |
| 286 | reka-flash-20240904 Proprietary | 1232.00 | +/-14 | 1,284 | — | — |
| 287 | Claude3-HaikuAnthropic | 1231.00 | +/-7 | 14,983 | Anthropic | Proprietary |
| 288 | command-r-plus-08-2024Cohere | 1231.00 | +/-14 | 1,467 | Cohere | CC-BY-NC-4.0 |
| 289 | gemini-1.5-flash-8b-001Google | 1230.00 | +/-9 | 5,036 | Proprietary | |
| 290 | Mixtral-8x22B-Instruct-v0.1MistralAI | 1228.00 | +/-9 | 6,778 | MistralAI | Apache 2.0 |
| 291 | olmo-2-0325-32b-instructAi2 | 1227.00 | +/-28 | 375 | Ai2 | Apache-2.0 |
| 292 | amazon-nova-micro-v1.0Amazon | 1223.00 | +/-11 | 2,455 | Amazon | Proprietary |
| 293 | Qwen1.5-110B-Chat阿里巴巴 | 1221.00 | +/-11 | 3,188 | 阿里巴巴 | Qianwen LICENSE |
| 294 | mistral-mediumMistral | 1220.00 | +/-11 | 4,406 | Mistral | Proprietary |
| 295 | gemma-2-9b-itGoogle | 1220.00 | +/-8 | 7,110 | Gemma license | |
| 296 | Phi-3-medium 14B-previewMicrosoft Azure | 1215.00 | +/-11 | 3,238 | Microsoft Azure | MIT |
| 297 | C4AI Command R+CohereAI | 1214.00 | +/-9 | 9,769 | CohereAI | CC-BY-NC-4.0 |
| 298 | ministral-8b-2410Mistral | 1214.00 | +/-20 | 683 | Mistral | MRL |
| 299 | Yi-1.5-34B零一万物 | 1212.00 | +/-11 | 2,985 | 零一万物 | — |
| 300 | reka-flash-21b-20240226-online Proprietary | 1212.00 | +/-14 | 2,028 | — | — |
| 301 | QwQ-32B-Preview阿里巴巴 | 1210.00 | +/-24 | 480 | 阿里巴巴 | Apache 2.0 |
| 302 | Qwen1.5-72B-Chat阿里巴巴 | 1209.00 | +/-10 | 5,327 | 阿里巴巴 | Qianwen LICENSE |
| 303 | gemma-2-9b-it-simpo MIT | 1207.00 | +/-15 | 1,285 | — | — |
| 304 | command-r-08-2024Cohere | 1207.00 | +/-14 | 1,601 | Cohere | CC-BY-NC-4.0 |
| 305 | InternLM2-Base-20B上海人工智能实验室 | 1206.00 | +/-15 | 1,387 | 上海人工智能实验室 | — |
| 306 | llama-3.1-tulu-3-8bAi2 | 1206.00 | +/-26 | 363 | Ai2 | Llama 3.1 |
| 307 | gpt-3.5-turbo-1106OpenAI | 1204.00 | +/-15 | 2,134 | OpenAI | Proprietary |
| 308 | Gemini-proDeepMind | 1201.00 | +/-19 | 993 | DeepMind | Proprietary |
| 309 | C4AI Aya Vision 8BCohereAI | 1201.00 | +/-15 | 1,307 | CohereAI | CC-BY-NC-4.0 |
| 310 | gpt-3.5-turbo-0125OpenAI | 1201.00 | +/-9 | 8,626 | OpenAI | Proprietary |
| 311 | qwen1.5-32b-chatAlibaba | 1201.00 | +/-12 | 2,649 | Alibaba | Qianwen LICENSE |
| 312 | reka-flash-21b-20240226 Proprietary | 1199.00 | +/-11 | 3,363 | — | — |
| 313 | granite-3.0-8b-instructIBM | 1198.00 | +/-19 | 873 | IBM | Apache 2.0 |
| 314 | gemini-pro-dev-apiGoogle | 1198.00 | +/-14 | 2,274 | Proprietary | |
| 315 | granite-3.1-2b-instructIBM | 1198.00 | +/-26 | 391 | IBM | Apache 2.0 |
| 316 | zephyr-orpo-141b-A35b-v0.1 Apache 2.0 | 1196.00 | +/-22 | 589 | — | — |
| 317 | dbrx-instruct-preview DBRX LICENSE | 1196.00 | +/-12 | 4,001 | — | — |
| 318 | Phi-3-mini 3.8BMicrosoft Azure | 1193.00 | +/-14 | 1,568 | Microsoft Azure | MIT |
| 319 | Phi-3-small 7BMicrosoft Azure | 1193.00 | +/-13 | 2,092 | Microsoft Azure | MIT |
| 320 | Llama3-8B-InstructFacebook AI研究实验室 | 1193.00 | +/-8 | 14,252 | Facebook AI研究实验室 | Llama 3 Community |
| 321 | mixtral-8x7b-instruct-v0.1Mistral | 1191.00 | +/-9 | 9,663 | Mistral | Apache 2.0 |
| 322 | Llama3.1-8B-InstructFacebook AI研究实验室 | 1190.00 | +/-28 | 382 | Facebook AI研究实验室 | Apache 2.0 |
| 323 | Llama3.1-8B-InstructFacebook AI研究实验室 | 1189.00 | +/-8 | 7,135 | Facebook AI研究实验室 | Llama 3.1 Community |
| 324 | jamba-1.5-mini Jamba Open | 1186.00 | +/-16 | 1,094 | — | — |
| 325 | command-rCohere | 1176.00 | +/-10 | 6,682 | Cohere | CC-BY-NC-4.0 |
| 326 | Qwen3-VL-2B阿里巴巴 | 1169.00 | +/-19 | 908 | 阿里巴巴 | Apache 2.0 |
| 327 | Qwen1.5-14B-Chat阿里巴巴 | 1167.00 | +/-14 | 2,184 | 阿里巴巴 | Qianwen LICENSE |
| 328 | llama-3.2-3b-instructMeta | 1165.00 | +/-16 | 1,136 | Meta | Llama 3.2 |
| 329 | gemma-2-2b-itGoogle | 1164.00 | +/-8 | 6,599 | Gemma license | |
| 330 | snowflake-arctic-instruct Apache 2.0 | 1163.00 | +/-11 | 4,793 | — | — |
| 331 | Gemma 1.1-7B-ITGoogle Research | 1162.00 | +/-11 | 3,039 | Google Research | Gemma license |
| 332 | openchat-3.5-0106 Apache-2.0 | 1158.00 | +/-14 | 1,726 | — | — |
| 333 | starling-lm-7b-beta Apache-2.0 | 1157.00 | +/-14 | 1,973 | — | — |
| 334 | WizardLM-70B-V1.0WizardLM Team | 1157.00 | +/-19 | 903 | WizardLM Team | Llama 2 Community |
| 335 | DeepSeek LLM 67B ChatDeepSeek-AI | 1156.00 | +/-24 | 576 | DeepSeek-AI | DeepSeek License |
| 336 | smollm2-1.7b-instruct Apache 2.0 | 1152.00 | +/-33 | 271 | — | — |
| 337 | openhermes-2.5-mistral-7b Apache-2.0 | 1152.00 | +/-20 | 697 | — | — |
| 338 | Yi-34B零一万物 | 1150.00 | +/-13 | 2,043 | 零一万物 | — |
| 339 | Phi-3-mini 3.8BMicrosoft Azure | 1150.00 | +/-12 | 2,564 | Microsoft Azure | MIT |
| 340 | LLaMA2 70BFacebook AI研究实验室 | 1144.00 | +/-19 | 888 | Facebook AI研究实验室 | — |
| 341 | Phi-3-mini 3.8BMicrosoft Azure | 1139.00 | +/-13 | 2,813 | Microsoft Azure | MIT |
| 342 | llama-2-70b-chatMeta | 1136.00 | +/-10 | 4,740 | Meta | Llama 2 Community |
| 343 | Mistral-7B-Instruct-v0.2MistralAI | 1127.00 | +/-12 | 2,605 | MistralAI | Apache-2.0 |
| 344 | starling-lm-7b-alpha CC-BY-NC-4.0 | 1126.00 | +/-16 | 1,300 | — | — |
| 345 | Qwen-14B-Chat阿里巴巴 | 1126.00 | +/-24 | 534 | 阿里巴巴 | Qianwen LICENSE |
| 346 | openchat-3.5 Apache-2.0 | 1125.00 | +/-18 | 945 | — | — |
| 347 | dolphin-2.2.1-mistral-7b Apache-2.0 | 1125.00 | +/-32 | 219 | — | — |
| 348 | llama-3.2-1b-instructMeta | 1123.00 | +/-16 | 1,162 | Meta | Llama 3.2 |
| 349 | Qwen1.5-7B-Chat阿里巴巴 | 1121.00 | +/-21 | 690 | 阿里巴巴 | Qianwen LICENSE |
| 350 | Gemma 7B - ItGoogle Research | 1119.00 | +/-17 | 1,120 | Google Research | Gemma license |
| 351 | PaLM 2Google Research | 1116.00 | +/-19 | 901 | Google Research | Proprietary |
| 352 | Vicuna 33BLM-SYS | 1116.00 | +/-13 | 2,663 | LM-SYS | — |
| 353 | llama2-70b-steerlm-chatNvidia | 1114.00 | +/-27 | 440 | Nvidia | Llama 2 Community |
| 354 | Baichuan2-13B-Chat百川智能 | 1110.00 | +/-13 | 2,218 | 百川智能 | Llama 2 Community |
| 355 | Gemma 1.1-2B-ITGoogle Research | 1110.00 | +/-16 | 1,355 | Google Research | Gemma license |
| 356 | CodeLLaMA-34BFacebook AI研究实验室 | 1109.00 | +/-19 | 770 | Facebook AI研究实验室 | Llama 2 Community |
| 357 | solar-10.7b-instruct-v1.0 CC-BY-NC-4.0 | 1109.00 | +/-22 | 604 | — | — |
| 358 | mpt-30b-chat CC-BY-NC-SA-4.0 | 1095.00 | +/-34 | 242 | — | — |
| 359 | nous-hermes-2-mixtral-8x7b-dpo Apache-2.0 | 1093.00 | +/-21 | 628 | — | — |
| 360 | Qwen1.5-4B-Chat阿里巴巴 | 1087.00 | +/-18 | 988 | 阿里巴巴 | Qianwen LICENSE |
| 361 | Baichuan2-7B-Chat百川智能 | 1086.00 | +/-14 | 1,656 | 百川智能 | Llama 2 Community |
| 362 | stripedhyena-nous-7b Apache 2.0 | 1085.00 | +/-20 | 676 | — | — |
| 363 | vicuna-13b Llama 2 Community | 1083.00 | +/-14 | 2,146 | — | — |
| 364 | Mistral 7B InstructMistralAI | 1082.00 | +/-19 | 974 | MistralAI | Apache 2.0 |
| 365 | zephyr-7b-beta MIT | 1082.00 | +/-17 | 1,250 | — | — |
| 366 | guanaco-33b Non-commercial | 1080.00 | +/-32 | 280 | — | — |
| 367 | Gemma 2B - ItGoogle Research | 1072.00 | +/-22 | 597 | Google Research | Gemma license |
| 368 | wizardlm-13bMicrosoft | 1064.00 | +/-21 | 669 | Microsoft | Llama 2 Community |
| 369 | olmo-7b-instructAi2 | 1053.00 | +/-19 | 848 | Ai2 | Apache-2.0 |
| 370 | vicuna-7b Llama 2 Community | 1048.00 | +/-22 | 658 | — | — |
| 371 | ChatGLM3-6B智谱AI | 1042.00 | +/-23 | 576 | 智谱AI | — |
| 372 | GPT4All 13BNomic AI | 999.00 | +/-37 | 211 | Nomic AI | — |
| 373 | alpaca-13b Non-commercial | 994.00 | +/-23 | 652 | — | — |
| 374 | mpt-7b-chat CC-BY-NC-SA-4.0 | 986.00 | +/-25 | 471 | — | — |
| 375 | RWKV-4-Raven-14B Apache 2.0 | 984.00 | +/-24 | 544 | — | — |
| 376 | Koala达摩院 | 980.00 | +/-21 | 751 | 达摩院 | — |
| 377 | ChatGLM-6B智谱AI | 977.00 | +/-26 | 525 | 智谱AI | — |
| 378 | ChatGLM2-6B智谱AI | 972.00 | +/-35 | 227 | 智谱AI | — |
| 379 | oasst-pythia-12b Apache 2.0 | 961.00 | +/-22 | 687 | — | — |
| 380 | dolly-v2-12b MIT | 953.00 | +/-29 | 370 | — | — |
| 381 | LLaMA 13BFacebook AI研究实验室 | 922.00 | +/-33 | 252 | Facebook AI研究实验室 | Non-commercial |
| 382 | fastchat-t5-3b Apache 2.0 | 920.00 | +/-26 | 462 | — | — |
| 383 | stablelm-tuned-alpha-7b CC-BY-NC-SA-4.0 | 890.00 | +/-29 | 353 | — | — |
数据仅供参考,以官方来源为准。模型名称旁的链接可跳转到 DataLearner 模型详情页。
常见问题 (FAQ)
什么是 LMArena Math Arena?
LMArena Math Arena 是 LMArena 旗下专注于数学推理能力的匿名评测平台。用户提交真实数学问题(如代数、几何、竞赛数学等),系统将不同模型的解题过程并排展示(隐藏模型名称),由用户投票选出更好的解答,最终通过 Elo 算法汇总形成动态排行榜。
Math Arena 与 MATH-500、AIME 等静态基准有什么区别?
MATH-500、AIME、AMC 等静态基准使用固定题目集和自动评分,可重现性强但容易被针对性优化("刷榜")。Math Arena 来自真实用户的开放式数学问题,测试内容不固定,更能反映模型在实际数学场景中的自然表现,两者互为补充。
思考模型(Thinking Model)在数学 Arena 中表现更好吗?
整体而言,具备思维链(Chain-of-Thought)或扩展推理能力的模型在数学 Arena 中往往排名更高。Claude Opus 系列 Thinking 模式、GPT 高算力模式以及 DeepSeek 思考版本均在榜单前列,说明延长推理时间对数学问题的解答质量有显著提升。
国产大模型在数学能力方面表现如何?
DeepSeek、Qwen3 系列、GLM 等国产模型在 Math Arena 表现亮眼,已跻身全球前列。DeepSeek 以 MIT 协议开源,Qwen3-235B 等系列支持中文数学场景,是选择开源数学推理模型的重要参考。



























