Text Generation Arena Leaderboard
The latest AI text generation leaderboard based on LMArena anonymous user voting. Covers Elo scores, confidence intervals, and vote counts for leading language models.
Top Model
Claude Fable 5
Top Score
1,507
Model Count
385
Data version
2026年08月06日
Data source: LM Arena
About This Leaderboard
This leaderboard ranks the strongest AI models for text generation. Data comes from LMArena (formerly LMSYS Chatbot Arena), the world's largest crowdsourced AI evaluation platform. Users chat with two anonymous models side-by-side and vote for the better response — rankings are determined entirely by real user preferences, not lab benchmarks.
Methodology Overview
Blind testing: Users chat with two anonymous models and vote based on response quality, eliminating brand bias.
Elo scoring: Using the Bradley-Terry model (adapted from chess Elo ratings) to calculate each model's strength score from battle outcomes. Higher scores mean users more frequently prefer that model.
Broad scenario coverage: Testing spans coding, creative writing, math reasoning, Q&A, role-playing, and more.
DataLearner provides in-depth analysis on top of the raw data, linking leaderboard models to the DataLearner model database so you can quickly access model details, API pricing, benchmark scores, and more.
Ranking Table
| Rank | Model | Score | 95% CI | Votes | Organization | License |
|---|---|---|---|---|---|---|
Claude Fable 5Anthropic | 1,507 | +/-6 | 19,390 | Anthropic | Proprietary | |
Claude Opus 4.6 (thinking)Anthropic | 1,505 | +/-4 | 69,336 | Anthropic | Proprietary | |
Opus 4.7 (thinking)Anthropic | 1,502 | +/-4 | 57,018 | Anthropic | Proprietary | |
| 4 | muse-spark-1.2 (xHigh)Meta | 1,498 | +/-13 | 2,057 | Meta | Proprietary |
| 5 | Claude Opus 4.6Anthropic | 1,497 | +/-3 | 73,158 | Anthropic | Proprietary |
| 6 | Qwen3.8-Max阿里巴巴 | 1,497 | +/-9 | 4,662 | 阿里巴巴 | Proprietary |
| 7 | Opus 4.7Anthropic | 1,493 | +/-4 | 58,135 | Anthropic | Proprietary |
| 8 | Claude Opus 5 (high)Anthropic | 1,493 | +/-6 | 14,902 | Anthropic | Proprietary |
| 9 | Claude Opus 5 (max)Anthropic | 1,488 | +/-8 | 7,137 | Anthropic | Proprietary |
| 10 | Muse SparkFacebook AI研究实验室 | 1,488 | +/-6 | 13,476 | Facebook AI研究实验室 | Proprietary |
| 11 | Muse Spark 1.1Facebook AI研究实验室 | 1,487 | +/-6 | 14,227 | Facebook AI研究实验室 | Proprietary |
| 12 | Gemini 3.1 Pro PreviewGoogle Deep Mind | 1,487 | +/-3 | 91,328 | Google Deep Mind | Proprietary |
| 13 | Gemini 3 ProGoogle Deep Mind | 1,486 | +/-4 | 41,242 | Google Deep Mind | Proprietary |
| 14 | Kimi K3 (max)Moonshot AI | 1,485 | +/-7 | 9,861 | Moonshot AI | Kimi K3 license |
| 15 | Gemini 3.6 FlashGoogle Deep Mind | 1,485 | +/-6 | 10,968 | Google Deep Mind | Proprietary |
| 16 | Claude Opus 4.8 (thinking)Anthropic | 1,483 | +/-5 | 37,546 | Anthropic | Proprietary |
| 17 | GPT-5.5 (high)OpenAI | 1,482 | +/-4 | 52,142 | OpenAI | Proprietary |
| 18 | gpt-5.6-sol-xhighOpenAI | 1,482 | +/-6 | 12,839 | OpenAI | Proprietary |
| 19 | GPT-5.4 (high)OpenAI | 1,477 | +/-4 | 60,156 | OpenAI | Proprietary |
| 20 | GPT-5.5OpenAI | 1,477 | +/-4 | 53,407 | OpenAI | Proprietary |
| 21 | gemini-3.5-flash-highGoogle | 1,476 | +/-5 | 22,797 | Proprietary | |
| 22 | GPT-5.2 Chat (0210)OpenAI | 1,476 | +/-4 | 34,020 | OpenAI | Proprietary |
| 23 | Qwen3.7-Max-Preview阿里巴巴 | 1,475 | +/-10 | 3,695 | 阿里巴巴 | Proprietary |
| 24 | 1,474 | +/-5 | 26,553 | xAI | Proprietary | |
| 25 | gemini-3.5-flash-mediumGoogle | 1,474 | +/-5 | 21,100 | Proprietary | |
| 26 | GPT-5.5 InstantOpenAI | 1,474 | +/-5 | 25,642 | OpenAI | Proprietary |
| 27 | Gemini 3.0 FlashGoogle Deep Mind | 1,473 | +/-4 | 30,625 | Google Deep Mind | Proprietary |
| 28 | Claude Opus 4.8Anthropic | 1,473 | +/-5 | 38,131 | Anthropic | Proprietary |
| 29 | Claude Opus 4 (thinking-32k)Anthropic | 1,473 | +/-4 | 36,960 | Anthropic | Proprietary |
| 30 | Claude Sonnet 4.6Anthropic | 1,472 | +/-4 | 63,453 | Anthropic | Proprietary |
| 31 | 1,472 | +/-4 | 61,639 | xAI | Proprietary | |
| 32 | 1,471 | +/-4 | 60,322 | xAI | Proprietary | |
| 33 | glm-5.2-maxZ.ai | 1,471 | +/-5 | 24,250 | Z.ai | MIT |
| 34 | Claude Opus 4Anthropic | 1,469 | +/-3 | 70,504 | Anthropic | Proprietary |
| 35 | 1,468 | +/-6 | 14,883 | xAI | Proprietary | |
| 36 | GLM 5.1智谱AI | 1,468 | +/-4 | 36,716 | 智谱AI | MIT |
| 37 | ERNIE-5.1-Preview百度 | 1,467 | +/-5 | 36,820 | 百度 | Proprietary |
| 38 | gpt-5.6-terra-xhighOpenAI | 1,467 | +/-6 | 13,275 | OpenAI | Proprietary |
| 39 | mimo-v2.5-proXiaomi | 1,467 | +/-4 | 47,805 | Xiaomi | MIT |
| 40 | 1,466 | +/-3 | 65,044 | xAI | Proprietary | |
| 41 | Qwen3.5 Max Preview阿里巴巴 | 1,465 | +/-5 | 21,274 | 阿里巴巴 | Proprietary |
| 42 | GPT-5.4OpenAI | 1,465 | +/-4 | 63,097 | OpenAI | Proprietary |
| 43 | Claude Sonnet 4.5 (high)Anthropic | 1,463 | +/-5 | 20,311 | Anthropic | Proprietary |
| 44 | Kimi K2.6Moonshot AI | 1,461 | +/-5 | 37,236 | Moonshot AI | Modified MIT |
| 45 | Qwen3.6-Max-Preview阿里巴巴 | 1,460 | +/-8 | 5,146 | 阿里巴巴 | Proprietary |
| 46 | 1,460 | +/-3 | 67,246 | xAI | Proprietary | |
| 47 | Gemini 3.5 Flash-LiteGoogle Deep Mind | 1,459 | +/-6 | 10,847 | Google Deep Mind | Proprietary |
| 48 | Gemini 3.0 Flash (minimal)Google Deep Mind | 1,459 | +/-3 | 85,546 | Google Deep Mind | Proprietary |
| 49 | Qwen3.7-Plus阿里巴巴 | 1,458 | +/-5 | 28,111 | 阿里巴巴 | Proprietary |
| 50 | DeepSeek-V4-ProDeepSeek-AI | 1,457 | +/-4 | 51,266 | DeepSeek-AI | MIT |
| 51 | GLM-5智谱AI | 1,457 | +/-4 | 27,591 | 智谱AI | MIT |
| 52 | DOLA Seed 2.0 Pro字节跳动Seed团队 | 1,456 | +/-4 | 73,874 | 字节跳动Seed团队 | Proprietary |
| 53 | Claude Sonnet 4.5 (thinking-32k)Anthropic | 1,456 | +/-3 | 81,888 | Anthropic | Proprietary |
| 54 | deepseek-v4-pro-high-previewDeepSeek | 1,456 | +/-4 | 48,756 | DeepSeek | MIT |
| 55 | Claude Sonnet 4.5Anthropic | 1,455 | +/-3 | 80,284 | Anthropic | Proprietary |
| 56 | GPT-5.1 Pro (high)OpenAI | 1,454 | +/-4 | 40,680 | OpenAI | Proprietary |
| 57 | Hy3腾讯AI实验室 | 1,453 | +/-10 | 3,941 | 腾讯AI实验室 | Apache 2.0 |
| 58 | Gemma 4 31BDeepMind | 1,451 | +/-8 | 5,839 | DeepMind | Apache 2.0 |
| 59 | Kimi K2 ThinkingMoonshot AI | 1,451 | +/-3 | 68,179 | Moonshot AI | Modified MIT |
| 60 | gpt-5.6-luna-xhighOpenAI | 1,450 | +/-6 | 13,651 | OpenAI | Proprietary |
| 61 | Opus 4.1 (thinking-16k)Anthropic | 1,449 | +/-3 | 49,720 | Anthropic | Proprietary |
| 62 | ERNIE 5.0 Preview (1203)百度 | 1,449 | +/-7 | 9,727 | 百度 | Proprietary |
| 63 | GPT-5.3 ChatOpenAI | 1,449 | +/-4 | 32,583 | OpenAI | Proprietary |
| 64 | GPT-5.4 mini (high)OpenAI | 1,448 | +/-4 | 58,984 | OpenAI | Proprietary |
| 65 | mimo-v2-proXiaomi | 1,448 | +/-5 | 24,211 | Xiaomi | Proprietary |
| 66 | Opus 4.1Anthropic | 1,447 | +/-3 | 77,192 | Anthropic | Proprietary |
| 67 | ERNIE 5.0 (0110)百度 | 1,447 | +/-4 | 35,015 | 百度 | Proprietary |
| 68 | Gemini 2.5 ProGoogle Deep Mind | 1,446 | +/-2 | 124,014 | Google Deep Mind | Proprietary |
| 69 | GPT-4.5OpenAI | 1,445 | +/-6 | 14,547 | OpenAI | Proprietary |
| 70 | 1,445 | +/-5 | 34,343 | MiniMaxAI | MiniMax Community License | |
| 71 | Qwen 3.6 Plus Preview阿里巴巴 | 1,443 | +/-4 | 45,033 | 阿里巴巴 | Proprietary |
| 72 | GPT-4o(2025-03-27)OpenAI | 1,443 | +/-3 | 82,353 | OpenAI | Proprietary |
| 73 | Qwen3.5-397B-A17B阿里巴巴 | 1,442 | +/-4 | 63,719 | 阿里巴巴 | Apache 2.0 |
| 74 | GLM-4.7智谱AI | 1,442 | +/-6 | 12,077 | 智谱AI | MIT |
| 75 | 1,442 | +/-4 | 53,234 | xAI | Proprietary | |
| 76 | InklingThinking Machines Lab | 1,442 | +/-6 | 12,086 | Thinking Machines Lab | Apache 2.0 |
| 77 | GPT-5.1OpenAI | 1,439 | +/-4 | 43,306 | OpenAI | Proprietary |
| 78 | Gemma 4 26B A4BDeepMind | 1,438 | +/-8 | 5,755 | DeepMind | Apache 2.0 |
| 79 | deepseek-v4-flash-high-previewDeepSeek | 1,438 | +/-4 | 48,316 | DeepSeek | MIT |
| 80 | GPT-5.2 Pro (high)OpenAI | 1,438 | +/-4 | 47,621 | OpenAI | Proprietary |
| 81 | DeepSeek-V4-FlashDeepSeek-AI | 1,436 | +/-4 | 48,561 | DeepSeek-AI | MIT |
| 82 | GPT-5.2OpenAI | 1,435 | +/-3 | 78,662 | OpenAI | Proprietary |
| 83 | longcat-flash-chat-2602-expMeituan | 1,435 | +/-5 | 27,706 | Meituan | Proprietary |
| 84 | Qwen3 Max (Preview)阿里巴巴 | 1,435 | +/-5 | 27,670 | 阿里巴巴 | Proprietary |
| 85 | mimo-v2.5Xiaomi | 1,434 | +/-4 | 44,200 | Xiaomi | MIT |
| 86 | GPT-5-Pro (high)OpenAI | 1,434 | +/-5 | 31,877 | OpenAI | Proprietary |
| 87 | GLM-5V-Turbo智谱AI | 1,433 | +/-7 | 9,288 | 智谱AI | Proprietary |
| 88 | Gemini 3.1 Flash-LiteGoogle Deep Mind | 1,432 | +/-4 | 60,041 | Google Deep Mind | Proprietary |
| 89 | Kimi K2.5 InstantMoonshot AI | 1,431 | +/-7 | 8,138 | Moonshot AI | Modified MIT |
| 90 | 1,431 | +/-3 | 56,402 | xAI | Proprietary | |
| 91 | OpenAI o3OpenAI | 1,431 | +/-4 | 59,661 | OpenAI | Proprietary |
| 92 | mimo-v2-omniXiaomi | 1,430 | +/-6 | 19,216 | Xiaomi | Proprietary |
| 93 | Kimi K2 Thinking (thinking-turbo)Moonshot AI | 1,430 | +/-3 | 61,677 | Moonshot AI | Modified MIT |
| 94 | Mistral Medium 3.5MistralAI | 1,427 | +/-7 | 10,887 | MistralAI | Modified MIT |
| 95 | GPT-5.2 Chat (0210)OpenAI | 1,427 | +/-4 | 31,517 | OpenAI | Proprietary |
| 96 | Nemotron 3 UltraNVIDIA | 1,426 | +/-7 | 10,564 | NVIDIA | OpenMDW-1.1 |
| 97 | amazon-nova-experimental-chat-26-02-10Amazon | 1,426 | +/-10 | 3,392 | Amazon | Proprietary |
| 98 | DeepSeek V3.2DeepSeek-AI | 1,425 | +/-4 | 46,999 | DeepSeek-AI | MIT |
| 99 | GLM-4.6智谱AI | 1,425 | +/-4 | 35,589 | 智谱AI | MIT |
| 100 | DeepSeek V3.2-Exp (thinking)DeepSeek-AI | 1,425 | +/-7 | 9,051 | DeepSeek-AI | MIT |
| 101 | Claude Opus 4 (thinking-16k)Anthropic | 1,425 | +/-4 | 36,832 | Anthropic | Proprietary |
| 102 | qwen3-max-2025-09-23Alibaba | 1,424 | +/-6 | 9,138 | Alibaba | Proprietary |
| 103 | Qwen3-235B-A22B-2507阿里巴巴 | 1,423 | +/-3 | 96,721 | 阿里巴巴 | Apache 2.0 |
| 104 | DeepSeek V3.2-Exp (thinking)DeepSeek-AI | 1,423 | +/-4 | 40,832 | DeepSeek-AI | MIT |
| 105 | DeepSeek V3.2-ExpDeepSeek-AI | 1,423 | +/-6 | 11,902 | DeepSeek-AI | MIT |
| 106 | DeepSeek-R1-0528DeepSeek-AI | 1,422 | +/-6 | 18,422 | DeepSeek-AI | MIT |
| 107 | 1,421 | +/-8 | 6,795 | xAI | Proprietary | |
| 108 | ERNIE 5.0百度 | 1,419 | +/-9 | 4,702 | 百度 | Proprietary |
| 109 | Kimi K2 0905Moonshot AI | 1,418 | +/-6 | 11,756 | Moonshot AI | Modified MIT |
| 110 | Kimi K2Moonshot AI | 1,418 | +/-5 | 27,589 | Moonshot AI | Modified MIT |
| 111 | DeepSeek-V3.1DeepSeek-AI | 1,418 | +/-6 | 14,942 | DeepSeek-AI | MIT |
| 112 | DeepSeek-V3.1 Terminus (thinking)DeepSeek-AI | 1,417 | +/-10 | 3,454 | DeepSeek-AI | MIT |
| 113 | Qwen3.5-122B-A10B阿里巴巴 | 1,417 | +/-4 | 28,208 | 阿里巴巴 | Apache 2.0 |
| 114 | DeepSeek-V3.1 (thinking)DeepSeek-AI | 1,417 | +/-7 | 11,715 | DeepSeek-AI | MIT |
| 115 | 1,416 | +/-4 | 55,513 | MiniMaxAI | Modified MIT | |
| 116 | amazon-nova-experimental-chat-26-01-10Amazon | 1,415 | +/-10 | 3,392 | Amazon | Proprietary |
| 117 | Mistral Large 3MistralAI | 1,415 | +/-3 | 56,503 | MistralAI | Apache 2.0 |
| 118 | Qwen3-VL-235B-A22B-Instruct阿里巴巴 | 1,415 | +/-6 | 11,483 | 阿里巴巴 | Apache 2.0 |
| 119 | DeepSeek-V3.1 TerminusDeepSeek-AI | 1,415 | +/-10 | 3,688 | DeepSeek-AI | MIT |
| 120 | GPT-4.1OpenAI | 1,414 | +/-4 | 50,899 | OpenAI | Proprietary |
| 121 | Claude Opus 4Anthropic | 1,413 | +/-4 | 44,134 | Anthropic | Proprietary |
| 122 | Haiku 4.5Anthropic | 1,412 | +/-3 | 114,721 | Anthropic | Proprietary |
| 123 | hunyuan-hy3-previewTencent | 1,412 | +/-8 | 6,555 | Tencent | tencent-hunyuan-community |
| 124 | 1,411 | +/-4 | 32,861 | xAI | Proprietary | |
| 125 | GLM-4.5智谱AI | 1,411 | +/-5 | 24,269 | 智谱AI | MIT |
| 126 | Gemini 2.5 FlashGoogle Deep Mind | 1,410 | +/-2 | 123,929 | Google Deep Mind | Proprietary |
| 127 | 1,409 | +/-4 | 41,304 | xAI | Proprietary | |
| 128 | Magistral-Medium-2506MistralAI | 1,409 | +/-3 | 93,390 | MistralAI | Proprietary |
| 129 | Qwen3.5-27B阿里巴巴 | 1,408 | +/-4 | 27,052 | 阿里巴巴 | Apache 2.0 |
| 130 | Gemini 2.5 Flash-Preview-09-2025Google Deep Mind | 1,404 | +/-4 | 32,849 | Google Deep Mind | Proprietary |
| 131 | 1,404 | +/-5 | 18,688 | xAI | Proprietary | |
| 132 | qwen3-235b-a22b-no-thinkingAlibaba | 1,403 | +/-5 | 38,137 | Alibaba | Apache 2.0 |
| 133 | GPT-5.4 nano (high)OpenAI | 1,403 | +/-4 | 57,979 | OpenAI | Proprietary |
| 134 | OpenAI o1OpenAI | 1,402 | +/-4 | 27,807 | OpenAI | Proprietary |
| 135 | Qwen3-Next阿里巴巴 | 1,401 | +/-5 | 22,846 | 阿里巴巴 | Apache 2.0 |
| 136 | longcat-flash-chatMeituan | 1,401 | +/-6 | 11,379 | Meituan | MIT |
| 137 | Claude Sonnet 4 (thinking-32k)Anthropic | 1,399 | +/-4 | 35,044 | Anthropic | Proprietary |
| 138 | qwen3-235b-a22b-thinking-2507Alibaba | 1,399 | +/-7 | 8,975 | Alibaba | Apache 2.0 |
| 139 | DeepSeek-R1DeepSeek-AI | 1,398 | +/-5 | 18,524 | DeepSeek-AI | MIT |
| 140 | Step 3.5 FlashStepFunAI | 1,397 | +/-4 | 57,847 | StepFunAI | Proprietary |
| 141 | DeepSeek-V3-0324DeepSeek-AI | 1,396 | +/-4 | 45,445 | DeepSeek-AI | MIT |
| 142 | Qwen3-VL-235B-A22B-Instruct (thinking)阿里巴巴 | 1,395 | +/-7 | 7,934 | 阿里巴巴 | Apache 2.0 |
| 143 | Qwen3.5-35B-A3B阿里巴巴 | 1,395 | +/-4 | 28,864 | 阿里巴巴 | Apache 2.0 |
| 144 | hunyuan-vision-1.5-thinkingTencent | 1,395 | +/-12 | 2,214 | Tencent | Proprietary |
| 145 | Step 3.5 FlashStepFunAI | 1,394 | +/-4 | 56,979 | StepFunAI | Apache 2.0 |
| 146 | amazon-nova-experimental-chat-12-10Amazon | 1,394 | +/-10 | 3,672 | Amazon | Proprietary |
| 147 | mimo-v2-flash (non-thinking)Xiaomi | 1,393 | +/-4 | 46,248 | Xiaomi | MIT |
| 148 | GPT-5-mini (high)OpenAI | 1,390 | +/-5 | 26,987 | OpenAI | Proprietary |
| 149 | OpenAI o4 - miniOpenAI | 1,390 | +/-4 | 45,383 | OpenAI | Proprietary |
| 150 | 1,390 | +/-4 | 40,688 | MiniMaxAI | Modified MIT | |
| 151 | Claude Sonnet 4Anthropic | 1,389 | +/-4 | 40,242 | Anthropic | Proprietary |
| 152 | OpenAI o1OpenAI | 1,389 | +/-5 | 31,122 | OpenAI | Proprietary |
| 153 | Claude3-Sonnet (thinking-32k)Anthropic | 1,388 | +/-4 | 38,785 | Anthropic | Proprietary |
| 154 | Qwen3-Coder-480B-A35B阿里巴巴 | 1,388 | +/-5 | 25,674 | 阿里巴巴 | Apache 2.0 |
| 155 | mimo-v2-flash (thinking)Xiaomi | 1,387 | +/-6 | 10,909 | Xiaomi | MIT |
| 156 | mistral-medium-2505Mistral | 1,387 | +/-5 | 33,178 | Mistral | Proprietary |
| 157 | Hunyuan-T1腾讯AI实验室 | 1,387 | +/-9 | 4,695 | 腾讯AI实验室 | Proprietary |
| 158 | 1,384 | +/-5 | 17,057 | MiniMaxAI | MIT | |
| 159 | Qwen3-30B-A3B-2507阿里巴巴 | 1,383 | +/-5 | 23,691 | 阿里巴巴 | Apache 2.0 |
| 160 | GPT-4.1 miniOpenAI | 1,383 | +/-4 | 39,277 | OpenAI | Proprietary |
| 161 | hunyuan-turbos-20250416Tencent | 1,382 | +/-6 | 10,715 | Tencent | Proprietary |
| 162 | Gemini 2.5 Flash-Lite-Preview-09-2025 (no-thinking)Google Deep Mind | 1,380 | +/-3 | 47,142 | Google Deep Mind | Proprietary |
| 163 | trinity-large-preview Apache 2.0 | 1,379 | +/-4 | 29,651 | — | — |
| 164 | GLM-4.6V智谱AI | 1,377 | +/-11 | 2,797 | 智谱AI | MIT |
| 165 | Qwen3-235B-A22B阿里巴巴 | 1,375 | +/-5 | 26,241 | 阿里巴巴 | Apache 2.0 |
| 166 | Gemini 2.5 Flash-Lite (thinking)Google Deep Mind | 1,375 | +/-5 | 32,839 | Google Deep Mind | Proprietary |
| 167 | Qwen2.5-Max阿里巴巴 | 1,374 | +/-4 | 32,590 | 阿里巴巴 | Proprietary |
| 168 | Claude 3.5 SonnetAnthropic | 1,373 | +/-3 | 88,291 | Anthropic | Proprietary |
| 169 | GLM-4.5-Air智谱AI | 1,373 | +/-4 | 31,040 | 智谱AI | MIT |
| 170 | Claude3-SonnetAnthropic | 1,372 | +/-4 | 43,140 | Anthropic | Proprietary |
| 171 | Qwen3-Next (thinking)阿里巴巴 | 1,369 | +/-6 | 13,661 | 阿里巴巴 | Apache 2.0 |
| 172 | trinity-large-thinking Apache 2.0 | 1,369 | +/-5 | 28,771 | — | — |
| 173 | GLM-4.7-Flash智谱AI | 1,368 | +/-6 | 11,678 | 智谱AI | MIT |
| 174 | amazon-nova-experimental-chat-11-10Amazon | 1,366 | +/-4 | 25,264 | Amazon | Proprietary |
| 175 | Gemma 3 - 27B (IT)Google Deep Mind | 1,366 | +/-4 | 47,453 | Google Deep Mind | Gemma |
| 176 | minimax-m1MiniMax | 1,364 | +/-4 | 35,117 | MiniMax | Apache 2.0 |
| 177 | OpenAI o3-mini (high)OpenAI | 1,364 | +/-5 | 18,589 | OpenAI | Proprietary |
| 178 | OpenAI o3-mini (high)OpenAI | 1,362 | +/-5 | 16,938 | OpenAI | Proprietary |
| 179 | nvidia-nemotron-3-super-120b-a12bNvidia | 1,361 | +/-7 | 7,484 | Nvidia | NVIDIA Open Model |
| 180 | Gemini 2.0 Flash ExperimentalDeepMind | 1,361 | +/-4 | 43,718 | DeepMind | Proprietary |
| 181 | DeepSeek-V3DeepSeek-AI | 1,359 | +/-5 | 21,770 | DeepSeek-AI | DeepSeek |
| 182 | Mistral-Small-3.2MistralAI | 1,358 | +/-5 | 17,685 | MistralAI | Apache 2.0 |
| 183 | 1,357 | +/-5 | 22,672 | xAI | Proprietary | |
| 184 | intellect-3 MIT | 1,356 | +/-8 | 5,323 | — | — |
| 185 | C4AI Command A (202503)CohereAI | 1,354 | +/-3 | 56,175 | CohereAI | CC-BY-NC-4.0 |
| 186 | Gemini 2.0 Flash-LiteDeepMind | 1,354 | +/-4 | 24,955 | DeepMind | Proprietary |
| 187 | GLM-4.5V智谱AI | 1,354 | +/-8 | 4,953 | 智谱AI | MIT |
| 188 | GPT OSS 120BOpenAI | 1,352 | +/-4 | 30,594 | OpenAI | Apache 2.0 |
| 189 | Gemini 1.5 ProGoogle Deep Mind | 1,351 | +/-3 | 55,606 | Google Deep Mind | Proprietary |
| 190 | amazon-nova-experimental-chat-10-20Amazon | 1,349 | +/-6 | 11,449 | Amazon | Proprietary |
| 191 | hunyuan-turbos-20250226Tencent | 1,349 | +/-12 | 2,220 | Tencent | Proprietary |
| 192 | Step3StepFunAI | 1,349 | +/-7 | 6,531 | StepFunAI | Apache 2.0 |
| 193 | OpenAI o3-miniOpenAI | 1,348 | +/-4 | 57,283 | OpenAI | Proprietary |
| 194 | llama-3.1-nemotron-ultra-253b-v1Nvidia | 1,348 | +/-12 | 2,549 | Nvidia | Nvidia Open Model |
| 195 | amazon-nova-experimental-chat-10-09Amazon | 1,348 | +/-11 | 2,823 | Amazon | Proprietary |
| 196 | Qwen3-32B阿里巴巴 | 1,347 | +/-9 | 3,926 | 阿里巴巴 | Apache 2.0 |
| 197 | mercury-2 InceptionAI | 1,347 | +/-11 | 3,099 | AI | Proprietary |
| 198 | ling-flash-2.0 AntGroup | 1,346 | +/-7 | 6,993 | Group | MIT |
| 199 | qwen-plus-0125Alibaba | 1,346 | +/-8 | 5,819 | Alibaba | Proprietary |
| 200 | 1,346 | +/-8 | 6,859 | MiniMaxAI | Apache 2.0 | |
| 201 | GPT-4oOpenAI | 1,346 | +/-3 | 112,881 | OpenAI | Proprietary |
| 202 | nvidia-llama-3.3-nemotron-super-49b-v1.5Nvidia | 1,343 | +/-10 | 3,340 | Nvidia | Nvidia Open |
| 203 | glm-4-plus-0111Zhipu | 1,343 | +/-8 | 5,760 | Zhipu | Proprietary |
| 204 | Claude 3.5 SonnetAnthropic | 1,343 | +/-3 | 82,419 | Anthropic | Proprietary |
| 205 | Gemma 3 - 12B (IT)Google Deep Mind | 1,342 | +/-10 | 3,829 | Google Deep Mind | Gemma |
| 206 | hunyuan-turbo-0110Tencent | 1,341 | +/-12 | 2,290 | Tencent | Proprietary |
| 207 | GPT-5-Nano (high)OpenAI | 1,337 | +/-7 | 8,254 | OpenAI | Proprietary |
| 208 | OpenAI o1-miniOpenAI | 1,337 | +/-4 | 51,981 | OpenAI | Proprietary |
| 209 | Nova 2 Lite亚马逊 | 1,337 | +/-6 | 12,213 | 亚马逊 | Proprietary |
| 210 | QwQ-32B阿里巴巴 | 1,336 | +/-4 | 25,361 | 阿里巴巴 | Apache 2.0 |
| 211 | 1,336 | +/-4 | 63,498 | xAI | Proprietary | |
| 212 | gemini-advanced-0514Google | 1,336 | +/-5 | 50,148 | Proprietary | |
| 213 | GPT-4oOpenAI | 1,335 | +/-4 | 45,499 | OpenAI | Proprietary |
| 214 | llama-3.1-405b-instruct-bf16Meta | 1,335 | +/-4 | 41,375 | Meta | Llama 3.1 Community |
| 215 | step-2-16k-exp-202412StepFun | 1,334 | +/-9 | 4,833 | StepFun | Proprietary |
| 216 | llama-3.1-405b-instruct-fp8Meta | 1,333 | +/-4 | 59,656 | Meta | Llama 3.1 Community |
| 217 | olmo-3.1-32b-instructAi2 | 1,330 | +/-6 | 12,195 | Ai2 | Apache 2.0 |
| 218 | molmo-2-8bAi2 | 1,329 | +/-21 | 800 | Ai2 | Apache 2.0 |
| 219 | yi-lightning Proprietary | 1,328 | +/-5 | 27,332 | — | — |
| 220 | llama-3.3-nemotron-49b-super-v1Nvidia | 1,328 | +/-12 | 2,218 | Nvidia | Nvidia |
| 221 | Qwen3-30B-A3B阿里巴巴 | 1,327 | +/-5 | 26,448 | 阿里巴巴 | Apache 2.0 |
| 222 | Llama 4 Maverick InstructFacebook AI研究实验室 | 1,327 | +/-4 | 39,928 | Facebook AI研究实验室 | Llama 4 |
| 223 | hunyuan-large-2025-02-10Tencent | 1,326 | +/-10 | 3,738 | Tencent | Proprietary |
| 224 | gpt-4-turbo-2024-04-09OpenAI | 1,324 | +/-4 | 98,114 | OpenAI | Proprietary |
| 225 | Claude 3.5 HaikuAnthropic | 1,324 | +/-3 | 69,899 | Anthropic | Proprietary |
| 226 | Gemini 1.5 ProGoogle Deep Mind | 1,324 | +/-4 | 79,138 | Google Deep Mind | Proprietary |
| 227 | deepseek-v2.5-1210DeepSeek | 1,324 | +/-8 | 6,795 | DeepSeek | DeepSeek |
| 228 | Llama 4 Scout InstructFacebook AI研究实验室 | 1,323 | +/-5 | 30,251 | Facebook AI研究实验室 | Llama |
| 229 | GPT-4.1 nanoOpenAI | 1,322 | +/-8 | 6,103 | OpenAI | Proprietary |
| 230 | Claude3-OpusAnthropic | 1,322 | +/-3 | 194,909 | Anthropic | Proprietary |
| 231 | ring-flash-2.0 AntGroup | 1,321 | +/-7 | 7,129 | Group | MIT |
| 232 | step-1o-turbo-202506StepFun | 1,320 | +/-7 | 9,023 | StepFun | Proprietary |
| 233 | glm-4-plusZhipu AI | 1,319 | +/-5 | 26,126 | Zhipu AI | Proprietary |
| 234 | Llama3.3-70B-InstructFacebook AI研究实验室 | 1,318 | +/-3 | 54,698 | Facebook AI研究实验室 | Llama-3.3 |
| 235 | Gemma-3n-E4BGoogle Deep Mind | 1,318 | +/-5 | 22,553 | Google Deep Mind | Gemma |
| 236 | qwen-max-0919Alibaba | 1,318 | +/-6 | 16,478 | Alibaba | Qwen |
| 237 | GPT-4o miniOpenAI | 1,318 | +/-4 | 68,707 | OpenAI | Proprietary |
| 238 | GPT OSS 20BOpenAI | 1,317 | +/-6 | 10,618 | OpenAI | Apache 2.0 |
| 239 | nvidia-nemotron-3-nano-30b-a3b-bf16Nvidia | 1,315 | +/-6 | 15,489 | Nvidia | NVIDIA Open Model |
| 240 | qwen2.5-plus-1127Alibaba | 1,315 | +/-6 | 10,187 | Alibaba | Proprietary |
| 241 | athene-v2-chat NexusFlow | 1,314 | +/-5 | 24,739 | — | — |
| 242 | mistral-large-2407Mistral | 1,314 | +/-4 | 45,459 | Mistral | Mistral Research |
| 243 | GPT-4OpenAI | 1,313 | +/-4 | 93,439 | OpenAI | Proprietary |
| 244 | GPT-4OpenAI | 1,312 | +/-4 | 100,105 | OpenAI | Proprietary |
| 245 | hunyuan-standard-2025-02-10Tencent | 1,311 | +/-10 | 3,904 | Tencent | Proprietary |
| 246 | gemini-1.5-flash-002Google | 1,309 | +/-4 | 34,902 | Proprietary | |
| 247 | 1,309 | +/-4 | 52,567 | xAI | Proprietary | |
| 248 | DeepSeek V2.5DeepSeek-AI | 1,307 | +/-5 | 24,572 | DeepSeek-AI | DeepSeek |
| 249 | athene-70b-0725 CC-BY-NC-4.0 | 1,307 | +/-6 | 19,621 | — | — |
| 250 | mercury InceptionAI | 1,306 | +/-14 | 1,953 | AI | Proprietary |
| 251 | granite-4.1-8bIBM | 1,306 | +/-10 | 3,995 | IBM | Apache 2.0 |
| 252 | olmo-3-32b-thinkAi2 | 1,305 | +/-8 | 5,930 | Ai2 | Apache 2.0 |
| 253 | mistral-large-2411Mistral | 1,305 | +/-4 | 28,073 | Mistral | MRL |
| 254 | Magistral-Medium-2506MistralAI | 1,304 | +/-6 | 11,617 | MistralAI | Proprietary |
| 255 | Gemma 3 - 4B (IT)Google Deep Mind | 1,303 | +/-9 | 4,171 | Google Deep Mind | Gemma |
| 256 | Mistral-Small-3.1-24B-Instruct-2503MistralAI | 1,303 | +/-5 | 33,180 | MistralAI | Apache 2.0 |
| 257 | Qwen2.5-VL-72B-Instruct阿里巴巴 | 1,303 | +/-4 | 39,406 | 阿里巴巴 | Qwen |
| 258 | Llama3.1-70B-InstructFacebook AI研究实验室 | 1,299 | +/-8 | 7,140 | Facebook AI研究实验室 | Llama 3.1 |
| 259 | hunyuan-large-visionTencent | 1,294 | +/-9 | 5,362 | Tencent | Proprietary |
| 260 | Llama3.1-70B-InstructFacebook AI研究实验室 | 1,293 | +/-4 | 55,240 | Facebook AI研究实验室 | Llama 3.1 Community |
| 261 | amazon-nova-pro-v1.0Amazon | 1,290 | +/-5 | 24,745 | Amazon | Proprietary |
| 262 | jamba-1.5-large Jamba Open | 1,289 | +/-7 | 8,662 | — | — |
| 263 | gemma-2-27b-itGoogle | 1,289 | +/-3 | 75,754 | Gemma license | |
| 264 | reka-core-20240904 Proprietary | 1,288 | +/-7 | 7,312 | — | — |
| 265 | ibm-granite-h-smallIBM | 1,287 | +/-8 | 5,677 | IBM | Apache 2.0 |
| 266 | GPT-4OpenAI | 1,287 | +/-5 | 54,173 | OpenAI | Proprietary |
| 267 | gemini-1.5-flash-001Google | 1,286 | +/-5 | 62,833 | Proprietary | |
| 268 | llama-3.1-nemotron-51b-instructNvidia | 1,286 | +/-10 | 3,749 | Nvidia | Llama 3.1 |
| 269 | llama-3.1-tulu-3-70bAi2 | 1,286 | +/-10 | 2,846 | Ai2 | Llama 3.1 |
| 270 | olmo-3.1-32b-thinkAi2 | 1,285 | +/-7 | 8,485 | Ai2 | Apache 2.0 |
| 271 | Claude3-SonnetAnthropic | 1,281 | +/-4 | 109,284 | Anthropic | Proprietary |
| 272 | gemma-2-9b-it-simpo MIT | 1,280 | +/-7 | 10,072 | — | — |
| 273 | nemotron-4-340b-instructNvidia | 1,277 | +/-5 | 19,659 | Nvidia | NVIDIA Open Model |
| 274 | Llama3-70B-InstructFacebook AI研究实验室 | 1,276 | +/-4 | 156,876 | Facebook AI研究实验室 | Llama 3 Community |
| 275 | command-r-plus-08-2024Cohere | 1,276 | +/-7 | 9,866 | Cohere | CC-BY-NC-4.0 |
| 276 | GPT-4OpenAI | 1,275 | +/-4 | 88,723 | OpenAI | Proprietary |
| 277 | Mistral Small 24B Instruct 2501MistralAI | 1,274 | +/-6 | 14,681 | MistralAI | Apache 2.0 |
| 278 | GLM4智谱AI | 1,273 | +/-7 | 9,788 | 智谱AI | Proprietary |
| 279 | reka-flash-20240904 Proprietary | 1,272 | +/-7 | 7,536 | — | — |
| 280 | Qwen2.5-Coder-32B-Instruct阿里巴巴 | 1,271 | +/-8 | 5,432 | 阿里巴巴 | Apache 2.0 |
| 281 | C4AI Aya Vision 32BCohereAI | 1,267 | +/-5 | 27,124 | CohereAI | CC-BY-NC-4.0 |
| 282 | gemma-2-9b-itGoogle | 1,267 | +/-4 | 54,611 | Gemma license | |
| 283 | deepseek-coder-v2DeepSeek | 1,265 | +/-6 | 15,147 | DeepSeek | DeepSeek License |
| 284 | Qwen2-72B-Instruct阿里巴巴 | 1,261 | +/-5 | 37,325 | 阿里巴巴 | Qianwen LICENSE |
| 285 | C4AI Command R+CohereAI | 1,261 | +/-4 | 77,554 | CohereAI | CC-BY-NC-4.0 |
| 286 | Claude3-HaikuAnthropic | 1,261 | +/-4 | 117,701 | Anthropic | Proprietary |
| 287 | amazon-nova-lite-v1.0Amazon | 1,260 | +/-5 | 19,372 | Amazon | Proprietary |
| 288 | gemini-1.5-flash-8b-001Google | 1,259 | +/-4 | 35,558 | Proprietary | |
| 289 | Phi-4-reasoningMicrosoft Azure | 1,256 | +/-5 | 24,126 | Microsoft Azure | MIT |
| 290 | olmo-2-0325-32b-instructAi2 | 1,251 | +/-11 | 3,334 | Ai2 | Apache-2.0 |
| 291 | command-r-08-2024Cohere | 1,250 | +/-7 | 10,140 | Cohere | CC-BY-NC-4.0 |
| 292 | mistral-large-2402Mistral | 1,242 | +/-5 | 62,436 | Mistral | Proprietary |
| 293 | amazon-nova-micro-v1.0Amazon | 1,241 | +/-5 | 19,364 | Amazon | Proprietary |
| 294 | jamba-1.5-mini Jamba Open | 1,240 | +/-7 | 8,858 | — | — |
| 295 | ministral-8b-2410Mistral | 1,237 | +/-9 | 4,781 | Mistral | MRL |
| 296 | gemini-pro-dev-apiGoogle | 1,236 | +/-7 | 18,354 | Proprietary | |
| 297 | Qwen1.5-110B-Chat阿里巴巴 | 1,234 | +/-6 | 26,195 | 阿里巴巴 | Qianwen LICENSE |
| 298 | hunyuan-standard-256kTencent | 1,233 | +/-12 | 2,728 | Tencent | Proprietary |
| 299 | reka-flash-21b-20240226-online Proprietary | 1,233 | +/-7 | 15,450 | — | — |
| 300 | Qwen1.5-72B-Chat阿里巴巴 | 1,233 | +/-5 | 39,302 | 阿里巴巴 | Qianwen LICENSE |
| 301 | Mixtral-8x22B-Instruct-v0.1MistralAI | 1,229 | +/-5 | 51,416 | MistralAI | Apache 2.0 |
| 302 | reka-flash-21b-20240226 Proprietary | 1,226 | +/-6 | 24,806 | — | — |
| 303 | command-rCohere | 1,226 | +/-5 | 54,036 | Cohere | CC-BY-NC-4.0 |
| 304 | gpt-3.5-turbo-0125OpenAI | 1,225 | +/-5 | 66,207 | OpenAI | Proprietary |
| 305 | Llama3-8B-InstructFacebook AI研究实验室 | 1,223 | +/-4 | 104,642 | Facebook AI研究实验室 | Llama 3 Community |
| 306 | C4AI Aya Vision 8BCohereAI | 1,223 | +/-7 | 9,818 | CohereAI | CC-BY-NC-4.0 |
| 307 | Gemini-proDeepMind | 1,223 | +/-12 | 6,390 | DeepMind | Proprietary |
| 308 | mistral-mediumMistral | 1,222 | +/-5 | 34,550 | Mistral | Proprietary |
| 309 | llama-3.1-tulu-3-8bAi2 | 1,221 | +/-11 | 2,896 | Ai2 | Llama 3.1 |
| 310 | Yi-1.5-34B零一万物 | 1,213 | +/-5 | 24,146 | 零一万物 | — |
| 311 | zephyr-orpo-141b-A35b-v0.1 Apache 2.0 | 1,212 | +/-11 | 4,652 | — | — |
| 312 | Llama3.1-8B-InstructFacebook AI研究实验室 | 1,211 | +/-4 | 49,605 | Facebook AI研究实验室 | Llama 3.1 Community |
| 313 | Llama3.1-8B-InstructFacebook AI研究实验室 | 1,208 | +/-11 | 3,090 | Facebook AI研究实验室 | Apache 2.0 |
| 314 | qwen1.5-32b-chatAlibaba | 1,203 | +/-6 | 21,741 | Alibaba | Qianwen LICENSE |
| 315 | gpt-3.5-turbo-1106OpenAI | 1,203 | +/-9 | 16,619 | OpenAI | Proprietary |
| 316 | gemma-2-2b-itGoogle | 1,200 | +/-4 | 46,616 | Gemma license | |
| 317 | Phi-3-medium 14B-previewMicrosoft Azure | 1,197 | +/-5 | 25,055 | Microsoft Azure | MIT |
| 318 | mixtral-8x7b-instruct-v0.1Mistral | 1,197 | +/-4 | 73,503 | Mistral | Apache 2.0 |
| 319 | dbrx-instruct-preview DBRX LICENSE | 1,195 | +/-6 | 32,191 | — | — |
| 320 | InternLM2-Base-20B上海人工智能实验室 | 1,191 | +/-7 | 9,901 | 上海人工智能实验室 | — |
| 321 | Qwen1.5-14B-Chat阿里巴巴 | 1,191 | +/-7 | 17,839 | 阿里巴巴 | Qianwen LICENSE |
| 322 | DeepSeek LLM 67B ChatDeepSeek-AI | 1,184 | +/-11 | 4,932 | DeepSeek-AI | DeepSeek License |
| 323 | WizardLM-70B-V1.0WizardLM Team | 1,184 | +/-9 | 8,214 | WizardLM Team | Llama 2 Community |
| 324 | Yi-34B零一万物 | 1,183 | +/-7 | 15,483 | 零一万物 | — |
| 325 | granite-3.0-8b-instructIBM | 1,182 | +/-9 | 6,638 | IBM | Apache 2.0 |
| 326 | openchat-3.5 Apache-2.0 | 1,182 | +/-10 | 7,968 | — | — |
| 327 | openchat-3.5-0106 Apache-2.0 | 1,182 | +/-8 | 12,637 | — | — |
| 328 | Gemma 1.1-7B-ITGoogle Research | 1,182 | +/-6 | 23,893 | Google Research | Gemma license |
| 329 | snowflake-arctic-instruct Apache 2.0 | 1,179 | +/-6 | 32,832 | — | — |
| 330 | granite-3.1-2b-instructIBM | 1,179 | +/-11 | 3,188 | IBM | Apache 2.0 |
| 331 | LLaMA2 70BFacebook AI研究实验室 | 1,177 | +/-10 | 6,535 | Facebook AI研究实验室 | — |
| 332 | openhermes-2.5-mistral-7b Apache-2.0 | 1,175 | +/-10 | 5,006 | — | — |
| 333 | Vicuna 33BLM-SYS | 1,172 | +/-6 | 22,479 | LM-SYS | — |
| 334 | starling-lm-7b-beta Apache-2.0 | 1,171 | +/-7 | 16,056 | — | — |
| 335 | Phi-3-small 7BMicrosoft Azure | 1,171 | +/-6 | 17,766 | Microsoft Azure | MIT |
| 336 | llama-2-70b-chatMeta | 1,170 | +/-5 | 38,492 | Meta | Llama 2 Community |
| 337 | starling-lm-7b-alpha CC-BY-NC-4.0 | 1,167 | +/-8 | 10,224 | — | — |
| 338 | llama-3.2-3b-instructMeta | 1,166 | +/-8 | 7,936 | Meta | Llama 3.2 |
| 339 | nous-hermes-2-mixtral-8x7b-dpo Apache-2.0 | 1,164 | +/-12 | 3,777 | — | — |
| 340 | Qwen3-VL-2B阿里巴巴 | 1,156 | +/-8 | 6,837 | 阿里巴巴 | Apache 2.0 |
| 341 | QwQ-32B-Preview阿里巴巴 | 1,155 | +/-11 | 3,231 | 阿里巴巴 | Apache 2.0 |
| 342 | llama2-70b-steerlm-chatNvidia | 1,154 | +/-13 | 3,585 | Nvidia | Llama 2 Community |
| 343 | solar-10.7b-instruct-v1.0 CC-BY-NC-4.0 | 1,152 | +/-13 | 4,155 | — | — |
| 344 | dolphin-2.2.1-mistral-7b Apache-2.0 | 1,152 | +/-15 | 1,679 | — | — |
| 345 | mpt-30b-chat CC-BY-NC-SA-4.0 | 1,150 | +/-12 | 2,572 | — | — |
| 346 | wizardlm-13bMicrosoft | 1,149 | +/-9 | 7,044 | Microsoft | Llama 2 Community |
| 347 | Mistral-7B-Instruct-v0.2MistralAI | 1,149 | +/-7 | 19,402 | MistralAI | Apache-2.0 |
| 348 | falcon-180b-chat Falcon-180B TII License | 1,147 | +/-17 | 1,295 | — | — |
| 349 | Qwen1.5-7B-Chat阿里巴巴 | 1,143 | +/-10 | 4,737 | 阿里巴巴 | Qianwen LICENSE |
| 350 | Phi-3-mini 3.8BMicrosoft Azure | 1,143 | +/-6 | 12,297 | Microsoft Azure | MIT |
| 351 | Baichuan2-13B-Chat百川智能 | 1,141 | +/-7 | 19,174 | 百川智能 | Llama 2 Community |
| 352 | vicuna-13b Llama 2 Community | 1,141 | +/-7 | 19,367 | — | — |
| 353 | Qwen-14B-Chat阿里巴巴 | 1,139 | +/-11 | 4,964 | 阿里巴巴 | Qianwen LICENSE |
| 354 | PaLM 2Google Research | 1,138 | +/-9 | 8,554 | Google Research | Proprietary |
| 355 | Gemma 7B - ItGoogle Research | 1,137 | +/-9 | 8,925 | Google Research | Gemma license |
| 356 | CodeLLaMA-34BFacebook AI研究实验室 | 1,136 | +/-9 | 7,366 | Facebook AI研究实验室 | Llama 2 Community |
| 357 | zephyr-7b-beta MIT | 1,130 | +/-9 | 11,118 | — | — |
| 358 | Phi-3-mini 3.8BMicrosoft Azure | 1,129 | +/-7 | 20,685 | Microsoft Azure | MIT |
| 359 | Phi-3-mini 3.8BMicrosoft Azure | 1,128 | +/-6 | 20,118 | Microsoft Azure | MIT |
| 360 | guanaco-33b Non-commercial | 1,127 | +/-12 | 2,921 | — | — |
| 361 | zephyr-7b-alpha MIT | 1,126 | +/-16 | 1,785 | — | — |
| 362 | stripedhyena-nous-7b Apache 2.0 | 1,121 | +/-11 | 5,182 | — | — |
| 363 | CodeLlama-70B-InstructFacebook AI研究实验室 | 1,119 | +/-18 | 1,143 | Facebook AI研究实验室 | Llama 2 Community |
| 364 | Gemma 1.1-2B-ITGoogle Research | 1,116 | +/-8 | 10,854 | Google Research | Gemma license |
| 365 | vicuna-7b Llama 2 Community | 1,115 | +/-9 | 6,923 | — | — |
| 366 | smollm2-1.7b-instruct Apache 2.0 | 1,114 | +/-14 | 2,199 | — | — |
| 367 | llama-3.2-1b-instructMeta | 1,111 | +/-8 | 8,045 | Meta | Llama 3.2 |
| 368 | Mistral 7B InstructMistralAI | 1,110 | +/-9 | 8,977 | MistralAI | Apache 2.0 |
| 369 | Baichuan2-7B-Chat百川智能 | 1,107 | +/-7 | 14,148 | 百川智能 | Llama 2 Community |
| 370 | Gemma 2B - ItGoogle Research | 1,093 | +/-11 | 4,780 | Google Research | Gemma license |
| 371 | Qwen1.5-4B-Chat阿里巴巴 | 1,090 | +/-9 | 7,597 | 阿里巴巴 | Qianwen LICENSE |
| 372 | olmo-7b-instructAi2 | 1,073 | +/-11 | 6,328 | Ai2 | Apache-2.0 |
| 373 | Koala达摩院 | 1,070 | +/-10 | 6,965 | 达摩院 | — |
| 374 | alpaca-13b Non-commercial | 1,069 | +/-11 | 5,745 | — | — |
| 375 | GPT4All 13BNomic AI | 1,067 | +/-15 | 1,743 | Nomic AI | — |
| 376 | mpt-7b-chat CC-BY-NC-SA-4.0 | 1,062 | +/-12 | 3,924 | — | — |
| 377 | ChatGLM3-6B智谱AI | 1,056 | +/-12 | 4,658 | 智谱AI | — |
| 378 | RWKV-4-Raven-14B Apache 2.0 | 1,041 | +/-11 | 4,845 | — | — |
| 379 | ChatGLM2-6B智谱AI | 1,024 | +/-14 | 2,658 | 智谱AI | — |
| 380 | oasst-pythia-12b Apache 2.0 | 1,023 | +/-11 | 6,310 | — | — |
| 381 | ChatGLM-6B智谱AI | 995 | +/-13 | 4,914 | 智谱AI | — |
| 382 | fastchat-t5-3b Apache 2.0 | 992 | +/-12 | 4,203 | — | — |
| 383 | dolly-v2-12b MIT | 981 | +/-13 | 3,412 | — | — |
| 384 | LLaMA 13BFacebook AI研究实验室 | 974 | +/-16 | 2,391 | Facebook AI研究实验室 | Non-commercial |
| 385 | stablelm-tuned-alpha-7b CC-BY-NC-SA-4.0 | 952 | +/-13 | 3,287 | — | — |
Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.
FAQ
What is Text Generation Arena (LMArena)?
Text Generation Arena, formerly LMSYS Chatbot Arena, is one of the most widely followed anonymous LLM evaluation platforms. Users compare answers from two hidden models and vote for the better response; Elo-style scoring aggregates those votes into a dynamic leaderboard.
How is the Arena Elo score calculated?
Arena Elo is adapted from chess rating systems. After each head-to-head comparison, the preferred model gains rating points and the other model loses points, with the size of the change depending on the rating gap. The 95% confidence interval reflects how much comparison data supports the estimate.
Why do some models have both Thinking and regular versions?
Some models offer an extended-thinking mode that spends more inference time reasoning before producing the final answer. This can improve scores on reasoning, math, and coding tasks, but usually increases latency and cost, so Arena tracks these variants separately.
How should I choose an LLM from this leaderboard?
Consider overall Elo, cost, language coverage, open-source availability, and latency. The top-ranked model is not always the best fit for every workflow.
















