DataLearner logo

Text Generation Arena Leaderboard

The latest AI text generation leaderboard based on LMArena anonymous user voting. Covers Elo scores, confidence intervals, and vote counts for leading language models.

Top Model

Kimi K3 (max)

Top Score

1,485

Model Count

385

Data version

2026年08月06日

Data source: LM Arena

About This Leaderboard

This leaderboard ranks the strongest AI models for text generation. Data comes from LMArena (formerly LMSYS Chatbot Arena), the world's largest crowdsourced AI evaluation platform. Users chat with two anonymous models side-by-side and vote for the better response — rankings are determined entirely by real user preferences, not lab benchmarks.

Methodology Overview

Blind testing: Users chat with two anonymous models and vote based on response quality, eliminating brand bias.

Elo scoring: Using the Bradley-Terry model (adapted from chess Elo ratings) to calculate each model's strength score from battle outcomes. Higher scores mean users more frequently prefer that model.

Broad scenario coverage: Testing spans coding, creative writing, math reasoning, Q&A, role-playing, and more.

DataLearner provides in-depth analysis on top of the raw data, linking leaderboard models to the DataLearner model database so you can quickly access model details, API pricing, benchmark scores, and more.

Origin:AllChina
Leaderboard snapshot month:

Ranking Table

RankModelScore95% CIVotesOrganizationLicense
14Moonshot AIKimi K3 (max)Moonshot AI1,485+/-79,861Moonshot AIKimi K3 license
44Moonshot AIKimi K2.6Moonshot AI1,461+/-537,236Moonshot AIModified MIT
50DeepSeek-AIDeepSeek-V4-ProDeepSeek-AI1,457+/-451,266DeepSeek-AIMIT
54DeepSeekdeepseek-v4-pro-high-previewDeepSeek1,456+/-448,756DeepSeekMIT
59Moonshot AIKimi K2 ThinkingMoonshot AI1,451+/-368,179Moonshot AIModified MIT
70MiniMaxAIMiniMax M3MiniMaxAI1,445+/-534,343MiniMaxAIMiniMax Community License
79DeepSeekdeepseek-v4-flash-high-previewDeepSeek1,438+/-448,316DeepSeekMIT
81DeepSeek-AIDeepSeek-V4-FlashDeepSeek-AI1,436+/-448,561DeepSeek-AIMIT
89Moonshot AIKimi K2.5 InstantMoonshot AI1,431+/-78,138Moonshot AIModified MIT
93Moonshot AIKimi K2 Thinking (thinking-turbo)Moonshot AI1,430+/-361,677Moonshot AIModified MIT
98DeepSeek-AIDeepSeek V3.2DeepSeek-AI1,425+/-446,999DeepSeek-AIMIT
100DeepSeek-AIDeepSeek V3.2-Exp (thinking)DeepSeek-AI1,425+/-79,051DeepSeek-AIMIT
102Alibabaqwen3-max-2025-09-23Alibaba1,424+/-69,138AlibabaProprietary
104DeepSeek-AIDeepSeek V3.2-Exp (thinking)DeepSeek-AI1,423+/-440,832DeepSeek-AIMIT
105DeepSeek-AIDeepSeek V3.2-ExpDeepSeek-AI1,423+/-611,902DeepSeek-AIMIT
106DeepSeek-AIDeepSeek-R1-0528DeepSeek-AI1,422+/-618,422DeepSeek-AIMIT
109Moonshot AIKimi K2 0905Moonshot AI1,418+/-611,756Moonshot AIModified MIT
110Moonshot AIKimi K2Moonshot AI1,418+/-527,589Moonshot AIModified MIT
111DeepSeek-AIDeepSeek-V3.1DeepSeek-AI1,418+/-614,942DeepSeek-AIMIT
112DeepSeek-AIDeepSeek-V3.1 Terminus (thinking)DeepSeek-AI1,417+/-103,454DeepSeek-AIMIT
114DeepSeek-AIDeepSeek-V3.1 (thinking)DeepSeek-AI1,417+/-711,715DeepSeek-AIMIT
115MiniMaxAIMiniMax-M2.7MiniMaxAI1,416+/-455,513MiniMaxAIModified MIT
119DeepSeek-AIDeepSeek-V3.1 TerminusDeepSeek-AI1,415+/-103,688DeepSeek-AIMIT
123Tencenthunyuan-hy3-previewTencent1,412+/-86,555Tencenttencent-hunyuan-community
132Alibabaqwen3-235b-a22b-no-thinkingAlibaba1,403+/-538,137AlibabaApache 2.0
138Alibabaqwen3-235b-a22b-thinking-2507Alibaba1,399+/-78,975AlibabaApache 2.0
139DeepSeek-AIDeepSeek-R1DeepSeek-AI1,398+/-518,524DeepSeek-AIMIT
140StepFunAIStep 3.5 FlashStepFunAI1,397+/-457,847StepFunAIProprietary
141DeepSeek-AIDeepSeek-V3-0324DeepSeek-AI1,396+/-445,445DeepSeek-AIMIT
144Tencenthunyuan-vision-1.5-thinkingTencent1,395+/-122,214TencentProprietary
145StepFunAIStep 3.5 FlashStepFunAI1,394+/-456,979StepFunAIApache 2.0
150MiniMaxAIMiniMax M2.5MiniMaxAI1,390+/-440,688MiniMaxAIModified MIT
158MiniMaxAIM2.1MiniMaxAI1,384+/-517,057MiniMaxAIMIT
161Tencenthunyuan-turbos-20250416Tencent1,382+/-610,715TencentProprietary
176MiniMaxminimax-m1MiniMax1,364+/-435,117MiniMaxApache 2.0
181DeepSeek-AIDeepSeek-V3DeepSeek-AI1,359+/-521,770DeepSeek-AIDeepSeek
191Tencenthunyuan-turbos-20250226Tencent1,349+/-122,220TencentProprietary
192StepFunAIStep3StepFunAI1,349+/-76,531StepFunAIApache 2.0
199Alibabaqwen-plus-0125Alibaba1,346+/-85,819AlibabaProprietary
200MiniMaxAIMiniMax M2MiniMaxAI1,346+/-86,859MiniMaxAIApache 2.0
203glm-4-plus-0111Zhipu1,343+/-85,760ZhipuProprietary
206Tencenthunyuan-turbo-0110Tencent1,341+/-122,290TencentProprietary
215StepFunstep-2-16k-exp-202412StepFun1,334+/-94,833StepFunProprietary
223Tencenthunyuan-large-2025-02-10Tencent1,326+/-103,738TencentProprietary
227DeepSeekdeepseek-v2.5-1210DeepSeek1,324+/-86,795DeepSeekDeepSeek
232StepFunstep-1o-turbo-202506StepFun1,320+/-79,023StepFunProprietary
233glm-4-plusZhipu AI1,319+/-526,126Zhipu AIProprietary
236Alibabaqwen-max-0919Alibaba1,318+/-616,478AlibabaQwen
240Alibabaqwen2.5-plus-1127Alibaba1,315+/-610,187AlibabaProprietary
245Tencenthunyuan-standard-2025-02-10Tencent1,311+/-103,904TencentProprietary
248DeepSeek-AIDeepSeek V2.5DeepSeek-AI1,307+/-524,572DeepSeek-AIDeepSeek
259Tencenthunyuan-large-visionTencent1,294+/-95,362TencentProprietary
283DeepSeekdeepseek-coder-v2DeepSeek1,265+/-615,147DeepSeekDeepSeek License
298Tencenthunyuan-standard-256kTencent1,233+/-122,728TencentProprietary
314Alibabaqwen1.5-32b-chatAlibaba1,203+/-621,741AlibabaQianwen LICENSE
322DeepSeek-AIDeepSeek LLM 67B ChatDeepSeek-AI1,184+/-114,932DeepSeek-AIDeepSeek License

Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.

FAQ

01

What is Text Generation Arena (LMArena)?

Text Generation Arena, formerly LMSYS Chatbot Arena, is one of the most widely followed anonymous LLM evaluation platforms. Users compare answers from two hidden models and vote for the better response; Elo-style scoring aggregates those votes into a dynamic leaderboard.

02

How is the Arena Elo score calculated?

Arena Elo is adapted from chess rating systems. After each head-to-head comparison, the preferred model gains rating points and the other model loses points, with the size of the change depending on the rating gap. The 95% confidence interval reflects how much comparison data supports the estimate.

03

Why do some models have both Thinking and regular versions?

Some models offer an extended-thinking mode that spends more inference time reasoning before producing the final answer. This can improve scores on reasoning, math, and coding tasks, but usually increases latency and cost, so Arena tracks these variants separately.

04

How should I choose an LLM from this leaderboard?

Consider overall Elo, cost, language coverage, open-source availability, and latency. The top-ranked model is not always the best fit for every workflow.