DataLearner logoDataLearnerAI
Latest AI Insights
Model Leaderboards
Benchmarks
Model Directory
Model Comparison
Resource Center
Tools
LanguageEnglish
DataLearner logoDataLearner AI

A knowledge platform focused on LLM benchmarking, datasets, and practical instruction with continuously updated capability maps.

Products

  • Leaderboards
  • Model comparison
  • Datasets

Resources

  • Tutorials
  • Editorial
  • Tool directory

Company

  • About
  • Privacy policy
  • Data methodology
  • Contact

© 2026 DataLearner AI. DataLearner curates industry data and case studies so researchers, enterprises, and developers can rely on trustworthy intelligence.

Privacy policyTerms of service

AI Model Leaderboards

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-05-02 07:14:49

As of 2026-05, AA Intelligence Index leaders include GPT-5.5 (xhigh), GPT-5.5 (high), Opus 4.7 (max), based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Opus 4.7 (thinking), Claude Opus 4.6 (thinking), Claude Opus 4.6 near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-05-10

#ModelScore
1
OpenAI
GPT-5.5 (xhigh)OpenAI
60
2
OpenAI
GPT-5.5 (high)OpenAI
59
3
Anthropic
Opus 4.7 (max)Anthropic
57
4
Google Deep Mind
Gemini 3.1 Pro PreviewGoogle Deep Mind
57
5
OpenAI
GPT-5.5 (medium)OpenAI
57
6
Moonshot AI
Kimi K2.6Moonshot AI
54
7
X
MiMo-V2.5-ProXiaomi
54
8
OpenAI
GPT-5.3 Codex (xhigh)OpenAI
54
9
xAI
Grok 4.3xAI
53
10
F
Muse SparkFacebook AI研究实验室
52
Source: Artificial Analysis

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-05-07

#ModelElo
1
Anthropic
Opus 4.7 (thinking)Anthropic
1503
2
Anthropic
Claude Opus 4.6 (thinking)Anthropic
1502
3
Anthropic
Claude Opus 4.6Anthropic
1498
4
Google Deep Mind
Gemini 3.1 Pro PreviewGoogle Deep Mind
1492
5
Anthropic
Opus 4.7Anthropic
1491
6
F
Muse SparkFacebook AI研究实验室
1490
7
Google Deep Mind
Gemini 3.0 Pro (Preview 11-2025)Google Deep Mind
1486
8
OpenAI
gpt-5.5-highOpenAI
1484
9
xAI
grok-4.20-beta1xAI
1480
10
OpenAI
gpt-5.2-chat-latest-20260210OpenAI
1477
Source: LMArena

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks
Overall
ARC-AGI-2HLEMMLU ProOpen Benchmark Directory
Math
AIME 2025FrontierMathMATH-500Open Math Leaderboard
Coding
SWE-bench VerifiedLiveCodeBenchSWE-Bench ProOpen Coding Leaderboard
Agent
τ²-BenchTerminal Bench 2.0Aider-PolyglotOpen Agent Leaderboard
Model Size:All3B and below7B13B34B65B100B and above
Model Type:AllReasoning ModelsFoundation ModelsInstruction/Chat ModelsCoding Models
Source:AllOpen SourceClosed Source
Origin:AllChina

LLM Performance Results

Data source: DataLearnerAI
Scores shown are the best result across all evaluation modes. Click a model name for the full breakdown.
RankModelLicense
Moonshot AI
Kimi K2.6
Moonshot AI
54.00——80.20—Free commercial
智谱AI
GLM 5.1
智谱AI
52.30————Free commercial
Moonshot AI
Kimi K2 Thinking
Moonshot AI
51.00——71.30—Free commercial
4
阿里巴巴
Qwen3-Max-Thinking
阿里巴巴
49.80——75.3082.10Proprietary
5
阿里巴巴
Qwen3.5-27B
阿里巴巴
48.50——72.4079.00Free commercial
6
DeepSeek-AI
DeepSeek-V4-Pro
DeepSeek-AI
48.20——80.60—Free commercial
7
DeepSeek-AI
DeepSeek-V4-Flash
DeepSeek-AI
45.10——79.00—Free commercial
8
DeepSeek-AI
DeepSeek V3.2 Speciale
DeepSeek-AI
30.60————Free commercial
9
MiniMaxAI
MiniMax-M2.7
MiniMaxAI
28.00————Non-commercial
10
DeepSeek-AI
DeepSeek V3.2
DeepSeek-AI
25.104.002.1073.1080.30Free commercial
11
阿里巴巴
Qwen3.6-27B
阿里巴巴
24.00——77.20—Free commercial
12
阿里巴巴
Qwen3.6-35B-A3B
阿里巴巴
21.40——73.40—Free commercial
13
DeepSeek-AI
DeepSeek V3.2-Exp
DeepSeek-AI
20.30——67.8066.70Free commercial
14
MiniMaxAI
MiniMax M2.5
MiniMaxAI
19.404.90—80.20—Free commercial
15
阿里巴巴
Qwen3-235B-A22B-Thinking
阿里巴巴
18.20————Free commercial
16
阿里巴巴
Qwen3-235B-A22B-Thinking-2507
阿里巴巴
18.20————Free commercial
17
DeepSeek-AI
DeepSeek-R1-0528
DeepSeek-AI
17.701.30—57.60—Free commercial
18
智谱AI
GLM-4.5
智谱AI
14.40——64.20—Free commercial
19
智谱AI
GLM-4.7-Flash
智谱AI
14.40——59.2079.50Free commercial
20
智谱AI
GLM-4.5-Air
智谱AI
10.60——57.60—Free commercial
21
MiniMaxAI
MiniMax-M1-80k
MiniMaxAI
8.40——56.00—Free commercial
22
阿里巴巴
Qwen3-235B-A22B
阿里巴巴
7.60——34.4034.40Free commercial
23
MiniMaxAI
MiniMax-M1-40k
MiniMaxAI
7.20——55.60—Free commercial
24
DeepSeek-AI
DeepSeek-R1-Distill-Llama-70B
DeepSeek-AI
—————Free commercial
25
DeepSeek-AI
DeepSeek-R1-Distill-Qwen-7B
DeepSeek-AI
—————Free commercial
26
Moonshot AI
Kimi-k1.6-IOI-high
Moonshot AI
—————Proprietary
27
Moonshot AI
Kimi-k1.6-IOI
Moonshot AI
—————Proprietary
28
阿里巴巴
QwQ-Max-Preview
阿里巴巴
—————Free commercial
29
Moonshot AI
Kimi k1.5 (Short-CoT)
Moonshot AI
—————Proprietary
30
阿里巴巴
Qwen3-32B
阿里巴巴
—————Free commercial
31
DeepSeek-AI
DeepSeek-R1
DeepSeek-AI
———49.20—Free commercial
32
腾讯AI实验室
Hunyuan-TurboS
腾讯AI实验室
—————Proprietary
33
阿里巴巴
QwQ-32B
阿里巴巴
—————Free commercial
34
腾讯AI实验室
Hunyuan-T1
腾讯AI实验室
—————Proprietary
35
阿里巴巴
Qwen3-8B
阿里巴巴
—————Free commercial
36
阿里巴巴
QwQ-32B-Preview
阿里巴巴
—————Free commercial
37
阿里巴巴
Qwen3-30B-A3B
阿里巴巴
—————Free commercial
Kimi K2.6
Moonshot AI
HLE54.00
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified80.20
τ²-Bench—
Free commercial
GLM 5.1
智谱AI
HLE52.30
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
Kimi K2 Thinking
Moonshot AI
HLE51.00
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified71.30
τ²-Bench—
Free commercial
4
Qwen3-Max-Thinking
阿里巴巴
HLE49.80
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified75.30
τ²-Bench82.10
Proprietary
5
Qwen3.5-27B
阿里巴巴
HLE48.50
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified72.40
τ²-Bench79.00
Free commercial
6
DeepSeek-V4-Pro
DeepSeek-AI
HLE48.20
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified80.60
τ²-Bench—
Free commercial
7
DeepSeek-V4-Flash
DeepSeek-AI
HLE45.10
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified79.00
τ²-Bench—
Free commercial
8
DeepSeek V3.2 Speciale
DeepSeek-AI
HLE30.60
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
9
MiniMax-M2.7
MiniMaxAI
HLE28.00
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Non-commercial
10
DeepSeek V3.2
DeepSeek-AI
HLE25.10
ARC-AGI-24.00
FrontierMath - Tier 42.10
SWE-bench Verified73.10
τ²-Bench80.30
Free commercial
11
Qwen3.6-27B
阿里巴巴
HLE24.00
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified77.20
τ²-Bench—
Free commercial
12
Qwen3.6-35B-A3B
阿里巴巴
HLE21.40
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified73.40
τ²-Bench—
Free commercial
13
DeepSeek V3.2-Exp
DeepSeek-AI
HLE20.30
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified67.80
τ²-Bench66.70
Free commercial
14
MiniMax M2.5
MiniMaxAI
HLE19.40
ARC-AGI-24.90
FrontierMath - Tier 4—
SWE-bench Verified80.20
τ²-Bench—
Free commercial
15
Qwen3-235B-A22B-Thinking
阿里巴巴
HLE18.20
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
16
Qwen3-235B-A22B-Thinking-2507
阿里巴巴
HLE18.20
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
17
DeepSeek-R1-0528
DeepSeek-AI
HLE17.70
ARC-AGI-21.30
FrontierMath - Tier 4—
SWE-bench Verified57.60
τ²-Bench—
Free commercial
18
GLM-4.5
智谱AI
HLE14.40
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified64.20
τ²-Bench—
Free commercial
19
GLM-4.7-Flash
智谱AI
HLE14.40
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified59.20
τ²-Bench79.50
Free commercial
20
GLM-4.5-Air
智谱AI
HLE10.60
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified57.60
τ²-Bench—
Free commercial
21
MiniMax-M1-80k
MiniMaxAI
HLE8.40
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified56.00
τ²-Bench—
Free commercial
22
Qwen3-235B-A22B
阿里巴巴
HLE7.60
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified34.40
τ²-Bench34.40
Free commercial
23
MiniMax-M1-40k
MiniMaxAI
HLE7.20
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified55.60
τ²-Bench—
Free commercial
24
DeepSeek-R1-Distill-Llama-70B
DeepSeek-AI
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
25
DeepSeek-R1-Distill-Qwen-7B
DeepSeek-AI
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
26
Kimi-k1.6-IOI-high
Moonshot AI
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Proprietary
27
Kimi-k1.6-IOI
Moonshot AI
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Proprietary
28
QwQ-Max-Preview
阿里巴巴
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
29
Kimi k1.5 (Short-CoT)
Moonshot AI
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Proprietary
30
Qwen3-32B
阿里巴巴
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
31
DeepSeek-R1
DeepSeek-AI
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified49.20
τ²-Bench—
Free commercial
32
Hunyuan-TurboS
腾讯AI实验室
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Proprietary
33
QwQ-32B
阿里巴巴
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
34
Hunyuan-T1
腾讯AI实验室
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Proprietary
35
Qwen3-8B
阿里巴巴
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
36
QwQ-32B-Preview
阿里巴巴
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
37
Qwen3-30B-A3B
阿里巴巴
HLE—
ARC-AGI-2—
FrontierMath - Tier 4—
SWE-bench Verified—
τ²-Bench—
Free commercial
Sort by:

Leaderboard FAQ

01

Where does the leaderboard data come from?

Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.

02

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

03

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

04

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

05

How do I compare an open-source model with a closed API model?

Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.