AI Model Leaderboards

Name: AI Model Performance Leaderboard
Creator: DataLearner
License: https://creativecommons.org/licenses/by/4.0/

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-05-02 07:14:49

As of 2026-05, AA Intelligence Index leaders include GPT-5.5 (xhigh), GPT-5.5 (high), Opus 4.7 (max), based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Opus 4.7 (thinking), Claude Opus 4.6 (thinking), Claude Opus 4.6 near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-05-10

#ModelScore

GPT-5.5 (xhigh)OpenAI

GPT-5.5 (high)OpenAI

Opus 4.7 (max)Anthropic

Gemini 3.1 Pro PreviewGoogle Deep Mind

GPT-5.5 (medium)OpenAI

Kimi K2.6Moonshot AI

MiMo-V2.5-ProXiaomi

GPT-5.3 Codex (xhigh)OpenAI

Grok 4.3xAI

Muse SparkFacebook AI研究实验室

Source: Artificial Analysis

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-05-07

#ModelElo

Opus 4.7 (thinking)Anthropic

1503

Claude Opus 4.6 (thinking)Anthropic

1502

Claude Opus 4.6Anthropic

1498

Gemini 3.1 Pro PreviewGoogle Deep Mind

1492

Opus 4.7Anthropic

1491

Muse SparkFacebook AI研究实验室

1490

Gemini 3.0 Pro (Preview 11-2025)Google Deep Mind

1486

gpt-5.5-highOpenAI

1484

grok-4.20-beta1xAI

1480

gpt-5.2-chat-latest-20260210OpenAI

1477

Source: LMArena

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks

Overall

ARC-AGI-2 HLE MMLU Pro Open Benchmark Directory

Math

AIME 2025 FrontierMath MATH-500 Open Math Leaderboard

Coding

SWE-bench Verified LiveCodeBench SWE-Bench Pro Open Coding Leaderboard

Agent

τ²-Bench Terminal Bench 2.0 Aider-Polyglot Open Agent Leaderboard

Model Size:All 3B and below 7B 13B 34B 65B 100B and above

Model Type:All Reasoning Models Foundation Models Instruction/Chat Models Coding Models

Source:All Open Source Closed Source

Origin:All China

LLM Performance Results

Data source: DataLearnerAI

Scores shown are the best result across all evaluation modes. Click a model name for the full breakdown.

Rank	Model						License
	Qwen 3.6 Plus Preview 阿里巴巴	50.60	—	—	78.80	—	Proprietary
	Claude Sonnet 4.5 Anthropic	33.60	13.60	4.20	82.00	84.70	Proprietary
	M2.1 MiniMaxAI	22.00	—	—	74.80	—	Free commercial
4	GPT-4.5 OpenAI	—	—	—	38.00	—	Proprietary
5	Gemma 4 31B DeepMind	26.50	—	—	—	76.90	Free commercial
6	DeepSeek-V3.1 Terminus DeepSeek-AI	21.70	—	—	68.40	37.00	Free commercial
7	DeepSeek-V3.1 DeepSeek-AI	15.90	—	—	66.00	—	Free commercial
8	GLM-4.7 智谱AI	42.80	—	2.10	73.80	87.40	Free commercial
9	Qwen3 Max (Preview) 阿里巴巴	11.10	—	—	69.60	74.00	Proprietary
10	GLM-4.6 智谱AI	30.40	—	2.10	68.00	75.90	Free commercial
11	Qwen3-235B-A22B-2507 阿里巴巴	—	1.30	—	—	—	Free commercial
12	Gemma 4 26B A4B DeepMind	17.20	—	—	—	68.20	Free commercial
13	Pangu Pro MoE 华为	—	—	—	—	—	Free commercial
14	MiniMax M2 MiniMaxAI	12.50	—	—	69.40	77.20	Free commercial
15	DeepSeek-V3-0324 DeepSeek-AI	5.20	—	—	38.80	38.80	Free commercial
16	Kimi K2 Moonshot AI	4.70	—	0.01	51.80	64.30	Free commercial
17	GPT-4.1 OpenAI	3.70	—	—	54.60	54.70	Proprietary
18	GPT-4o(2025-03-27) OpenAI	—	—	—	—	—	Proprietary
19	Gemini 2.0 Pro Experimental DeepMind	—	—	—	—	—	Proprietary
20	Pangu Embedded 华为	—	—	—	—	—	Free commercial
21	Qwen3-30B-A3B-2507 阿里巴巴	9.80	—	—	22.00	49.00	Free commercial
22	ERNIE-4.5-300B-A47B 百度	—	—	—	—	—	Free commercial
23	Claude 3.5 Sonnet New Anthropic	—	—	—	49.00	—	Proprietary
24	GPT-4o(2024-11-20) OpenAI	—	—	—	31.00	—	Proprietary
25	Qwen2.5-Max 阿里巴巴	—	—	—	—	—	Proprietary
26	DeepSeek-V3 DeepSeek-AI	—	—	—	—	—	Free commercial
27	Grok 2 xAI	—	—	—	—	—	Free commercial
28	GLM-4-9B-Chat 智谱AI	—	—	—	—	—	Free commercial
29	Gemini 2.0 Flash-Lite DeepMind	—	—	—	—	—	Proprietary
30	Mistral-Small-3.2 MistralAI	—	—	—	—	—	Free commercial
31	Llama3.3-70B-Instruct Facebook AI研究实验室	—	—	—	—	—	Free commercial
32	Gemma 3 - 27B (IT) Google Deep Mind	—	—	—	—	—	Free commercial
33	Qwen3-Next 阿里巴巴	—	—	—	—	—	Free commercial
34	Mixtral-8x22B-Instruct-v0.1 MistralAI	—	—	—	—	—	Free commercial
35	Llama3-70B-Instruct Facebook AI研究实验室	—	—	—	—	—	Free commercial
36	Phi-4-mini-instruct (3.8B) Microsoft Azure	—	—	—	—	—	Free commercial
37	Llama3-70B Facebook AI研究实验室	—	—	—	—	—	Free commercial
38	Grok-1.5 xAI	—	—	—	—	—	Proprietary
39	Llama3.1-8B-Instruct Facebook AI研究实验室	—	—	—	—	—	Free commercial
40	Moonlight-16B-A3B-Instruct Moonshot AI	—	—	—	—	—	Free commercial
41	Mistral-7B-Instruct-v0.3 MistralAI	—	—	—	—	—	Free commercial
42	Claude Mythos Preview Anthropic	64.70	—	—	93.90	—	Proprietary
43	GLM-5 智谱AI	50.40	4.90	2.10	77.80	89.70	Free commercial
44	Claude Sonnet 4.6 Anthropic	49.00	58.30	8.30	79.60	—	Proprietary
45	GPT-5.2 OpenAI	45.50	54.20	18.80	80.00	82.00	Proprietary
46	Grok 4 Heavy xAI	44.40	—	2.10	73.50	—	Proprietary
47	Gemini 3.0 Flash Google Deep Mind	43.50	33.60	4.20	68.70	90.20	Proprietary
48	Gemini 2.5 Deep Think Google Deep Mind	34.80	—	10.40	—	—	Proprietary
49	Kimi K2 0905 Moonshot AI	21.70	—	—	69.20	—	Free commercial
50	Grok 4 Fast xAI	20.00	—	—	—	—	Proprietary

Qwen 3.6 Plus Preview

阿里巴巴

HLE50.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.80

τ²-Bench—

Proprietary

Claude Sonnet 4.5

Anthropic

HLE33.60

ARC-AGI-213.60

FrontierMath - Tier 44.20

SWE-bench Verified82.00

τ²-Bench84.70

Proprietary

M2.1

MiniMaxAI

HLE22.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified74.80

τ²-Bench—

Free commercial

GPT-4.5

OpenAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified38.00

τ²-Bench—

Proprietary

Gemma 4 31B

DeepMind

HLE26.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench76.90

Free commercial

DeepSeek-V3.1 Terminus

DeepSeek-AI

HLE21.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified68.40

τ²-Bench37.00

Free commercial

DeepSeek-V3.1

DeepSeek-AI

HLE15.90

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified66.00

τ²-Bench—

Free commercial

GLM-4.7

智谱AI

HLE42.80

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.80

τ²-Bench87.40

Free commercial

Qwen3 Max (Preview)

阿里巴巴

HLE11.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.60

τ²-Bench74.00

Proprietary

GLM-4.6

智谱AI

HLE30.40

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified68.00

τ²-Bench75.90

Free commercial

Qwen3-235B-A22B-2507

阿里巴巴

HLE—

ARC-AGI-21.30

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemma 4 26B A4B

DeepMind

HLE17.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench68.20

Free commercial

Pangu Pro MoE

华为

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

MiniMax M2

MiniMaxAI

HLE12.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.40

τ²-Bench77.20

Free commercial

DeepSeek-V3-0324

DeepSeek-AI

HLE5.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified38.80

τ²-Bench38.80

Free commercial

Kimi K2

Moonshot AI

HLE4.70

ARC-AGI-2—

FrontierMath - Tier 40.01

SWE-bench Verified51.80

τ²-Bench64.30

Free commercial

GPT-4.1

OpenAI

HLE3.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified54.60

τ²-Bench54.70

Proprietary

GPT-4o(2025-03-27)

OpenAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 2.0 Pro Experimental

DeepMind

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Pangu Embedded

华为

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-30B-A3B-2507

阿里巴巴

HLE9.80

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified22.00

τ²-Bench49.00

Free commercial

ERNIE-4.5-300B-A47B

百度

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Claude 3.5 Sonnet New

Anthropic

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified49.00

τ²-Bench—

Proprietary

GPT-4o(2024-11-20)

OpenAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified31.00

τ²-Bench—

Proprietary

Qwen2.5-Max

阿里巴巴

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

DeepSeek-V3

DeepSeek-AI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Grok 2

xAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

GLM-4-9B-Chat

智谱AI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemini 2.0 Flash-Lite

DeepMind

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Mistral-Small-3.2

MistralAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama3.3-70B-Instruct

Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemma 3 - 27B (IT)

Google Deep Mind

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-Next

阿里巴巴

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Mixtral-8x22B-Instruct-v0.1

MistralAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama3-70B-Instruct

Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Phi-4-mini-instruct (3.8B)

Microsoft Azure

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama3-70B

Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Grok-1.5

xAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Llama3.1-8B-Instruct

Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Moonlight-16B-A3B-Instruct

Moonshot AI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Mistral-7B-Instruct-v0.3

MistralAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Claude Mythos Preview

Anthropic

HLE64.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified93.90

τ²-Bench—

Proprietary

GLM-5

智谱AI

HLE50.40

ARC-AGI-24.90

FrontierMath - Tier 42.10

SWE-bench Verified77.80

τ²-Bench89.70

Free commercial

Claude Sonnet 4.6

Anthropic

HLE49.00

ARC-AGI-258.30

FrontierMath - Tier 48.30

SWE-bench Verified79.60

τ²-Bench—

Proprietary

GPT-5.2

OpenAI

HLE45.50

ARC-AGI-254.20

FrontierMath - Tier 418.80

SWE-bench Verified80.00

τ²-Bench82.00

Proprietary

Grok 4 Heavy

xAI

HLE44.40

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.50

τ²-Bench—

Proprietary

Gemini 3.0 Flash

Google Deep Mind

HLE43.50

ARC-AGI-233.60

FrontierMath - Tier 44.20

SWE-bench Verified68.70

τ²-Bench90.20

Proprietary

Gemini 2.5 Deep Think

Google Deep Mind

HLE34.80

ARC-AGI-2—

FrontierMath - Tier 410.40

SWE-bench Verified—

τ²-Bench—

Proprietary

Kimi K2 0905

Moonshot AI

HLE21.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.20

τ²-Bench—

Free commercial

Grok 4 Fast

xAI

HLE20.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Sort by:

Showing 50 of 60 modelsView MMLU Pro benchmark page

Leaderboard FAQ

Where does the leaderboard data come from?

Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

How do I compare an open-source model with a closed API model?

Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.

Composite Rankings

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

LLM Performance Results

Data source: DataLearnerAI

Scores shown are the best result across all evaluation modes. Click a model name for the full breakdown.

Rank	Model						License
	Qwen 3.6 Plus Preview 阿里巴巴	50.60	—	—	78.80	—	Proprietary
	Claude Sonnet 4.5 Anthropic	33.60	13.60	4.20	82.00	84.70	Proprietary
	M2.1 MiniMaxAI	22.00	—	—	74.80	—	Free commercial
4	GPT-4.5 OpenAI	—	—	—	38.00	—	Proprietary
5	Gemma 4 31B DeepMind	26.50	—	—	—	76.90	Free commercial
6	DeepSeek-V3.1 Terminus DeepSeek-AI	21.70	—	—	68.40	37.00	Free commercial
7	DeepSeek-V3.1 DeepSeek-AI	15.90	—	—	66.00	—	Free commercial
8	GLM-4.7 智谱AI	42.80	—	2.10	73.80	87.40	Free commercial
9	Qwen3 Max (Preview) 阿里巴巴	11.10	—	—	69.60	74.00	Proprietary
10	GLM-4.6 智谱AI	30.40	—	2.10	68.00	75.90	Free commercial
11	Qwen3-235B-A22B-2507 阿里巴巴	—	1.30	—	—	—	Free commercial
12	Gemma 4 26B A4B DeepMind	17.20	—	—	—	68.20	Free commercial
13	Pangu Pro MoE 华为	—	—	—	—	—	Free commercial
14	MiniMax M2 MiniMaxAI	12.50	—	—	69.40	77.20	Free commercial
15	DeepSeek-V3-0324 DeepSeek-AI	5.20	—	—	38.80	38.80	Free commercial
16	Kimi K2 Moonshot AI	4.70	—	0.01	51.80	64.30	Free commercial
17	GPT-4.1 OpenAI	3.70	—	—	54.60	54.70	Proprietary
18	GPT-4o(2025-03-27) OpenAI	—	—	—	—	—	Proprietary
19	Gemini 2.0 Pro Experimental DeepMind	—	—	—	—	—	Proprietary
20	Pangu Embedded 华为	—	—	—	—	—	Free commercial
21	Qwen3-30B-A3B-2507 阿里巴巴	9.80	—	—	22.00	49.00	Free commercial
22	ERNIE-4.5-300B-A47B 百度	—	—	—	—	—	Free commercial
23	Claude 3.5 Sonnet New Anthropic	—	—	—	49.00	—	Proprietary
24	GPT-4o(2024-11-20) OpenAI	—	—	—	31.00	—	Proprietary
25	Qwen2.5-Max 阿里巴巴	—	—	—	—	—	Proprietary
26	DeepSeek-V3 DeepSeek-AI	—	—	—	—	—	Free commercial
27	Grok 2 xAI	—	—	—	—	—	Free commercial
28	GLM-4-9B-Chat 智谱AI	—	—	—	—	—	Free commercial
29	Gemini 2.0 Flash-Lite DeepMind	—	—	—	—	—	Proprietary
30	Mistral-Small-3.2 MistralAI	—	—	—	—	—	Free commercial
31	Llama3.3-70B-Instruct Facebook AI研究实验室	—	—	—	—	—	Free commercial
32	Gemma 3 - 27B (IT) Google Deep Mind	—	—	—	—	—	Free commercial
33	Qwen3-Next 阿里巴巴	—	—	—	—	—	Free commercial
34	Mixtral-8x22B-Instruct-v0.1 MistralAI	—	—	—	—	—	Free commercial
35	Llama3-70B-Instruct Facebook AI研究实验室	—	—	—	—	—	Free commercial
36	Phi-4-mini-instruct (3.8B) Microsoft Azure	—	—	—	—	—	Free commercial
37	Llama3-70B Facebook AI研究实验室	—	—	—	—	—	Free commercial
38	Grok-1.5 xAI	—	—	—	—	—	Proprietary
39	Llama3.1-8B-Instruct Facebook AI研究实验室	—	—	—	—	—	Free commercial
40	Moonlight-16B-A3B-Instruct Moonshot AI	—	—	—	—	—	Free commercial
41	Mistral-7B-Instruct-v0.3 MistralAI	—	—	—	—	—	Free commercial
42	Claude Mythos Preview Anthropic	64.70	—	—	93.90	—	Proprietary
43	GLM-5 智谱AI	50.40	4.90	2.10	77.80	89.70	Free commercial
44	Claude Sonnet 4.6 Anthropic	49.00	58.30	8.30	79.60	—	Proprietary
45	GPT-5.2 OpenAI	45.50	54.20	18.80	80.00	82.00	Proprietary
46	Grok 4 Heavy xAI	44.40	—	2.10	73.50	—	Proprietary
47	Gemini 3.0 Flash Google Deep Mind	43.50	33.60	4.20	68.70	90.20	Proprietary
48	Gemini 2.5 Deep Think Google Deep Mind	34.80	—	10.40	—	—	Proprietary
49	Kimi K2 0905 Moonshot AI	21.70	—	—	69.20	—	Free commercial
50	Grok 4 Fast xAI	20.00	—	—	—	—	Proprietary

Qwen 3.6 Plus Preview

阿里巴巴

HLE50.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.80

τ²-Bench—

Proprietary

Claude Sonnet 4.5

Anthropic

HLE33.60

ARC-AGI-213.60

FrontierMath - Tier 44.20

SWE-bench Verified82.00

τ²-Bench84.70

Proprietary

M2.1

MiniMaxAI

HLE22.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified74.80

τ²-Bench—

Free commercial

GPT-4.5

OpenAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified38.00

τ²-Bench—

Proprietary

Gemma 4 31B

DeepMind

HLE26.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench76.90

Free commercial

DeepSeek-V3.1 Terminus

DeepSeek-AI

HLE21.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified68.40

τ²-Bench37.00

Free commercial

DeepSeek-V3.1

DeepSeek-AI

HLE15.90

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified66.00

τ²-Bench—

Free commercial

GLM-4.7

智谱AI

HLE42.80

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.80

τ²-Bench87.40

Free commercial

Qwen3 Max (Preview)

阿里巴巴

HLE11.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.60

τ²-Bench74.00

Proprietary

GLM-4.6

智谱AI

HLE30.40

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified68.00

τ²-Bench75.90

Free commercial

Qwen3-235B-A22B-2507

阿里巴巴

HLE—

ARC-AGI-21.30

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemma 4 26B A4B

DeepMind

HLE17.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench68.20

Free commercial

Pangu Pro MoE

华为

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

MiniMax M2

MiniMaxAI

HLE12.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.40

τ²-Bench77.20

Free commercial

DeepSeek-V3-0324

DeepSeek-AI

HLE5.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified38.80

τ²-Bench38.80

Free commercial

Kimi K2

Moonshot AI

HLE4.70

ARC-AGI-2—

FrontierMath - Tier 40.01

SWE-bench Verified51.80

τ²-Bench64.30

Free commercial

GPT-4.1

OpenAI

HLE3.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified54.60

τ²-Bench54.70

Proprietary

GPT-4o(2025-03-27)

OpenAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 2.0 Pro Experimental

DeepMind

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Pangu Embedded

华为

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-30B-A3B-2507

阿里巴巴

HLE9.80

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified22.00

τ²-Bench49.00

Free commercial

ERNIE-4.5-300B-A47B

百度

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Claude 3.5 Sonnet New

Anthropic

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified49.00

τ²-Bench—

Proprietary

GPT-4o(2024-11-20)

OpenAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified31.00

τ²-Bench—

Proprietary

Qwen2.5-Max

阿里巴巴

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

DeepSeek-V3

DeepSeek-AI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Grok 2

xAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

GLM-4-9B-Chat

智谱AI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemini 2.0 Flash-Lite

DeepMind

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Mistral-Small-3.2

MistralAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama3.3-70B-Instruct

Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemma 3 - 27B (IT)

Google Deep Mind

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-Next

阿里巴巴

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Mixtral-8x22B-Instruct-v0.1

MistralAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama3-70B-Instruct

Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Phi-4-mini-instruct (3.8B)

Microsoft Azure

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama3-70B

Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Grok-1.5

xAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Llama3.1-8B-Instruct

Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Moonlight-16B-A3B-Instruct

Moonshot AI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Mistral-7B-Instruct-v0.3

MistralAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Claude Mythos Preview

Anthropic

HLE64.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified93.90

τ²-Bench—

Proprietary

GLM-5

智谱AI

HLE50.40

ARC-AGI-24.90

FrontierMath - Tier 42.10

SWE-bench Verified77.80

τ²-Bench89.70

Free commercial

Claude Sonnet 4.6

Anthropic

HLE49.00

ARC-AGI-258.30

FrontierMath - Tier 48.30

SWE-bench Verified79.60

τ²-Bench—

Proprietary

GPT-5.2

OpenAI

HLE45.50

ARC-AGI-254.20

FrontierMath - Tier 418.80

SWE-bench Verified80.00

τ²-Bench82.00

Proprietary

Grok 4 Heavy

xAI

HLE44.40

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.50

τ²-Bench—

Proprietary

Gemini 3.0 Flash

Google Deep Mind

HLE43.50

ARC-AGI-233.60

FrontierMath - Tier 44.20

SWE-bench Verified68.70

τ²-Bench90.20

Proprietary

Gemini 2.5 Deep Think

Google Deep Mind

HLE34.80

ARC-AGI-2—

FrontierMath - Tier 410.40

SWE-bench Verified—

τ²-Bench—

Proprietary

Kimi K2 0905

Moonshot AI

HLE21.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.20

τ²-Bench—

Free commercial

Grok 4 Fast

xAI

HLE20.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Sort by:

Showing 50 of 60 modelsView MMLU Pro benchmark page

Leaderboard FAQ

Where does the leaderboard data come from?

Why do scores for the same model differ across benchmarks?

How often is the leaderboard updated?

How should I read the composite ranking?

How do I compare an open-source model with a closed API model?