AI Model Leaderboards

Name: AI Model Performance Leaderboard
Creator: DataLearner
License: https://creativecommons.org/licenses/by/4.0/

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-07-28 08:43:41

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-08-02

#ModelScore

Claude Opus 5 (max)Anthropic

Claude Opus 5 (xhigh)Anthropic

Claude Fable 5Anthropic

GPT-5.6 Sol (max)OpenAI

Claude Opus 5 (high)Anthropic

GPT-5.6 Sol (xhigh)OpenAI

Kimi K3 (max)Kimi

Claude Opus 5 (medium)Anthropic

GPT-5.6 Sol (high)OpenAI

GPT-5.6 Terra (max)OpenAI

Source: Artificial Analysis

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-08-01

#ModelElo

Claude Fable 5Anthropic

1509

Claude Opus 4.6 (thinking)Anthropic

1505

Opus 4.7 (thinking)Anthropic

1502

Claude Opus 4.6Anthropic

1497

Opus 4.7Anthropic

1492

claude-opus-5-highAnthropic

1492

claude-opus-5-maxAnthropic

1490

Muse Spark 1.1Facebook AI研究实验室

1490

Muse SparkFacebook AI研究实验室

1488

Gemini 3 ProGoogle Deep Mind

1486

Source: LMArena

Recent Rank Changes

Risers, decliners, and new entrants across the coding, math, and agent leaderboards over the last 30 days.

Coding

Full ranking

Agent

Full ranking

View the full AI model changelog

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks

Overall

ARC-AGI-2 HLE MMLU Pro Open Benchmark Directory

Math

AIME 2025 FrontierMath MATH-500 Open Math Leaderboard

Coding

SWE-bench Verified LiveCodeBench SWE-Bench Pro Open Coding Leaderboard

Agent

τ²-Bench Terminal Bench 2.0 Aider-Polyglot Open Agent Leaderboard

Model Size:All 3B and below 7B 13B 34B 65B 100B and above

Model Type:All Reasoning Models Foundation Models Instruction/Chat Models Coding Models

License:All Open Source Closed Source

Region:All China

Recommended models

Ranked by LiveCodeBench

Current SOTA

Gemini 2.5 Deep Think

Google Deep Mind

87.60LiveCodeBench

View model

Best Open-Source

Step 3.5 Flash

StepFunAI

86.40LiveCodeBench−1.20

View model

Best China-Made

Qwen 3.6 Plus Preview

阿里巴巴

87.10LiveCodeBench−0.50

View model

LLM Performance Results

Data source: DataLearnerAI

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.

Rank	Model							License
	Gemini 2.5 Deep Think Google Deep Mind	87.60	34.80	—	10.40	—	—	Proprietary	Details
	Qwen 3.6 Plus Preview 阿里巴巴	87.10	50.60	—	—	78.80	—	Proprietary	Details
	Qwen3.6-Max-Preview 阿里巴巴	87.10	50.20	—	—	78.80	—	Proprietary	Details
4	Step 3.5 Flash StepFunAI	86.40	—	—	—	74.40	88.20	Free commercial	Details
5	GLM-4.7 智谱AI	84.90	42.80	—	2.10	73.80	87.40	Free commercial	Details
6	GLM-4.6 智谱AI	84.50	30.40	—	2.10	68.00	75.90	Free commercial	Details
7	MiniMax M2 MiniMaxAI	83.00	12.50	—	—	69.40	77.20	Free commercial	Details
8	Gemma 4 31B DeepMind	80.00	26.50	—	—	—	76.90	Free commercial	Details
9	DeepSeek-V3.1 Terminus DeepSeek-AI	80.00	21.70	—	—	68.40	37.00	Free commercial	Details
10	Grok 4 Fast xAI	80.00	20.00	—	—	—	—	Proprietary	Details
11	Gemma 4 26B A4B DeepMind	77.10	17.20	—	—	—	68.20	Free commercial	Details
12	DeepSeek-V3.1 DeepSeek-AI	74.80	15.90	—	—	66.00	—	Free commercial	Details
13	Claude Sonnet 4.5 Anthropic	71.00	33.60	13.60	4.20	82.00	84.70	Proprietary	Details
14	Grok 3 xAI	70.60	—	—	—	—	—	Proprietary	Details
15	Pangu Embedded 华为	67.10	—	—	—	—	—	Free commercial	Details
16	Pangu Pro MoE 华为	59.60	—	—	—	—	—	Free commercial	Details
17	Qwen3 Max (Preview) 阿里巴巴	57.50	11.10	—	—	69.60	74.00	Proprietary	Details
18	Hunyuan-7B Tencent ARC	57.00	—	—	—	—	—	Free commercial	Details
19	Qwen3-Next 阿里巴巴	56.60	—	—	—	—	—	Free commercial	Details
20	Qwen3-4B-Thinking-2507 阿里巴巴	55.20	—	—	—	—	—	Free commercial	Details
21	Kimi K2 Moonshot AI	53.70	4.70	—	0.01	51.80	64.30	Free commercial	Details
22	Qwen3-235B-A22B-2507 阿里巴巴	51.80	—	1.30	—	—	—	Free commercial	Details
23	GLM-4-9B-Chat 智谱AI	51.80	—	—	—	—	—	Free commercial	Details
24	DeepSeek-V3-0324 DeepSeek-AI	49.20	5.20	—	—	38.80	38.80	Free commercial	Details
25	GPT-4.5 OpenAI	46.40	—	—	—	38.00	—	Proprietary	Details
26	Qwen3-30B-A3B-2507 阿里巴巴	43.20	9.80	—	—	22.00	49.00	Free commercial	Details
27	GPT-4.1 OpenAI	40.50	3.70	—	—	54.60	54.70	Proprietary	Details
28	ERNIE-4.5-300B-A47B 百度	38.80	—	—	—	—	—	Free commercial	Details
29	Claude 3.5 Sonnet New Anthropic	38.70	—	—	—	49.00	—	Proprietary	Details
30	GPT-4o(2025-03-27) OpenAI	35.80	—	—	—	—	—	Proprietary	Details
31	Qwen3-4B-2507 阿里巴巴	35.10	—	—	—	—	—	Free commercial	Details
32	DeepSeek-V3 DeepSeek-AI	34.60	—	—	—	—	—	Free commercial	Details
33	Llama3.3-70B-Instruct Facebook AI研究实验室	33.30	—	—	—	—	—	Free commercial	Details
34	Gemma 3 - 27B (IT) Google Deep Mind	29.70	—	—	—	—	—	Free commercial	Details
35	Gemini 2.0 Flash-Lite DeepMind	28.90	—	—	—	—	—	Proprietary	Details
36	Claude Mythos Preview Anthropic	—	64.70	—	—	93.90	—	Proprietary	Details
37	GLM-5 智谱AI	—	50.40	4.90	2.10	77.80	89.70	Free commercial	Details
38	Claude Sonnet 4.6 Anthropic	—	49.00	58.30	8.30	79.60	—	Proprietary	Details
39	GPT-5.2 OpenAI	—	45.50	54.20	18.80	80.00	82.00	Proprietary	Details
40	Grok 4 Heavy xAI	—	44.40	—	2.10	73.50	—	Proprietary	Details
41	Gemini 3.0 Flash Google Deep Mind	—	43.50	33.60	4.20	68.70	90.20	Proprietary	Details
42	M2.1 MiniMaxAI	—	22.00	—	—	74.80	—	Free commercial	Details
43	Kimi K2 0905 Moonshot AI	—	21.70	—	—	69.20	—	Free commercial	Details
44	Claude Sonnet 3.7 Anthropic	—	10.30	—	—	70.30	61.80	Proprietary	Details
45	Mistral-7B-Instruct-v0.3 MistralAI	—	—	—	—	—	—	Free commercial	Details
46	GPT-4.1 nano OpenAI	—	—	—	—	—	—	Proprietary	Details
47	Moonlight-16B-A3B-Instruct Moonshot AI	—	—	—	—	—	—	Free commercial	Details
48	Gemini 2.5 Flash-Preview-09-2025 Google Deep Mind	—	—	—	—	54.00	—	Proprietary	Details
49	GPT-4.1 mini OpenAI	—	—	—	—	23.60	53.00	Proprietary	Details
50	GPT-4o(2025-01-29) OpenAI	—	—	—	—	—	—	Proprietary	Details

Gemini 2.5 Deep Think Google Deep Mind

LiveCodeBench87.60

HLE34.80

ARC-AGI-2—

FrontierMath - Tier 410.40

SWE-bench Verified—

τ²-Bench—

Proprietary

Qwen 3.6 Plus Preview 阿里巴巴

LiveCodeBench87.10

HLE50.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.80

τ²-Bench—

Proprietary

Qwen3.6-Max-Preview 阿里巴巴

LiveCodeBench87.10

HLE50.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.80

τ²-Bench—

Proprietary

Step 3.5 Flash StepFunAI

LiveCodeBench86.40

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified74.40

τ²-Bench88.20

Free commercial

GLM-4.7 智谱AI

LiveCodeBench84.90

HLE42.80

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.80

τ²-Bench87.40

Free commercial

GLM-4.6 智谱AI

LiveCodeBench84.50

HLE30.40

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified68.00

τ²-Bench75.90

Free commercial

MiniMax M2 MiniMaxAI

LiveCodeBench83.00

HLE12.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.40

τ²-Bench77.20

Free commercial

Gemma 4 31B DeepMind

LiveCodeBench80.00

HLE26.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench76.90

Free commercial

DeepSeek-V3.1 Terminus DeepSeek-AI

LiveCodeBench80.00

HLE21.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified68.40

τ²-Bench37.00

Free commercial

Grok 4 Fast xAI

LiveCodeBench80.00

HLE20.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemma 4 26B A4B DeepMind

LiveCodeBench77.10

HLE17.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench68.20

Free commercial

DeepSeek-V3.1 DeepSeek-AI

LiveCodeBench74.80

HLE15.90

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified66.00

τ²-Bench—

Free commercial

Claude Sonnet 4.5 Anthropic

LiveCodeBench71.00

HLE33.60

ARC-AGI-213.60

FrontierMath - Tier 44.20

SWE-bench Verified82.00

τ²-Bench84.70

Proprietary

Grok 3 xAI

LiveCodeBench70.60

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Pangu Embedded 华为

LiveCodeBench67.10

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Pangu Pro MoE 华为

LiveCodeBench59.60

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3 Max (Preview)阿里巴巴

LiveCodeBench57.50

HLE11.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.60

τ²-Bench74.00

Proprietary

Hunyuan-7B Tencent ARC

LiveCodeBench57.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-Next 阿里巴巴

LiveCodeBench56.60

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-4B-Thinking-2507 阿里巴巴

LiveCodeBench55.20

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Kimi K2 Moonshot AI

LiveCodeBench53.70

HLE4.70

ARC-AGI-2—

FrontierMath - Tier 40.01

SWE-bench Verified51.80

τ²-Bench64.30

Free commercial

Qwen3-235B-A22B-2507 阿里巴巴

LiveCodeBench51.80

HLE—

ARC-AGI-21.30

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

GLM-4-9B-Chat 智谱AI

LiveCodeBench51.80

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

DeepSeek-V3-0324 DeepSeek-AI

LiveCodeBench49.20

HLE5.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified38.80

τ²-Bench38.80

Free commercial

GPT-4.5 OpenAI

LiveCodeBench46.40

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified38.00

τ²-Bench—

Proprietary

Qwen3-30B-A3B-2507 阿里巴巴

LiveCodeBench43.20

HLE9.80

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified22.00

τ²-Bench49.00

Free commercial

GPT-4.1 OpenAI

LiveCodeBench40.50

HLE3.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified54.60

τ²-Bench54.70

Proprietary

ERNIE-4.5-300B-A47B 百度

LiveCodeBench38.80

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Claude 3.5 Sonnet New Anthropic

LiveCodeBench38.70

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified49.00

τ²-Bench—

Proprietary

GPT-4o(2025-03-27)OpenAI

LiveCodeBench35.80

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Qwen3-4B-2507 阿里巴巴

LiveCodeBench35.10

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

DeepSeek-V3 DeepSeek-AI

LiveCodeBench34.60

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama3.3-70B-Instruct Facebook AI研究实验室

LiveCodeBench33.30

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemma 3 - 27B (IT)Google Deep Mind

LiveCodeBench29.70

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemini 2.0 Flash-Lite DeepMind

LiveCodeBench28.90

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude Mythos Preview Anthropic

LiveCodeBench—

HLE64.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified93.90

τ²-Bench—

Proprietary

GLM-5 智谱AI

LiveCodeBench—

HLE50.40

ARC-AGI-24.90

FrontierMath - Tier 42.10

SWE-bench Verified77.80

τ²-Bench89.70

Free commercial

Claude Sonnet 4.6 Anthropic

LiveCodeBench—

HLE49.00

ARC-AGI-258.30

FrontierMath - Tier 48.30

SWE-bench Verified79.60

τ²-Bench—

Proprietary

GPT-5.2 OpenAI

LiveCodeBench—

HLE45.50

ARC-AGI-254.20

FrontierMath - Tier 418.80

SWE-bench Verified80.00

τ²-Bench82.00

Proprietary

Grok 4 Heavy xAI

LiveCodeBench—

HLE44.40

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.50

τ²-Bench—

Proprietary

Gemini 3.0 Flash Google Deep Mind

LiveCodeBench—

HLE43.50

ARC-AGI-233.60

FrontierMath - Tier 44.20

SWE-bench Verified68.70

τ²-Bench90.20

Proprietary

M2.1 MiniMaxAI

LiveCodeBench—

HLE22.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified74.80

τ²-Bench—

Free commercial

Kimi K2 0905 Moonshot AI

LiveCodeBench—

HLE21.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.20

τ²-Bench—

Free commercial

Claude Sonnet 3.7 Anthropic

LiveCodeBench—

HLE10.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified70.30

τ²-Bench61.80

Proprietary

Mistral-7B-Instruct-v0.3 MistralAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

GPT-4.1 nano OpenAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Moonlight-16B-A3B-Instruct Moonshot AI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemini 2.5 Flash-Preview-09-2025 Google Deep Mind

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified54.00

τ²-Bench—

Proprietary

GPT-4.1 mini OpenAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified23.60

τ²-Bench53.00

Proprietary

GPT-4o(2025-01-29)OpenAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Sort by:

Showing 50 of 63 modelsView LiveCodeBench benchmark page

Leaderboard FAQ

Where does the leaderboard data come from?

Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

How do I compare an open-source model with a closed API model?

Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.

As of 2026-07, AA Intelligence Index leaders include Claude Opus 5 (max), Claude Opus 5 (xhigh), Claude Fable 5, based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Claude Fable 5, Claude Opus 4.6 (thinking), Opus 4.7 (thinking) near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.