AI Model Leaderboards

Name: AI Model Performance Leaderboard
Creator: DataLearner
License: https://creativecommons.org/licenses/by/4.0/

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-07-28 08:43:41

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-08-02

#ModelScore

Claude Opus 5 (max)Anthropic

Claude Opus 5 (xhigh)Anthropic

Claude Fable 5Anthropic

GPT-5.6 Sol (max)OpenAI

Claude Opus 5 (high)Anthropic

GPT-5.6 Sol (xhigh)OpenAI

Kimi K3 (max)Kimi

Claude Opus 5 (medium)Anthropic

GPT-5.6 Sol (high)OpenAI

GPT-5.6 Terra (max)OpenAI

Source: Artificial Analysis

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-08-01

#ModelElo

Claude Fable 5Anthropic

1509

Claude Opus 4.6 (thinking)Anthropic

1505

Opus 4.7 (thinking)Anthropic

1502

Claude Opus 4.6Anthropic

1497

Opus 4.7Anthropic

1492

claude-opus-5-highAnthropic

1492

claude-opus-5-maxAnthropic

1490

Muse Spark 1.1Facebook AI研究实验室

1490

Muse SparkFacebook AI研究实验室

1488

Gemini 3 ProGoogle Deep Mind

1486

Source: LMArena

Recent Rank Changes

Risers, decliners, and new entrants across the coding, math, and agent leaderboards over the last 30 days.

Coding

Full ranking

Agent

Full ranking

View the full AI model changelog

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks

Overall

ARC-AGI-2 HLE MMLU Pro Open Benchmark Directory

Math

AIME 2025 FrontierMath MATH-500 Open Math Leaderboard

Coding

SWE-bench Verified LiveCodeBench SWE-Bench Pro Open Coding Leaderboard

Agent

τ²-Bench Terminal Bench 2.0 Aider-Polyglot Open Agent Leaderboard

Model Size:All 3B and below 7B 13B 34B 65B 100B and above

Model Type:All Reasoning Models Foundation Models Instruction/Chat Models Coding Models

License:All Open Source Closed Source

Region:All China

Recommended models

Ranked by MMLU Pro

Current SOTA

OpenAI o1

OpenAI

91.04MMLU Pro

View model

Best Open-Source

No qualifying model on this benchmark.

Best China-Made

Qwen3.7-Max-Preview

阿里巴巴

89.60MMLU Pro−1.44

View model

LLM Performance Results

Data source: DataLearnerAI

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.

Rank	Model							License
	OpenAI o1 OpenAI	91.04	9.10	—	—	48.90	—	Proprietary	Details
	Gemini 3.0 Pro (Preview 11-2025) Google Deep Mind	90.00	45.80	45.10	18.80	76.20	85.40	Proprietary	Details
	Opus 4.5 Anthropic	90.00	43.20	37.60	4.20	80.90	81.99	Proprietary	Details
4	Qwen3.7-Max-Preview 阿里巴巴	89.60	53.50	—	—	80.40	—	Proprietary	Details
5	Qwen 3.6 Plus Preview 阿里巴巴	88.50	50.60	—	—	78.80	—	Proprietary	Details
6	Qwen3.6-Max-Preview 阿里巴巴	88.50	50.20	—	—	78.80	—	Proprietary	Details
7	Claude Sonnet 4.5 Anthropic	88.00	33.60	13.60	4.20	82.00	84.70	Proprietary	Details
8	Opus 4.1 Anthropic	88.00	—	—	4.20	74.50	—	Proprietary	Details
9	Hunyuan-T1 腾讯AI实验室	87.20	—	—	—	—	—	Proprietary	Details
10	Grok 4 xAI	87.00	38.60	15.90	2.10	58.60	—	Proprietary	Details
11	Doubao Seed 2.0 Pro 字节跳动Seed团队	87.00	—	—	—	76.50	—	Proprietary	Details
12	GPT-4.5 OpenAI	86.10	—	—	—	38.00	—	Proprietary	Details
13	Gemini 2.5-Pro Google Deep Mind	86.00	21.60	4.90	2.10	67.20	—	Proprietary	Details
14	Qwen3-Max-Thinking 阿里巴巴	85.70	49.80	—	—	75.30	82.10	Proprietary	Details
15	OpenAI o3 OpenAI	85.60	20.32	6.50	2.10	69.10	—	Proprietary	Details
16	Grok 4.1 Fast xAI	85.00	17.60	—	—	—	82.71	Proprietary	Details
17	Claude Opus 4 Anthropic	85.00	10.70	8.60	4.20	72.50	72.50	Proprietary	Details
18	Claude Sonnet 4 Anthropic	84.00	9.60	5.90	—	80.20	52.00	Proprietary	Details
19	Qwen3 Max (Preview) 阿里巴巴	84.00	11.10	—	—	69.60	74.00	Proprietary	Details
20	ERNIE 5.0 百度	83.80	25.81	—	—	—	78.79	Proprietary	Details
21	OpenAI o4 - mini OpenAI	80.60	17.70	—	6.30	68.10	56.90	Proprietary	Details
22	GPT-4.1 OpenAI	80.50	3.70	—	—	54.60	54.70	Proprietary	Details
23	OpenAI o1-mini OpenAI	80.30	—	—	—	—	—	Proprietary	Details
24	Haiku 4.5 Anthropic	80.00	9.70	4.50	2.10	73.30	33.00	Proprietary	Details
25	GPT-4o(2025-03-27) OpenAI	79.80	—	—	—	—	—	Proprietary	Details
26	Gemini 2.0 Pro Experimental DeepMind	79.10	—	—	—	—	—	Proprietary	Details
27	Hunyuan-TurboS 腾讯AI实验室	79.00	—	—	—	—	—	Proprietary	Details
28	GPT-5-mini OpenAI	78.00	5.00	—	6.30	—	—	Proprietary	Details
29	Claude 3.5 Sonnet New Anthropic	78.00	—	—	—	49.00	—	Proprietary	Details
30	GPT-4o OpenAI	77.90	5.30	—	—	31.00	—	Proprietary	Details
31	GPT-4o(2024-11-20) OpenAI	77.90	—	—	—	31.00	—	Proprietary	Details
32	Claude 3.5 Sonnet Anthropic	77.64	—	—	—	—	—	Proprietary	Details
33	Gemini 2.0 Flash Experimental DeepMind	76.24	5.10	—	—	21.40	—	Proprietary	Details
34	Qwen2.5-Max 阿里巴巴	76.10	—	—	—	—	—	Proprietary	Details
35	Gemini 1.5 Pro Google Deep Mind	76.10	—	—	—	—	—	Proprietary	Details
36	Gemini 2.0 Flash-Lite DeepMind	71.60	—	—	—	—	—	Proprietary	Details
37	Claude3-Opus Anthropic	68.45	—	—	—	—	—	Proprietary	Details
38	Claude 3.5 Haiku Anthropic	65.00	—	—	—	—	—	Proprietary	Details
39	GPT-4o mini OpenAI	61.70	—	—	—	—	—	Proprietary	Details
40	Claude3-Sonnet Anthropic	56.80	—	—	—	—	—	Proprietary	Details
41	Grok-1.5 xAI	51.00	—	—	—	—	—	Proprietary	Details
42	Gemini 2.5 Flash Google Deep Mind	—	11.00	—	4.20	50.00	—	Proprietary	Details
43	Claude Mythos Preview Anthropic	—	64.70	—	—	93.90	—	Proprietary	Details
44	Claude Opus 5 Anthropic	—	64.70	90.40	—	96.00	—	Proprietary	Details
45	Muse Spark 1.1 Facebook AI研究实验室	—	62.10	—	—	—	—	Proprietary	Details
46	Gemini 2.5 Flash-Lite Google Deep Mind	—	6.90	—	—	27.60	—	Proprietary	Details
47	GPT-5 OpenAI	—	35.20	9.90	12.50	72.80	80.00	Proprietary	Details
48	Claude Fable 5 Anthropic	—	59.00	—	—	95.00	—	Proprietary	Details
49	GPT-5.4 Pro OpenAI	—	58.70	83.30	38.00	—	—	Proprietary	Details
50	Muse Spark Facebook AI研究实验室	—	58.00	42.50	14.60	77.40	—	Proprietary	Details

OpenAI o1 OpenAI

MMLU Pro91.04

HLE9.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified48.90

τ²-Bench—

Proprietary

Gemini 3.0 Pro (Preview 11-2025)Google Deep Mind

MMLU Pro90.00

HLE45.80

ARC-AGI-245.10

FrontierMath - Tier 418.80

SWE-bench Verified76.20

τ²-Bench85.40

Proprietary

Opus 4.5 Anthropic

MMLU Pro90.00

HLE43.20

ARC-AGI-237.60

FrontierMath - Tier 44.20

SWE-bench Verified80.90

τ²-Bench81.99

Proprietary

Qwen3.7-Max-Preview 阿里巴巴

MMLU Pro89.60

HLE53.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.40

τ²-Bench—

Proprietary

Qwen 3.6 Plus Preview 阿里巴巴

MMLU Pro88.50

HLE50.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.80

τ²-Bench—

Proprietary

Qwen3.6-Max-Preview 阿里巴巴

MMLU Pro88.50

HLE50.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.80

τ²-Bench—

Proprietary

Claude Sonnet 4.5 Anthropic

MMLU Pro88.00

HLE33.60

ARC-AGI-213.60

FrontierMath - Tier 44.20

SWE-bench Verified82.00

τ²-Bench84.70

Proprietary

Opus 4.1 Anthropic

MMLU Pro88.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 44.20

SWE-bench Verified74.50

τ²-Bench—

Proprietary

Hunyuan-T1 腾讯AI实验室

MMLU Pro87.20

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Grok 4 xAI

MMLU Pro87.00

HLE38.60

ARC-AGI-215.90

FrontierMath - Tier 42.10

SWE-bench Verified58.60

τ²-Bench—

Proprietary

Doubao Seed 2.0 Pro 字节跳动Seed团队

MMLU Pro87.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified76.50

τ²-Bench—

Proprietary

GPT-4.5 OpenAI

MMLU Pro86.10

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified38.00

τ²-Bench—

Proprietary

Gemini 2.5-Pro Google Deep Mind

MMLU Pro86.00

HLE21.60

ARC-AGI-24.90

FrontierMath - Tier 42.10

SWE-bench Verified67.20

τ²-Bench—

Proprietary

Qwen3-Max-Thinking 阿里巴巴

MMLU Pro85.70

HLE49.80

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified75.30

τ²-Bench82.10

Proprietary

OpenAI o3 OpenAI

MMLU Pro85.60

HLE20.32

ARC-AGI-26.50

FrontierMath - Tier 42.10

SWE-bench Verified69.10

τ²-Bench—

Proprietary

Grok 4.1 Fast xAI

MMLU Pro85.00

HLE17.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench82.71

Proprietary

Claude Opus 4 Anthropic

MMLU Pro85.00

HLE10.70

ARC-AGI-28.60

FrontierMath - Tier 44.20

SWE-bench Verified72.50

τ²-Bench72.50

Proprietary

Claude Sonnet 4 Anthropic

MMLU Pro84.00

HLE9.60

ARC-AGI-25.90

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench52.00

Proprietary

Qwen3 Max (Preview)阿里巴巴

MMLU Pro84.00

HLE11.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.60

τ²-Bench74.00

Proprietary

ERNIE 5.0 百度

MMLU Pro83.80

HLE25.81

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench78.79

Proprietary

OpenAI o4 - mini OpenAI

MMLU Pro80.60

HLE17.70

ARC-AGI-2—

FrontierMath - Tier 46.30

SWE-bench Verified68.10

τ²-Bench56.90

Proprietary

GPT-4.1 OpenAI

MMLU Pro80.50

HLE3.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified54.60

τ²-Bench54.70

Proprietary

OpenAI o1-mini OpenAI

MMLU Pro80.30

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Haiku 4.5 Anthropic

MMLU Pro80.00

HLE9.70

ARC-AGI-24.50

FrontierMath - Tier 42.10

SWE-bench Verified73.30

τ²-Bench33.00

Proprietary

GPT-4o(2025-03-27)OpenAI

MMLU Pro79.80

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 2.0 Pro Experimental DeepMind

MMLU Pro79.10

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Hunyuan-TurboS 腾讯AI实验室

MMLU Pro79.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

GPT-5-mini OpenAI

MMLU Pro78.00

HLE5.00

ARC-AGI-2—

FrontierMath - Tier 46.30

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude 3.5 Sonnet New Anthropic

MMLU Pro78.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified49.00

τ²-Bench—

Proprietary

GPT-4o OpenAI

MMLU Pro77.90

HLE5.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified31.00

τ²-Bench—

Proprietary

GPT-4o(2024-11-20)OpenAI

MMLU Pro77.90

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified31.00

τ²-Bench—

Proprietary

Claude 3.5 Sonnet Anthropic

MMLU Pro77.64

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 2.0 Flash Experimental DeepMind

MMLU Pro76.24

HLE5.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified21.40

τ²-Bench—

Proprietary

Qwen2.5-Max 阿里巴巴

MMLU Pro76.10

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 1.5 Pro Google Deep Mind

MMLU Pro76.10

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 2.0 Flash-Lite DeepMind

MMLU Pro71.60

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude3-Opus Anthropic

MMLU Pro68.45

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude 3.5 Haiku Anthropic

MMLU Pro65.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

GPT-4o mini OpenAI

MMLU Pro61.70

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude3-Sonnet Anthropic

MMLU Pro56.80

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Grok-1.5 xAI

MMLU Pro51.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 2.5 Flash Google Deep Mind

MMLU Pro—

HLE11.00

ARC-AGI-2—

FrontierMath - Tier 44.20

SWE-bench Verified50.00

τ²-Bench—

Proprietary

Claude Mythos Preview Anthropic

MMLU Pro—

HLE64.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified93.90

τ²-Bench—

Proprietary

Claude Opus 5 Anthropic

MMLU Pro—

HLE64.70

ARC-AGI-290.40

FrontierMath - Tier 4—

SWE-bench Verified96.00

τ²-Bench—

Proprietary

Muse Spark 1.1 Facebook AI研究实验室

MMLU Pro—

HLE62.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 2.5 Flash-Lite Google Deep Mind

MMLU Pro—

HLE6.90

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified27.60

τ²-Bench—

Proprietary

GPT-5 OpenAI

MMLU Pro—

HLE35.20

ARC-AGI-29.90

FrontierMath - Tier 412.50

SWE-bench Verified72.80

τ²-Bench80.00

Proprietary

Claude Fable 5 Anthropic

MMLU Pro—

HLE59.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified95.00

τ²-Bench—

Proprietary

GPT-5.4 Pro OpenAI

MMLU Pro—

HLE58.70

ARC-AGI-283.30

FrontierMath - Tier 438.00

SWE-bench Verified—

τ²-Bench—

Proprietary

Muse Spark Facebook AI研究实验室

MMLU Pro—

HLE58.00

ARC-AGI-242.50

FrontierMath - Tier 414.60

SWE-bench Verified77.40

τ²-Bench—

Proprietary

Sort by:

Showing 50 of 117 modelsView MMLU Pro benchmark page

Leaderboard FAQ

Where does the leaderboard data come from?

Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

How do I compare an open-source model with a closed API model?

Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.

As of 2026-07, AA Intelligence Index leaders include Claude Opus 5 (max), Claude Opus 5 (xhigh), Claude Fable 5, based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Claude Fable 5, Claude Opus 4.6 (thinking), Opus 4.7 (thinking) near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.