AI Model Leaderboards

Name: AI Model Performance Leaderboard
Creator: DataLearner
License: https://creativecommons.org/licenses/by/4.0/

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-07-28 08:43:41

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-08-02

#ModelScore

Claude Opus 5 (max)Anthropic

Claude Opus 5 (xhigh)Anthropic

Claude Fable 5Anthropic

GPT-5.6 Sol (max)OpenAI

Claude Opus 5 (high)Anthropic

GPT-5.6 Sol (xhigh)OpenAI

Kimi K3 (max)Kimi

Claude Opus 5 (medium)Anthropic

GPT-5.6 Sol (high)OpenAI

GPT-5.6 Terra (max)OpenAI

Source: Artificial Analysis

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-08-01

#ModelElo

Claude Fable 5Anthropic

1509

Claude Opus 4.6 (thinking)Anthropic

1505

Opus 4.7 (thinking)Anthropic

1502

Claude Opus 4.6Anthropic

1497

Opus 4.7Anthropic

1492

claude-opus-5-highAnthropic

1492

claude-opus-5-maxAnthropic

1490

Muse Spark 1.1Facebook AI研究实验室

1490

Muse SparkFacebook AI研究实验室

1488

Gemini 3 ProGoogle Deep Mind

1486

Source: LMArena

Recent Rank Changes

Risers, decliners, and new entrants across the coding, math, and agent leaderboards over the last 30 days.

Coding

Full ranking

Agent

Full ranking

View the full AI model changelog

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks

Overall

ARC-AGI-2 HLE MMLU Pro Open Benchmark Directory

Math

AIME 2025 FrontierMath MATH-500 Open Math Leaderboard

Coding

SWE-bench Verified LiveCodeBench SWE-Bench Pro Open Coding Leaderboard

Agent

τ²-Bench Terminal Bench 2.0 Aider-Polyglot Open Agent Leaderboard

Model Size:All 3B and below 7B 13B 34B 65B 100B and above

Model Type:All Reasoning Models Foundation Models Instruction/Chat Models Coding Models

License:All Open Source Closed Source

Region:All China

Recommended models

Ranked by LiveCodeBench

Current SOTA

DeepSeek-V4-Pro

DeepSeek-AI

93.50LiveCodeBench

View model

Best Open-Source

DeepSeek-V4-Pro

DeepSeek-AI

93.50LiveCodeBench

View model

Best China-Made

DeepSeek-V4-Pro

DeepSeek-AI

93.50LiveCodeBench

View model

LLM Performance Results

Data source: DataLearnerAI

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.

Rank	Model							License
	DeepSeek-V4-Pro DeepSeek-AI	93.50	48.20	—	—	80.60	—	Free commercial	Details
	Qwen3.7-Max-Preview 阿里巴巴	91.60	53.50	—	—	80.40	—	Proprietary	Details
	DeepSeek-V4-Flash DeepSeek-AI	91.60	45.10	—	—	79.00	—	Free commercial	Details
4	Kimi K2.6 Moonshot AI	89.60	54.00	—	—	80.20	—	Free commercial	Details
5	Opus 4.5 Anthropic	87.00	43.20	37.60	4.20	80.90	81.99	Proprietary	Details
6	Qwen3-Max-Thinking 阿里巴巴	85.90	49.80	—	—	75.30	82.10	Proprietary	Details
7	Qwen3.6-27B 阿里巴巴	83.90	24.00	—	—	77.20	—	Free commercial	Details
8	DeepSeek V3.2 DeepSeek-AI	83.30	25.10	4.00	2.10	73.10	80.30	Free commercial	Details
9	Kimi K2 Thinking Moonshot AI	83.10	51.00	—	—	71.30	—	Free commercial	Details
10	Grok 4 xAI	82.00	38.60	15.90	2.10	58.60	—	Proprietary	Details
11	Grok 4.1 Fast xAI	82.00	17.60	—	—	—	82.71	Proprietary	Details
12	Qwen3.5-27B 阿里巴巴	80.70	48.50	—	—	72.40	79.00	Free commercial	Details
13	Qwen3.6-35B-A3B 阿里巴巴	80.40	21.40	—	—	73.40	—	Free commercial	Details
14	Gemini 2.5 Pro Deep Think Google Deep Mind	80.40	—	—	10.40	—	—	Proprietary	Details
15	Grok-3 - Reasoning Beta xAI	79.40	—	—	—	—	—	Proprietary	Details
16	Gemini-2.5-Pro-Preview-05-06 Google Deep Mind	77.10	21.60	—	2.10	63.20	—	Proprietary	Details
17	Gemini 2.5-Pro Google Deep Mind	77.10	21.60	4.90	2.10	67.20	—	Proprietary	Details
18	Claude Opus 4.6 Anthropic	76.00	53.00	66.30	22.90	80.84	91.89	Proprietary	Details
19	OpenAI o3 OpenAI	75.80	20.32	6.50	2.10	69.10	—	Proprietary	Details
20	DeepSeek V3.2-Exp DeepSeek-AI	74.10	20.30	—	—	67.80	66.70	Free commercial	Details
21	Qwen3-235B-A22B-Thinking 阿里巴巴	74.10	18.20	—	—	—	—	Free commercial	Details
22	Qwen3-235B-A22B-Thinking-2507 阿里巴巴	74.10	18.20	—	—	—	—	Free commercial	Details
23	Kimi-k1.6-IOI-high Moonshot AI	73.80	—	—	—	—	—	Proprietary	Details
24	DeepSeek-R1-0528 DeepSeek-AI	73.30	17.70	1.30	—	57.60	—	Free commercial	Details
25	GLM-4.5 智谱AI	72.90	14.40	—	—	64.20	—	Free commercial	Details
26	OpenAI o1 OpenAI	71.00	9.10	—	—	48.90	—	Proprietary	Details
27	Qwen3-235B-A22B 阿里巴巴	70.70	7.60	—	—	34.40	34.40	Free commercial	Details
28	GLM-4.5-Air 智谱AI	70.70	10.60	—	—	57.60	—	Free commercial	Details
29	Gemini 2.5 Pro Experimental 03-25 Google Deep Mind	70.40	18.80	—	4.20	63.80	—	Proprietary	Details
30	OpenAI o3-mini (high) OpenAI	69.50	—	—	4.20	49.30	—	Proprietary	Details
31	OpenAI o3-mini (medium) OpenAI	67.40	—	—	—	—	—	Proprietary	Details
32	Claude Sonnet 4 Anthropic	66.00	9.60	5.90	—	80.20	52.00	Proprietary	Details
33	Kimi-k1.6-IOI Moonshot AI	65.90	—	—	—	—	—	Proprietary	Details
34	DeepSeek-R1 DeepSeek-AI	65.90	—	—	—	49.20	—	Free commercial	Details
35	Qwen3-32B 阿里巴巴	65.70	—	—	—	—	—	Free commercial	Details
36	QwQ-Max-Preview 阿里巴巴	65.60	—	—	—	—	—	Free commercial	Details
37	MiniMax-M1-80k MiniMaxAI	65.00	8.40	—	—	56.00	—	Free commercial	Details
38	Hunyuan-T1 腾讯AI实验室	64.90	—	—	—	—	—	Proprietary	Details
39	MiniMax-M1-40k MiniMaxAI	62.30	7.20	—	—	55.60	—	Free commercial	Details
40	Qwen3-8B 阿里巴巴	61.80	—	—	—	—	—	Free commercial	Details
41	Magistral-Medium-2506 MistralAI	59.36	—	—	—	—	—	Proprietary	Details
42	Claude Opus 4 Anthropic	56.60	10.70	8.60	4.20	72.50	72.50	Proprietary	Details
43	Magistral-Small-2506 MistralAI	55.84	—	—	—	—	—	Free commercial	Details
44	Gemini 2.5 Flash Google Deep Mind	55.40	11.00	—	4.20	50.00	—	Proprietary	Details
45	OpenAI o1-mini OpenAI	52.00	—	—	—	—	—	Proprietary	Details
46	Gemini 2.5 Flash-Lite Google Deep Mind	34.30	6.90	—	—	27.60	—	Proprietary	Details
47	Hunyuan-TurboS 腾讯AI实验室	32.00	—	—	—	—	—	Proprietary	Details
48	Qwen3-30B-A3B 阿里巴巴	29.00	—	—	—	—	—	Free commercial	Details
49	Claude Opus 5 Anthropic	—	64.70	90.40	—	96.00	—	Proprietary	Details
50	Muse Spark 1.1 Facebook AI研究实验室	—	62.10	—	—	—	—	Proprietary	Details

DeepSeek-V4-Pro DeepSeek-AI

LiveCodeBench93.50

HLE48.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.60

τ²-Bench—

Free commercial

Qwen3.7-Max-Preview 阿里巴巴

LiveCodeBench91.60

HLE53.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.40

τ²-Bench—

Proprietary

DeepSeek-V4-Flash DeepSeek-AI

LiveCodeBench91.60

HLE45.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified79.00

τ²-Bench—

Free commercial

Kimi K2.6 Moonshot AI

LiveCodeBench89.60

HLE54.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench—

Free commercial

Opus 4.5 Anthropic

LiveCodeBench87.00

HLE43.20

ARC-AGI-237.60

FrontierMath - Tier 44.20

SWE-bench Verified80.90

τ²-Bench81.99

Proprietary

Qwen3-Max-Thinking 阿里巴巴

LiveCodeBench85.90

HLE49.80

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified75.30

τ²-Bench82.10

Proprietary

Qwen3.6-27B 阿里巴巴

LiveCodeBench83.90

HLE24.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified77.20

τ²-Bench—

Free commercial

DeepSeek V3.2 DeepSeek-AI

LiveCodeBench83.30

HLE25.10

ARC-AGI-24.00

FrontierMath - Tier 42.10

SWE-bench Verified73.10

τ²-Bench80.30

Free commercial

Kimi K2 Thinking Moonshot AI

LiveCodeBench83.10

HLE51.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified71.30

τ²-Bench—

Free commercial

Grok 4 xAI

LiveCodeBench82.00

HLE38.60

ARC-AGI-215.90

FrontierMath - Tier 42.10

SWE-bench Verified58.60

τ²-Bench—

Proprietary

Grok 4.1 Fast xAI

LiveCodeBench82.00

HLE17.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench82.71

Proprietary

Qwen3.5-27B 阿里巴巴

LiveCodeBench80.70

HLE48.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified72.40

τ²-Bench79.00

Free commercial

Qwen3.6-35B-A3B 阿里巴巴

LiveCodeBench80.40

HLE21.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified73.40

τ²-Bench—

Free commercial

Gemini 2.5 Pro Deep Think Google Deep Mind

LiveCodeBench80.40

HLE—

ARC-AGI-2—

FrontierMath - Tier 410.40

SWE-bench Verified—

τ²-Bench—

Proprietary

Grok-3 - Reasoning Beta xAI

LiveCodeBench79.40

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini-2.5-Pro-Preview-05-06 Google Deep Mind

LiveCodeBench77.10

HLE21.60

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified63.20

τ²-Bench—

Proprietary

Gemini 2.5-Pro Google Deep Mind

LiveCodeBench77.10

HLE21.60

ARC-AGI-24.90

FrontierMath - Tier 42.10

SWE-bench Verified67.20

τ²-Bench—

Proprietary

Claude Opus 4.6 Anthropic

LiveCodeBench76.00

HLE53.00

ARC-AGI-266.30

FrontierMath - Tier 422.90

SWE-bench Verified80.84

τ²-Bench91.89

Proprietary

OpenAI o3 OpenAI

LiveCodeBench75.80

HLE20.32

ARC-AGI-26.50

FrontierMath - Tier 42.10

SWE-bench Verified69.10

τ²-Bench—

Proprietary

DeepSeek V3.2-Exp DeepSeek-AI

LiveCodeBench74.10

HLE20.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified67.80

τ²-Bench66.70

Free commercial

Qwen3-235B-A22B-Thinking 阿里巴巴

LiveCodeBench74.10

HLE18.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-235B-A22B-Thinking-2507 阿里巴巴

LiveCodeBench74.10

HLE18.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Kimi-k1.6-IOI-high Moonshot AI

LiveCodeBench73.80

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

DeepSeek-R1-0528 DeepSeek-AI

LiveCodeBench73.30

HLE17.70

ARC-AGI-21.30

FrontierMath - Tier 4—

SWE-bench Verified57.60

τ²-Bench—

Free commercial

GLM-4.5 智谱AI

LiveCodeBench72.90

HLE14.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified64.20

τ²-Bench—

Free commercial

OpenAI o1 OpenAI

LiveCodeBench71.00

HLE9.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified48.90

τ²-Bench—

Proprietary

Qwen3-235B-A22B 阿里巴巴

LiveCodeBench70.70

HLE7.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified34.40

τ²-Bench34.40

Free commercial

GLM-4.5-Air 智谱AI

LiveCodeBench70.70

HLE10.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified57.60

τ²-Bench—

Free commercial

Gemini 2.5 Pro Experimental 03-25 Google Deep Mind

LiveCodeBench70.40

HLE18.80

ARC-AGI-2—

FrontierMath - Tier 44.20

SWE-bench Verified63.80

τ²-Bench—

Proprietary

OpenAI o3-mini (high)OpenAI

LiveCodeBench69.50

HLE—

ARC-AGI-2—

FrontierMath - Tier 44.20

SWE-bench Verified49.30

τ²-Bench—

Proprietary

OpenAI o3-mini (medium)OpenAI

LiveCodeBench67.40

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude Sonnet 4 Anthropic

LiveCodeBench66.00

HLE9.60

ARC-AGI-25.90

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench52.00

Proprietary

Kimi-k1.6-IOI Moonshot AI

LiveCodeBench65.90

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

DeepSeek-R1 DeepSeek-AI

LiveCodeBench65.90

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified49.20

τ²-Bench—

Free commercial

Qwen3-32B 阿里巴巴

LiveCodeBench65.70

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

QwQ-Max-Preview 阿里巴巴

LiveCodeBench65.60

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

MiniMax-M1-80k MiniMaxAI

LiveCodeBench65.00

HLE8.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified56.00

τ²-Bench—

Free commercial

Hunyuan-T1 腾讯AI实验室

LiveCodeBench64.90

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

MiniMax-M1-40k MiniMaxAI

LiveCodeBench62.30

HLE7.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified55.60

τ²-Bench—

Free commercial

Qwen3-8B 阿里巴巴

LiveCodeBench61.80

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Magistral-Medium-2506 MistralAI

LiveCodeBench59.36

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude Opus 4 Anthropic

LiveCodeBench56.60

HLE10.70

ARC-AGI-28.60

FrontierMath - Tier 44.20

SWE-bench Verified72.50

τ²-Bench72.50

Proprietary

Magistral-Small-2506 MistralAI

LiveCodeBench55.84

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Gemini 2.5 Flash Google Deep Mind

LiveCodeBench55.40

HLE11.00

ARC-AGI-2—

FrontierMath - Tier 44.20

SWE-bench Verified50.00

τ²-Bench—

Proprietary

OpenAI o1-mini OpenAI

LiveCodeBench52.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 2.5 Flash-Lite Google Deep Mind

LiveCodeBench34.30

HLE6.90

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified27.60

τ²-Bench—

Proprietary

Hunyuan-TurboS 腾讯AI实验室

LiveCodeBench32.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Qwen3-30B-A3B 阿里巴巴

LiveCodeBench29.00

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Claude Opus 5 Anthropic

LiveCodeBench—

HLE64.70

ARC-AGI-290.40

FrontierMath - Tier 4—

SWE-bench Verified96.00

τ²-Bench—

Proprietary

Muse Spark 1.1 Facebook AI研究实验室

LiveCodeBench—

HLE62.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Sort by:

Showing 50 of 94 modelsView LiveCodeBench benchmark page

Leaderboard FAQ

Where does the leaderboard data come from?

Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

How do I compare an open-source model with a closed API model?

Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.

As of 2026-07, AA Intelligence Index leaders include Claude Opus 5 (max), Claude Opus 5 (xhigh), Claude Fable 5, based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Claude Fable 5, Claude Opus 4.6 (thinking), Opus 4.7 (thinking) near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.