AI Model Leaderboards

Name: AI Model Performance Leaderboard
Creator: DataLearner
License: https://creativecommons.org/licenses/by/4.0/

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-07-28 08:43:41

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-08-02

#ModelScore

Claude Opus 5 (max)Anthropic

Claude Opus 5 (xhigh)Anthropic

Claude Fable 5Anthropic

GPT-5.6 Sol (max)OpenAI

Claude Opus 5 (high)Anthropic

GPT-5.6 Sol (xhigh)OpenAI

Kimi K3 (max)Kimi

Claude Opus 5 (medium)Anthropic

GPT-5.6 Sol (high)OpenAI

GPT-5.6 Terra (max)OpenAI

Source: Artificial Analysis

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-08-01

#ModelElo

Claude Fable 5Anthropic

1509

Claude Opus 4.6 (thinking)Anthropic

1505

Opus 4.7 (thinking)Anthropic

1502

Claude Opus 4.6Anthropic

1497

Opus 4.7Anthropic

1492

claude-opus-5-highAnthropic

1492

claude-opus-5-maxAnthropic

1490

Muse Spark 1.1Facebook AI研究实验室

1490

Muse SparkFacebook AI研究实验室

1488

Gemini 3 ProGoogle Deep Mind

1486

Source: LMArena

Recent Rank Changes

Risers, decliners, and new entrants across the coding, math, and agent leaderboards over the last 30 days.

Coding

Full ranking

Agent

Full ranking

View the full AI model changelog

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks

Overall

ARC-AGI-2 HLE MMLU Pro Open Benchmark Directory

Math

AIME 2025 FrontierMath MATH-500 Open Math Leaderboard

Coding

SWE-bench Verified LiveCodeBench SWE-Bench Pro Open Coding Leaderboard

Agent

τ²-Bench Terminal Bench 2.0 Aider-Polyglot Open Agent Leaderboard

Model Size:All 3B and below 7B 13B 34B 65B 100B and above

Model Type:All Reasoning Models Foundation Models Instruction/Chat Models Coding Models

License:All Open Source Closed Source

Region:All China

Recommended models

Ranked by HLE

Current SOTA

Claude Mythos Preview

Anthropic

64.70HLE

View model

Best Open-Source

Kimi K3

Moonshot AI

56.00HLE−8.70

View model

Best China-Made

Kimi K3

Moonshot AI

56.00HLE−8.70

View model

LLM Performance Results

Data source: DataLearnerAI

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.

Rank	Model						License
	Claude Mythos Preview Anthropic	64.70	—	—	93.90	—	Proprietary	Details
	Claude Opus 5 Anthropic	64.70	90.40	—	96.00	—	Proprietary	Details
	Muse Spark 1.1 Facebook AI研究实验室	62.10	—	—	—	—	Proprietary	Details
4	Claude Fable 5 Anthropic	59.00	—	—	95.00	—	Proprietary	Details
5	GPT-5.4 Pro OpenAI	58.70	83.30	38.00	—	—	Proprietary	Details
6	Muse Spark Facebook AI研究实验室	58.00	42.50	14.60	77.40	—	Proprietary	Details
7	Claude Opus 4.8 Anthropic	57.90	—	—	88.60	—	Proprietary	Details
8	Claude Sonnet 5 Anthropic	57.40	—	—	85.20	—	Proprietary	Details
9	GPT-5.5 Pro OpenAI	57.20	84.60	39.60	—	—	Proprietary	Details
10	Kimi K3 Moonshot AI	56.00	—	—	—	—	Free commercial	Details
11	Opus 4.7 Anthropic	54.70	75.80	22.90	87.60	—	Proprietary	Details
12	GLM-5.2 智谱AI	54.70	—	—	—	—	Free commercial	Details
13	Kimi K2.6 Moonshot AI	54.00	—	—	80.20	—	Free commercial	Details
14	Qwen3.7-Max-Preview 阿里巴巴	53.50	—	—	80.40	—	Proprietary	Details
15	Hy3 腾讯AI实验室	53.20	—	—	78.00	—	Free commercial	Details
16	Claude Opus 4.6 Anthropic	53.00	66.30	22.90	80.84	91.89	Proprietary	Details
17	GLM 5.1 智谱AI	52.30	—	—	—	—	Free commercial	Details
18	GPT-5.5 OpenAI	52.20	85.00	35.40	—	—	Proprietary	Details
19	GPT-5.4 OpenAI	52.10	77.10	27.10	—	—	Proprietary	Details
20	Gemini 3.1 Pro Preview Google Deep Mind	51.40	77.10	16.70	80.60	90.80	Proprietary	Details
21	Kimi K2 Thinking Moonshot AI	51.00	—	—	71.30	—	Free commercial	Details
22	Qwen 3.6 Plus Preview 阿里巴巴	50.60	—	—	78.80	—	Proprietary	Details
23	GLM-5 智谱AI	50.40	4.90	2.10	77.80	89.70	Free commercial	Details
24	Kimi K2.5 Moonshot AI	50.20	11.80	4.20	76.80	—	Free commercial	Details
25	Qwen3.6-Max-Preview 阿里巴巴	50.20	—	—	78.80	—	Proprietary	Details
26	GPT-5.2 Pro OpenAI	50.00	54.20	31.30	—	—	Proprietary	Details
27	Qwen3-Max-Thinking 阿里巴巴	49.80	—	—	75.30	82.10	Proprietary	Details
28	Claude Sonnet 4.6 Anthropic	49.00	58.30	8.30	79.60	—	Proprietary	Details
29	Qwen3.5-27B 阿里巴巴	48.50	—	—	72.40	79.00	Free commercial	Details
30	Gemini 3 Deep Think - 2620 Google Deep Mind	48.40	84.60	—	—	—	Proprietary	Details
31	Qwen3.5-397B-A17B 阿里巴巴	48.30	—	—	76.40	86.70	Free commercial	Details
32	DeepSeek-V4-Pro DeepSeek-AI	48.20	—	—	80.60	—	Free commercial	Details
33	Step 3.7 Flash StepFunAI	47.20	—	—	—	—	Free commercial	Details
34	Inkling Thinking Machines Lab	46.00	—	—	77.60	—	Free commercial	Details
35	Gemini 3.0 Pro (Preview 11-2025) Google Deep Mind	45.80	45.10	18.80	76.20	85.40	Proprietary	Details
36	GPT-5.2 OpenAI	45.50	54.20	18.80	80.00	82.00	Proprietary	Details
37	DeepSeek-V4-Flash DeepSeek-AI	45.10	—	—	79.00	—	Free commercial	Details
38	Grok 4 Heavy xAI	44.40	—	2.10	73.50	—	Proprietary	Details
39	Gemini 3.0 Flash Google Deep Mind	43.50	33.60	4.20	68.70	90.20	Proprietary	Details
40	Opus 4.5 Anthropic	43.20	37.60	4.20	80.90	81.99	Proprietary	Details
41	GLM-4.7 智谱AI	42.80	—	2.10	73.80	87.40	Free commercial	Details
42	GPT-5.1 OpenAI	42.70	17.60	12.50	76.30	—	Proprietary	Details
43	GPT-5-Pro OpenAI	42.00	18.00	14.60	—	—	Proprietary	Details
44	GPT-5.4 mini OpenAI	41.50	—	2.10	—	—	Proprietary	Details
45	Gemini 3.5 Flash Google Deep Mind	40.20	72.10	—	—	—	Proprietary	Details
46	Grok 4 xAI	38.60	15.90	2.10	58.60	—	Proprietary	Details
47	GPT-5.4 nano OpenAI	37.70	—	6.30	—	—	Proprietary	Details
48	GPT-5 OpenAI	35.20	9.90	12.50	72.80	80.00	Proprietary	Details
49	Gemini 2.5 Deep Think Google Deep Mind	34.80	—	10.40	—	—	Proprietary	Details
50	Claude Sonnet 4.5 Anthropic	33.60	13.60	4.20	82.00	84.70	Proprietary	Details

Claude Mythos Preview Anthropic

HLE64.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified93.90

τ²-Bench—

Proprietary

Claude Opus 5 Anthropic

HLE64.70

ARC-AGI-290.40

FrontierMath - Tier 4—

SWE-bench Verified96.00

τ²-Bench—

Proprietary

Muse Spark 1.1 Facebook AI研究实验室

HLE62.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude Fable 5 Anthropic

HLE59.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified95.00

τ²-Bench—

Proprietary

GPT-5.4 Pro OpenAI

HLE58.70

ARC-AGI-283.30

FrontierMath - Tier 438.00

SWE-bench Verified—

τ²-Bench—

Proprietary

Muse Spark Facebook AI研究实验室

HLE58.00

ARC-AGI-242.50

FrontierMath - Tier 414.60

SWE-bench Verified77.40

τ²-Bench—

Proprietary

Claude Opus 4.8 Anthropic

HLE57.90

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified88.60

τ²-Bench—

Proprietary

Claude Sonnet 5 Anthropic

HLE57.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified85.20

τ²-Bench—

Proprietary

GPT-5.5 Pro OpenAI

HLE57.20

ARC-AGI-284.60

FrontierMath - Tier 439.60

SWE-bench Verified—

τ²-Bench—

Proprietary

Kimi K3 Moonshot AI

HLE56.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Opus 4.7 Anthropic

HLE54.70

ARC-AGI-275.80

FrontierMath - Tier 422.90

SWE-bench Verified87.60

τ²-Bench—

Proprietary

GLM-5.2 智谱AI

HLE54.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Kimi K2.6 Moonshot AI

HLE54.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench—

Free commercial

Qwen3.7-Max-Preview 阿里巴巴

HLE53.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.40

τ²-Bench—

Proprietary

Hy3 腾讯AI实验室

HLE53.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.00

τ²-Bench—

Free commercial

Claude Opus 4.6 Anthropic

HLE53.00

ARC-AGI-266.30

FrontierMath - Tier 422.90

SWE-bench Verified80.84

τ²-Bench91.89

Proprietary

GLM 5.1 智谱AI

HLE52.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

GPT-5.5 OpenAI

HLE52.20

ARC-AGI-285.00

FrontierMath - Tier 435.40

SWE-bench Verified—

τ²-Bench—

Proprietary

GPT-5.4 OpenAI

HLE52.10

ARC-AGI-277.10

FrontierMath - Tier 427.10

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 3.1 Pro Preview Google Deep Mind

HLE51.40

ARC-AGI-277.10

FrontierMath - Tier 416.70

SWE-bench Verified80.60

τ²-Bench90.80

Proprietary

Kimi K2 Thinking Moonshot AI

HLE51.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified71.30

τ²-Bench—

Free commercial

Qwen 3.6 Plus Preview 阿里巴巴

HLE50.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.80

τ²-Bench—

Proprietary

GLM-5 智谱AI

HLE50.40

ARC-AGI-24.90

FrontierMath - Tier 42.10

SWE-bench Verified77.80

τ²-Bench89.70

Free commercial

Kimi K2.5 Moonshot AI

HLE50.20

ARC-AGI-211.80

FrontierMath - Tier 44.20

SWE-bench Verified76.80

τ²-Bench—

Free commercial

Qwen3.6-Max-Preview 阿里巴巴

HLE50.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified78.80

τ²-Bench—

Proprietary

GPT-5.2 Pro OpenAI

HLE50.00

ARC-AGI-254.20

FrontierMath - Tier 431.30

SWE-bench Verified—

τ²-Bench—

Proprietary

Qwen3-Max-Thinking 阿里巴巴

HLE49.80

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified75.30

τ²-Bench82.10

Proprietary

Claude Sonnet 4.6 Anthropic

HLE49.00

ARC-AGI-258.30

FrontierMath - Tier 48.30

SWE-bench Verified79.60

τ²-Bench—

Proprietary

Qwen3.5-27B 阿里巴巴

HLE48.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified72.40

τ²-Bench79.00

Free commercial

Gemini 3 Deep Think - 2620 Google Deep Mind

HLE48.40

ARC-AGI-284.60

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Qwen3.5-397B-A17B 阿里巴巴

HLE48.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified76.40

τ²-Bench86.70

Free commercial

DeepSeek-V4-Pro DeepSeek-AI

HLE48.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.60

τ²-Bench—

Free commercial

Step 3.7 Flash StepFunAI

HLE47.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Inkling Thinking Machines Lab

HLE46.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified77.60

τ²-Bench—

Free commercial

Gemini 3.0 Pro (Preview 11-2025)Google Deep Mind

HLE45.80

ARC-AGI-245.10

FrontierMath - Tier 418.80

SWE-bench Verified76.20

τ²-Bench85.40

Proprietary

GPT-5.2 OpenAI

HLE45.50

ARC-AGI-254.20

FrontierMath - Tier 418.80

SWE-bench Verified80.00

τ²-Bench82.00

Proprietary

DeepSeek-V4-Flash DeepSeek-AI

HLE45.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified79.00

τ²-Bench—

Free commercial

Grok 4 Heavy xAI

HLE44.40

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.50

τ²-Bench—

Proprietary

Gemini 3.0 Flash Google Deep Mind

HLE43.50

ARC-AGI-233.60

FrontierMath - Tier 44.20

SWE-bench Verified68.70

τ²-Bench90.20

Proprietary

Opus 4.5 Anthropic

HLE43.20

ARC-AGI-237.60

FrontierMath - Tier 44.20

SWE-bench Verified80.90

τ²-Bench81.99

Proprietary

GLM-4.7 智谱AI

HLE42.80

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.80

τ²-Bench87.40

Free commercial

GPT-5.1 OpenAI

HLE42.70

ARC-AGI-217.60

FrontierMath - Tier 412.50

SWE-bench Verified76.30

τ²-Bench—

Proprietary

GPT-5-Pro OpenAI

HLE42.00

ARC-AGI-218.00

FrontierMath - Tier 414.60

SWE-bench Verified—

τ²-Bench—

Proprietary

GPT-5.4 mini OpenAI

HLE41.50

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified—

τ²-Bench—

Proprietary

Gemini 3.5 Flash Google Deep Mind

HLE40.20

ARC-AGI-272.10

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Grok 4 xAI

HLE38.60

ARC-AGI-215.90

FrontierMath - Tier 42.10

SWE-bench Verified58.60

τ²-Bench—

Proprietary

GPT-5.4 nano OpenAI

HLE37.70

ARC-AGI-2—

FrontierMath - Tier 46.30

SWE-bench Verified—

τ²-Bench—

Proprietary

GPT-5 OpenAI

HLE35.20

ARC-AGI-29.90

FrontierMath - Tier 412.50

SWE-bench Verified72.80

τ²-Bench80.00

Proprietary

Gemini 2.5 Deep Think Google Deep Mind

HLE34.80

ARC-AGI-2—

FrontierMath - Tier 410.40

SWE-bench Verified—

τ²-Bench—

Proprietary

Claude Sonnet 4.5 Anthropic

HLE33.60

ARC-AGI-213.60

FrontierMath - Tier 44.20

SWE-bench Verified82.00

τ²-Bench84.70

Proprietary

Sort by:

Showing 50 of 233 modelsView HLE benchmark page

Leaderboard FAQ

Where does the leaderboard data come from?

Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

How do I compare an open-source model with a closed API model?

Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.

As of 2026-07, AA Intelligence Index leaders include Claude Opus 5 (max), Claude Opus 5 (xhigh), Claude Fable 5, based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Claude Fable 5, Claude Opus 4.6 (thinking), Opus 4.7 (thinking) near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.