AI Model Leaderboards

Name: AI Model Performance Leaderboard
Creator: DataLearner
License: https://creativecommons.org/licenses/by/4.0/

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-07-28 08:43:41

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-08-02

#ModelScore

Claude Opus 5 (max)Anthropic

Claude Opus 5 (xhigh)Anthropic

Claude Fable 5Anthropic

GPT-5.6 Sol (max)OpenAI

Claude Opus 5 (high)Anthropic

GPT-5.6 Sol (xhigh)OpenAI

Kimi K3 (max)Kimi

Claude Opus 5 (medium)Anthropic

GPT-5.6 Sol (high)OpenAI

GPT-5.6 Terra (max)OpenAI

Source: Artificial Analysis

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-08-01

#ModelElo

Claude Fable 5Anthropic

1509

Claude Opus 4.6 (thinking)Anthropic

1505

Opus 4.7 (thinking)Anthropic

1502

Claude Opus 4.6Anthropic

1497

Opus 4.7Anthropic

1492

claude-opus-5-highAnthropic

1492

claude-opus-5-maxAnthropic

1490

Muse Spark 1.1Facebook AI研究实验室

1490

Muse SparkFacebook AI研究实验室

1488

Gemini 3 ProGoogle Deep Mind

1486

Source: LMArena

Recent Rank Changes

Risers, decliners, and new entrants across the coding, math, and agent leaderboards over the last 30 days.

Coding

Full ranking

Agent

Full ranking

View the full AI model changelog

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks

Overall

ARC-AGI-2 HLE MMLU Pro Open Benchmark Directory

Math

AIME 2025 FrontierMath MATH-500 Open Math Leaderboard

Coding

SWE-bench Verified LiveCodeBench SWE-Bench Pro Open Coding Leaderboard

Agent

τ²-Bench Terminal Bench 2.0 Aider-Polyglot Open Agent Leaderboard

Model Size:All 3B and below 7B 13B 34B 65B 100B and above

Model Type:All Reasoning Models Foundation Models Instruction/Chat Models Coding Models

License:All Open Source Closed Source

Region:All China

Recommended models

Ranked by LiveCodeBench

Current SOTA

GPT-5.1 Codex

OpenAI

85.50LiveCodeBench

View model

Best Open-Source

Codestral

MistralAI

31.50LiveCodeBench−54.00

View model

Best China-Made

Qwen3-Coder-Flash

阿里巴巴

—LiveCodeBench

View model

LLM Performance Results

Data source: DataLearnerAI

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.

Rank	Model							License
	GPT-5.1 Codex OpenAI	85.50	—	—	—	70.40	—	Proprietary	Details
	Codestral 25.01 MistralAI	37.90	—	—	—	—	—	Proprietary	Details
	Codestral MistralAI	31.50	—	—	—	—	—	Non-commercial	Details
4	Devstral Small 1.1 MistralAI	—	—	—	—	53.60	—	Free commercial	Details
5	Composer 2 Cursor	—	—	—	—	—	—	Proprietary	Details
6	Composer 2.5 Cursor	—	—	—	—	—	—	Proprietary	Details
7	GPT-5.3 Codex OpenAI	—	—	—	—	—	—	Proprietary	Details
8	Grok 4.5 xAI	—	—	—	—	—	—	Proprietary	Details
9	Devstral Small 1.0 MistralAI	—	—	—	—	46.80	—	Free commercial	Details
10	Qwen3-Coder-Flash 阿里巴巴	—	—	—	—	51.60	—	Free commercial	Details
11	GPT-5.1-Codex-Max OpenAI	—	—	—	—	76.80	—	Proprietary	Details
12	Devstral Medium MistralAI	—	—	—	—	61.60	—	Proprietary	Details
13	Qwen3-Coder-480B-A35B 阿里巴巴	—	—	—	—	67.00	—	Free commercial	Details
14	Qwen3-Coder-Next 阿里巴巴	—	—	—	—	70.60	—	Free commercial	Details
15	Grok Code Fast 1 xAI	—	—	—	—	70.80	—	Proprietary	Details
16	Grok 4 Code xAI	—	—	—	—	72.00	—	Proprietary	Details
17	GPT-5 Codex OpenAI	—	—	—	—	74.50	—	Proprietary	Details

GPT-5.1 Codex OpenAI

LiveCodeBench85.50

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified70.40

τ²-Bench—

Proprietary

Codestral 25.01 MistralAI

LiveCodeBench37.90

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Codestral MistralAI

LiveCodeBench31.50

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Non-commercial

Devstral Small 1.1 MistralAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified53.60

τ²-Bench—

Free commercial

Composer 2 Cursor

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Composer 2.5 Cursor

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

GPT-5.3 Codex OpenAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Grok 4.5 xAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

Devstral Small 1.0 MistralAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified46.80

τ²-Bench—

Free commercial

Qwen3-Coder-Flash 阿里巴巴

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified51.60

τ²-Bench—

Free commercial

GPT-5.1-Codex-Max OpenAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified76.80

τ²-Bench—

Proprietary

Devstral Medium MistralAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified61.60

τ²-Bench—

Proprietary

Qwen3-Coder-480B-A35B 阿里巴巴

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified67.00

τ²-Bench—

Free commercial

Qwen3-Coder-Next 阿里巴巴

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified70.60

τ²-Bench—

Free commercial

Grok Code Fast 1 xAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified70.80

τ²-Bench—

Proprietary

Grok 4 Code xAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified72.00

τ²-Bench—

Proprietary

GPT-5 Codex OpenAI

LiveCodeBench—

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified74.50

τ²-Bench—

Proprietary

Sort by:

Leaderboard FAQ

Where does the leaderboard data come from?

Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

How do I compare an open-source model with a closed API model?

Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.

As of 2026-07, AA Intelligence Index leaders include Claude Opus 5 (max), Claude Opus 5 (xhigh), Claude Fable 5, based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Claude Fable 5, Claude Opus 4.6 (thinking), Opus 4.7 (thinking) near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.