DataLearner logo

AI Model Leaderboards

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-07-28 08:43:41

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-08-02

#ModelScore
1
Anthropic
Claude Opus 5 (max)
61
2
Anthropic
Claude Opus 5 (xhigh)
60
3
60
5
Anthropic
Claude Opus 5 (high)
59
7
Kimi
Kimi K3 (max)
57
8
Anthropic
Claude Opus 5 (medium)
56

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-08-01

#ModelElo
1
1509
3
1502
4
1497
5
Anthropic
Opus 4.7
1492
6
Anthropic
claude-opus-5-high
1492
7
Anthropic
claude-opus-5-max
1490
8
F
Muse Spark 1.1
1490
9
F
Muse Spark
1488
10
Google Deep Mind
Gemini 3 Pro
1486
Source: LMArena

Recent Rank Changes

Risers, decliners, and new entrants across the coding, math, and agent leaderboards over the last 30 days.

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

LLM Performance Results

Data source: DataLearnerAI
No chart data available

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.

LiveCodeBench
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified51.60
τ²-Bench
Free commercial
LiveCodeBench
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified67.00
τ²-Bench
Free commercial
LiveCodeBench
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified70.60
τ²-Bench
Free commercial
Sort by:

Leaderboard FAQ

01

Where does the leaderboard data come from?

Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.

02

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

03

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

04

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

05

How do I compare an open-source model with a closed API model?

Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.

As of 2026-07, AA Intelligence Index leaders include Claude Opus 5 (max), Claude Opus 5 (xhigh), Claude Fable 5, based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Claude Fable 5, Claude Opus 4.6 (thinking), Opus 4.7 (thinking) near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.

Leading model developers

View all 100 organizations

Jump to a developer to explore its full model lineup, series, and product lines.

Today's picksRotates daily · discover more labs

Explore more

The leaderboard covers benchmarked models. Browse the full catalog by model, organization, or benchmark.