Where does the leaderboard data come from?

Scores are aggregated from official model cards, technical reports, papers, vendor announcements, and reproducible third-party evaluations.

Why do scores for the same model differ across benchmarks?

Different benchmarks measure different capabilities — reasoning, math, coding, or agent use — so the same model often gets very different scores across benchmarks.

How often is the leaderboard updated?

Data is revalidated every 5 minutes; new models and evaluations are added on publication. The "Updated on" indicator shows the most recent refresh time.

How should I read the composite ranking?

The composite ranking aggregates a model's standing across several core benchmarks. For production decisions, drill into the specific benchmark closest to your workload.

How do I compare an open-source model with a closed API model?

Use the license filter to view open and closed models together, then compare the same benchmark column. Also consider API pricing vs. self-hosting cost.

AI Model Leaderboards

Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.

View benchmark detailsUpdated on 2026-05-02 07:14:49

As of 2026-05, AA Intelligence Index leaders include GPT-5.5 (xhigh), GPT-5.5 (high), Opus 4.7 (max), based on 10 standardized capability benchmarks.

On the user-preference side, LMArena Text Generation currently ranks Opus 4.7 (thinking), Claude Opus 4.6 (thinking), Claude Opus 4.6 near the top via anonymous A/B voting.

Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-05-10

#ModelScore

GPT-5.5 (xhigh)OpenAI

GPT-5.5 (high)OpenAI

Opus 4.7 (max)Anthropic

Gemini 3.1 Pro PreviewGoogle Deep Mind

GPT-5.5 (medium)OpenAI

Kimi K2.6Moonshot AI

MiMo-V2.5-ProXiaomi

GPT-5.3 Codex (xhigh)OpenAI

Grok 4.3xAI

Muse SparkFacebook AI研究实验室

Source: Artificial Analysis

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks

Overall

ARC-AGI-2 HLE MMLU Pro Open Benchmark Directory

Math

AIME 2025 FrontierMath MATH-500

Scores shown are the best result across all evaluation modes. Click a model name for the full breakdown.

Rank	Model						License
	Muse Spark Facebook AI研究实验室	58.00	42.50	14.60	77.40	—	Proprietary
	GPT-5.5 Pro OpenAI	57.20	84.60	39.60	—	—	Proprietary
	Opus 4.7 Anthropic	54.70	75.80	22.90	87.60	—	Proprietary
4	Kimi K2.6 Moonshot AI	54.00	—	—	80.20	—	Free commercial
5	Claude Opus 4.6 Anthropic	53.00	66.30	22.90	80.84	91.89	Proprietary
6	GLM 5.1 智谱AI	52.30	—	—	—	—	Free commercial
7	GPT-5.5 OpenAI	52.20	85.00	35.40	—	—	Proprietary
8	Kimi K2 Thinking Moonshot AI	51.00	—	—	71.30	—	Free commercial
9	GPT-5.2 Pro OpenAI	50.00	54.20	31.30	—	—	Proprietary
10	Qwen3-Max-Thinking 阿里巴巴	49.80	—	—	75.30	82.10	Proprietary
11	Qwen3.5-27B 阿里巴巴	48.50	—	—	72.40	79.00	Free commercial
12	Gemini 3 Deep Think - 2620 Google Deep Mind	48.40	84.60	—	—	—	Proprietary
13	DeepSeek-V4-Pro DeepSeek-AI	48.20	—	—	80.60	—	Free commercial
14	DeepSeek-V4-Flash DeepSeek-AI	45.10	—	—	79.00	—	Free commercial
15	Opus 4.5 Anthropic	43.20	37.60	4.20	80.90	81.99	Proprietary
16	GPT-5.1 OpenAI	42.70	17.60	12.50	76.30	—	Proprietary
17	GPT-5-Pro OpenAI	42.00	18.00	14.60	—	—	Proprietary
18	GPT-5.4 mini OpenAI	41.50	—	2.10	—	—	Proprietary
19	Grok 4 xAI	38.60	15.90	2.10	58.60	—	Proprietary
20	DeepSeek V3.2 Speciale DeepSeek-AI	30.60	—	—	—	—	Free commercial
21	MiniMax-M2.7 MiniMaxAI	28.00	—	—	—	—	Non-commercial
22	DeepSeek V3.2 DeepSeek-AI	25.10	4.00	2.10	73.10	80.30	Free commercial
23	Qwen3.6-27B 阿里巴巴	24.00	—	—	77.20	—	Free commercial
24	Gemini 2.5-Pro Google Deep Mind	21.60	4.90	2.10	67.20	—	Proprietary
25	Gemini-2.5-Pro-Preview-05-06 Google Deep Mind	21.60	—	2.10	63.20	—	Proprietary
26	Qwen3.6-35B-A3B 阿里巴巴	21.40	—	—	73.40	—	Free commercial
27	o3-pro OpenAI	21.00	—	—	75.00	—	Proprietary
28	OpenAI o3 OpenAI	20.32	6.50	2.10	69.10	—	Proprietary
29	DeepSeek V3.2-Exp DeepSeek-AI	20.30	—	—	67.80	66.70	Free commercial
30	MiniMax M2.5 MiniMaxAI	19.40	4.90	—	80.20	—	Free commercial
31	GPT OSS 120B OpenAI	19.00	—	—	60.10	—	Free commercial
32	Gemini 2.5 Pro Experimental 03-25 Google Deep Mind	18.80	—	4.20	63.80	—	Proprietary
33	Qwen3-235B-A22B-Thinking 阿里巴巴	18.20	—	—	—	—	Free commercial
34	Qwen3-235B-A22B-Thinking-2507 阿里巴巴	18.20	—	—	—	—	Free commercial
35	DeepSeek-R1-0528 DeepSeek-AI	17.70	1.30	—	57.60	—	Free commercial
36	OpenAI o4 - mini OpenAI	17.70	—	6.30	68.10	56.90	Proprietary
37	Grok 4.1 Fast xAI	17.60	—	—	—	82.71	Proprietary
38	GPT OSS 20B OpenAI	17.30	—	—	34.00	47.70	Free commercial
39	GLM-4.5 智谱AI	14.40	—	—	64.20	—	Free commercial
40	GLM-4.7-Flash 智谱AI	14.40	—	—	59.20	79.50	Free commercial
41	OpenAI o3-mini OpenAI	13.40	—	4.20	40.80	—	Proprietary
42	Gemini 2.5 Flash Google Deep Mind	11.00	—	4.20	50.00	—	Proprietary
43	Claude Opus 4 Anthropic	10.70	8.60	4.20	72.50	72.50	Proprietary
44	GLM-4.5-Air 智谱AI	10.60	—	—	57.60	—	Free commercial
45	Claude Sonnet 4 Anthropic	9.60	5.90	—	80.20	52.00	Proprietary
46	OpenAI o1 OpenAI	9.10	—	—	48.90	—	Proprietary
47	MiniMax-M1-80k MiniMaxAI	8.40	—	—	56.00	—	Free commercial
48	Qwen3-235B-A22B 阿里巴巴	7.60	—	—	34.40	34.40	Free commercial
49	MiniMax-M1-40k MiniMaxAI	7.20	—	—	55.60	—	Free commercial
50	Gemini 2.5 Flash-Lite Google Deep Mind	6.90	—	—	27.60	—	Proprietary

Muse Spark

Facebook AI研究实验室

HLE58.00

ARC-AGI-242.50

FrontierMath - Tier 414.60

SWE-bench Verified77.40

τ²-Bench—

Proprietary

GPT-5.5 Pro

OpenAI

HLE57.20

ARC-AGI-284.60

FrontierMath - Tier 439.60

SWE-bench Verified—

τ²-Bench—

Proprietary

Opus 4.7

Anthropic

HLE54.70

ARC-AGI-275.80

FrontierMath - Tier 422.90

SWE-bench Verified87.60

τ²-Bench—

Proprietary

Kimi K2.6

Moonshot AI

HLE54.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench—

Free commercial

Claude Opus 4.6

Anthropic

HLE53.00

ARC-AGI-266.30

FrontierMath - Tier 422.90

SWE-bench Verified80.84

τ²-Bench91.89

Proprietary

GLM 5.1

智谱AI

HLE52.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

GPT-5.5

OpenAI

HLE52.20

ARC-AGI-285.00

FrontierMath - Tier 435.40

SWE-bench Verified—

τ²-Bench—

Proprietary

Kimi K2 Thinking

Moonshot AI

HLE51.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified71.30

τ²-Bench—

Free commercial

GPT-5.2 Pro

OpenAI

HLE50.00

ARC-AGI-254.20

FrontierMath - Tier 431.30

SWE-bench Verified—

τ²-Bench—

Proprietary

Qwen3-Max-Thinking

阿里巴巴

HLE49.80

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified75.30

τ²-Bench82.10

Proprietary

Qwen3.5-27B

阿里巴巴

HLE48.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified72.40

τ²-Bench79.00

Free commercial

Gemini 3 Deep Think - 2620

Google Deep Mind

HLE48.40

ARC-AGI-284.60

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Proprietary

DeepSeek-V4-Pro

DeepSeek-AI

HLE48.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.60

τ²-Bench—

Free commercial

DeepSeek-V4-Flash

DeepSeek-AI

HLE45.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified79.00

τ²-Bench—

Free commercial

Opus 4.5

Anthropic

HLE43.20

ARC-AGI-237.60

FrontierMath - Tier 44.20

SWE-bench Verified80.90

τ²-Bench81.99

Proprietary

GPT-5.1

OpenAI

HLE42.70

ARC-AGI-217.60

FrontierMath - Tier 412.50

SWE-bench Verified76.30

τ²-Bench—

Proprietary

GPT-5-Pro

OpenAI

HLE42.00

ARC-AGI-218.00

FrontierMath - Tier 414.60

SWE-bench Verified—

τ²-Bench—

Proprietary

GPT-5.4 mini

OpenAI

HLE41.50

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified—

τ²-Bench—

Proprietary

Grok 4

xAI

HLE38.60

ARC-AGI-215.90

FrontierMath - Tier 42.10

SWE-bench Verified58.60

τ²-Bench—

Proprietary

DeepSeek V3.2 Speciale

DeepSeek-AI

HLE30.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

MiniMax-M2.7

MiniMaxAI

HLE28.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Non-commercial

DeepSeek V3.2

DeepSeek-AI

HLE25.10

ARC-AGI-24.00

FrontierMath - Tier 42.10

SWE-bench Verified73.10

τ²-Bench80.30

Free commercial

Qwen3.6-27B

阿里巴巴

HLE24.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified77.20

τ²-Bench—

Free commercial

Gemini 2.5-Pro

Google Deep Mind

HLE21.60

ARC-AGI-24.90

FrontierMath - Tier 42.10

SWE-bench Verified67.20

τ²-Bench—

Proprietary

Gemini-2.5-Pro-Preview-05-06

Google Deep Mind

HLE21.60

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified63.20

τ²-Bench—

Proprietary

Qwen3.6-35B-A3B

阿里巴巴

HLE21.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified73.40

τ²-Bench—

Free commercial

o3-pro

OpenAI

HLE21.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified75.00

τ²-Bench—

Proprietary

OpenAI o3

OpenAI

HLE20.32

ARC-AGI-26.50

FrontierMath - Tier 42.10

SWE-bench Verified69.10

τ²-Bench—

Proprietary

DeepSeek V3.2-Exp

DeepSeek-AI

HLE20.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified67.80

τ²-Bench66.70

Free commercial

MiniMax M2.5

MiniMaxAI

HLE19.40

ARC-AGI-24.90

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench—

Free commercial

GPT OSS 120B

OpenAI

HLE19.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified60.10

τ²-Bench—

Free commercial

Gemini 2.5 Pro Experimental 03-25

Google Deep Mind

HLE18.80

ARC-AGI-2—

FrontierMath - Tier 44.20

SWE-bench Verified63.80

τ²-Bench—

Proprietary

Qwen3-235B-A22B-Thinking

阿里巴巴

HLE18.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-235B-A22B-Thinking-2507

阿里巴巴

HLE18.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

DeepSeek-R1-0528

DeepSeek-AI

HLE17.70

ARC-AGI-21.30

FrontierMath - Tier 4—

SWE-bench Verified57.60

τ²-Bench—

Free commercial

OpenAI o4 - mini

OpenAI

HLE17.70

ARC-AGI-2—

FrontierMath - Tier 46.30

SWE-bench Verified68.10

τ²-Bench56.90

Proprietary

Grok 4.1 Fast

xAI

HLE17.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench82.71

Proprietary

GPT OSS 20B

OpenAI

HLE17.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified34.00

τ²-Bench47.70

Free commercial

GLM-4.5

智谱AI

HLE14.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified64.20

τ²-Bench—

Free commercial

GLM-4.7-Flash

智谱AI

HLE14.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified59.20

τ²-Bench79.50

Free commercial

OpenAI o3-mini

OpenAI

HLE13.40

ARC-AGI-2—

FrontierMath - Tier 44.20

SWE-bench Verified40.80

τ²-Bench—

Proprietary

Gemini 2.5 Flash

Google Deep Mind

HLE11.00

ARC-AGI-2—

FrontierMath - Tier 44.20

SWE-bench Verified50.00

τ²-Bench—

Proprietary

Claude Opus 4

Anthropic

HLE10.70

ARC-AGI-28.60

FrontierMath - Tier 44.20

SWE-bench Verified72.50

τ²-Bench72.50

Proprietary

GLM-4.5-Air

智谱AI

HLE10.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified57.60

τ²-Bench—

Free commercial

Claude Sonnet 4

Anthropic

HLE9.60

ARC-AGI-25.90

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench52.00

Proprietary

OpenAI o1

OpenAI

HLE9.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified48.90

τ²-Bench—

Proprietary

MiniMax-M1-80k

MiniMaxAI

HLE8.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified56.00

τ²-Bench—

Free commercial

Qwen3-235B-A22B

阿里巴巴

HLE7.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified34.40

τ²-Bench34.40

Free commercial

MiniMax-M1-40k

MiniMaxAI

HLE7.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified55.60

τ²-Bench—

Free commercial

Gemini 2.5 Flash-Lite

Google Deep Mind

HLE6.90

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified27.60

τ²-Bench—

Proprietary

Sort by:

Showing 50 of 80 modelsView HLE benchmark page

AI Model Leaderboards

Composite Rankings

AA Intelligence Index

Per-Benchmark Rankings

LMArena Text Generation

LLM Performance Results

Leaderboard FAQ