Open Source LLM Leaderboard

Name: Open Source LLM Leaderboard
Creator: DataLearner
License: https://creativecommons.org/licenses/by/4.0/

Track benchmark rankings for open-weight and open-source AI models, then compare score, size, and license signals in one place.

View benchmark detailsUpdated on 2026-06-17 07:42:33

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-06-13

#ModelScore

Claude Fable 5 (with fallback)Anthropic

Claude Opus 4.8 (max)Anthropic

GPT-5.5 (xhigh)OpenAI

GPT-5.5 (high)OpenAI

Opus 4.7 (max)Anthropic

Gemini 3.1 Pro PreviewGoogle Deep Mind

GPT-5.5 (medium)OpenAI

阿

Qwen3.7 Max阿里巴巴

Gemini 3.5 FlashGoogle Deep Mind

Gemini 3.5 Flash (medium)Google

Source: Artificial Analysis

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-06-10

#ModelElo

claude-fable-5Anthropic

1510

Claude Opus 4.6 (thinking)Anthropic

1504

Opus 4.7 (thinking)Anthropic

1502

Claude Opus 4.6Anthropic

1498

Opus 4.7Anthropic

1492

Muse SparkFacebook AI研究实验室

1487

Gemini 3.1 Pro PreviewGoogle Deep Mind

1487

Gemini 3.0 Pro (Preview 11-2025)Google Deep Mind

1486

claude-opus-4-8-thinkingAnthropic

1486

GPT-5.5 (high)OpenAI

1481

Source: LMArena

Leading model developers

View all 99 organizations

Jump to a developer to explore its full model lineup, series, and product lines.

xAI

百度

Today's picksRotates daily · discover more labs

达摩院2 models · Overseas

上海人工智能实验室10 models · Overseas

BigScience2 models · Overseas

Moonshot AI14 models · Overseas

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Benchmark Tracks

Overall

ARC-AGI-2 HLE MMLU Pro Open Benchmark Directory

Math

AIME 2025 FrontierMath MATH-500 Open Math Leaderboard

Coding

SWE-bench Verified LiveCodeBench SWE-Bench Pro Open Coding Leaderboard

Agent

τ²-Bench Terminal Bench 2.0 Aider-Polyglot Open Agent Leaderboard

Model Size:All 3B and below 7B 13B 34B 65B 100B and above

Model Type:All Reasoning Models Foundation Models Instruction/Chat Models Coding Models

Source:All Open Source Closed Source

Origin:All China

Recommended models

Ranked by FrontierMath

Current SOTA

Kimi K2

Moonshot AI

2.10FrontierMath

View model

Best Open-Source

Kimi K2

Moonshot AI

2.10FrontierMath

View model

Best China-Made

Kimi K2

Moonshot AI

2.10FrontierMath

View model

LLM Performance Results

Data source: DataLearnerAI

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.

Rank	Model						License
	Kimi K2 Moonshot AI	4.70	—	0.01	51.80	64.30	Free commercial	Details
	DeepSeek-V3 DeepSeek-AI	—	—	—	—	—	Free commercial	Details
	Llama 4 Maverick Facebook AI研究实验室	—	—	—	—	—	Free commercial	Details
4	Grok 2 xAI	—	—	—	—	—	Free commercial	Details
5	Qwen3-30B-A3B-2507 阿里巴巴	9.80	—	—	22.00	49.00	Free commercial	Details
6	Gemma 4 26B A4B DeepMind	17.20	—	—	—	68.20	Free commercial	Details
7	DeepSeek V3.2-Exp DeepSeek-AI	20.30	—	—	67.80	66.70	Free commercial	Details
8	MiniMax-M1-80k MiniMaxAI	8.40	—	—	56.00	—	Free commercial	Details
9	DeepSeek-V4-Flash DeepSeek-AI	45.10	—	—	79.00	—	Free commercial	Details
10	DeepSeek-V4-Pro DeepSeek-AI	48.20	—	—	80.60	—	Free commercial	Details
11	Qwen3-235B-A22B 阿里巴巴	7.60	—	—	34.40	34.40	Free commercial	Details
12	MiniMax-M1-40k MiniMaxAI	7.20	—	—	55.60	—	Free commercial	Details
13	GLM-4.7-Flash 智谱AI	14.40	—	—	59.20	79.50	Free commercial	Details
14	GLM-5.2 智谱AI	54.70	—	—	—	—	Free commercial	Details
15	Kimi K2.6 Moonshot AI	54.00	—	—	80.20	—	Free commercial	Details
16	GLM 5.1 智谱AI	52.30	—	—	—	—	Free commercial	Details
17	Kimi K2 Thinking Moonshot AI	51.00	—	—	71.30	—	Free commercial	Details
18	GLM-5 智谱AI	50.40	4.90	2.10	77.80	89.70	Free commercial	Details
19	Kimi K2.5 Moonshot AI	50.20	11.80	4.20	76.80	—	Free commercial	Details
20	GLM-4.6 智谱AI	30.40	—	2.10	68.00	75.90	Free commercial	Details
21	DeepSeek-V3-0324 DeepSeek-AI	5.20	—	—	38.80	38.80	Free commercial	Details
22	Qwen3.5-27B 阿里巴巴	48.50	—	—	72.40	79.00	Free commercial	Details
23	Qwen3.5-397B-A17B 阿里巴巴	48.30	—	—	76.40	86.70	Free commercial	Details
24	GLM-4.7 智谱AI	42.80	—	2.10	73.80	87.40	Free commercial	Details
25	DeepSeek V3.2 Speciale DeepSeek-AI	30.60	—	—	—	—	Free commercial	Details
26	MiniMax-M2.7 MiniMaxAI	28.00	—	—	—	—	Non-commercial	Details
27	Gemma 4 31B DeepMind	26.50	—	—	—	76.90	Free commercial	Details
28	DeepSeek V3.2 DeepSeek-AI	25.10	4.00	2.10	73.10	80.30	Free commercial	Details
29	Qwen3.6-27B 阿里巴巴	24.00	—	—	77.20	—	Free commercial	Details
30	M2.1 MiniMaxAI	22.00	—	—	74.80	—	Free commercial	Details
31	DeepSeek-V3.1 Terminus DeepSeek-AI	21.70	—	—	68.40	37.00	Free commercial	Details
32	Kimi K2 0905 Moonshot AI	21.70	—	—	69.20	—	Free commercial	Details
33	Qwen3.6-35B-A3B 阿里巴巴	21.40	—	—	73.40	—	Free commercial	Details
34	MiniMax M2.5 MiniMaxAI	19.40	4.90	—	80.20	—	Free commercial	Details
35	GPT OSS 120B OpenAI	19.00	—	—	60.10	—	Free commercial	Details
36	Qwen3-235B-A22B-Thinking 阿里巴巴	18.20	—	—	—	—	Free commercial	Details
37	Qwen3-235B-A22B-Thinking-2507 阿里巴巴	18.20	—	—	—	—	Free commercial	Details
38	DeepSeek-R1-0528 DeepSeek-AI	17.70	1.30	—	57.60	—	Free commercial	Details
39	GPT OSS 20B OpenAI	17.30	—	—	34.00	47.70	Free commercial	Details
40	DeepSeek-V3.1 DeepSeek-AI	15.90	—	—	66.00	—	Free commercial	Details
41	GLM-4.5 智谱AI	14.40	—	—	64.20	—	Free commercial	Details
42	MiniMax M2 MiniMaxAI	12.50	—	—	69.40	77.20	Free commercial	Details
43	GLM-4.5-Air 智谱AI	10.60	—	—	57.60	—	Free commercial	Details
44	Qwen3-32B 阿里巴巴	—	—	—	—	—	Free commercial	Details
45	Llama 4 Scout Facebook AI研究实验室	—	—	—	—	—	Free commercial	Details
46	MiniMax M3 MiniMaxAI	—	—	—	—	—	Non-commercial	Details
47	Llama3.1-405B Instruct Facebook AI研究实验室	—	—	—	—	—	Free commercial	Details
48	QwQ-32B 阿里巴巴	—	—	—	—	—	Free commercial	Details
49	Pangu Embedded 华为	—	—	—	—	—	Free commercial	Details
50	ERNIE-4.5-300B-A47B 百度	—	—	—	—	—	Free commercial	Details

Kimi K2 Moonshot AI

HLE4.70

ARC-AGI-2—

FrontierMath - Tier 40.01

SWE-bench Verified51.80

τ²-Bench64.30

Free commercial

DeepSeek-V3 DeepSeek-AI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama 4 Maverick Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Grok 2 xAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-30B-A3B-2507 阿里巴巴

HLE9.80

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified22.00

τ²-Bench49.00

Free commercial

Gemma 4 26B A4B DeepMind

HLE17.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench68.20

Free commercial

DeepSeek V3.2-Exp DeepSeek-AI

HLE20.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified67.80

τ²-Bench66.70

Free commercial

MiniMax-M1-80k MiniMaxAI

HLE8.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified56.00

τ²-Bench—

Free commercial

DeepSeek-V4-Flash DeepSeek-AI

HLE45.10

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified79.00

τ²-Bench—

Free commercial

DeepSeek-V4-Pro DeepSeek-AI

HLE48.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.60

τ²-Bench—

Free commercial

Qwen3-235B-A22B 阿里巴巴

HLE7.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified34.40

τ²-Bench34.40

Free commercial

MiniMax-M1-40k MiniMaxAI

HLE7.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified55.60

τ²-Bench—

Free commercial

GLM-4.7-Flash 智谱AI

HLE14.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified59.20

τ²-Bench79.50

Free commercial

GLM-5.2 智谱AI

HLE54.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Kimi K2.6 Moonshot AI

HLE54.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench—

Free commercial

GLM 5.1 智谱AI

HLE52.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Kimi K2 Thinking Moonshot AI

HLE51.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified71.30

τ²-Bench—

Free commercial

GLM-5 智谱AI

HLE50.40

ARC-AGI-24.90

FrontierMath - Tier 42.10

SWE-bench Verified77.80

τ²-Bench89.70

Free commercial

Kimi K2.5 Moonshot AI

HLE50.20

ARC-AGI-211.80

FrontierMath - Tier 44.20

SWE-bench Verified76.80

τ²-Bench—

Free commercial

GLM-4.6 智谱AI

HLE30.40

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified68.00

τ²-Bench75.90

Free commercial

DeepSeek-V3-0324 DeepSeek-AI

HLE5.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified38.80

τ²-Bench38.80

Free commercial

Qwen3.5-27B 阿里巴巴

HLE48.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified72.40

τ²-Bench79.00

Free commercial

Qwen3.5-397B-A17B 阿里巴巴

HLE48.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified76.40

τ²-Bench86.70

Free commercial

GLM-4.7 智谱AI

HLE42.80

ARC-AGI-2—

FrontierMath - Tier 42.10

SWE-bench Verified73.80

τ²-Bench87.40

Free commercial

DeepSeek V3.2 Speciale DeepSeek-AI

HLE30.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

MiniMax-M2.7 MiniMaxAI

HLE28.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Non-commercial

Gemma 4 31B DeepMind

HLE26.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench76.90

Free commercial

DeepSeek V3.2 DeepSeek-AI

HLE25.10

ARC-AGI-24.00

FrontierMath - Tier 42.10

SWE-bench Verified73.10

τ²-Bench80.30

Free commercial

Qwen3.6-27B 阿里巴巴

HLE24.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified77.20

τ²-Bench—

Free commercial

M2.1 MiniMaxAI

HLE22.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified74.80

τ²-Bench—

Free commercial

DeepSeek-V3.1 Terminus DeepSeek-AI

HLE21.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified68.40

τ²-Bench37.00

Free commercial

Kimi K2 0905 Moonshot AI

HLE21.70

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.20

τ²-Bench—

Free commercial

Qwen3.6-35B-A3B 阿里巴巴

HLE21.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified73.40

τ²-Bench—

Free commercial

MiniMax M2.5 MiniMaxAI

HLE19.40

ARC-AGI-24.90

FrontierMath - Tier 4—

SWE-bench Verified80.20

τ²-Bench—

Free commercial

GPT OSS 120B OpenAI

HLE19.00

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified60.10

τ²-Bench—

Free commercial

Qwen3-235B-A22B-Thinking 阿里巴巴

HLE18.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Qwen3-235B-A22B-Thinking-2507 阿里巴巴

HLE18.20

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

DeepSeek-R1-0528 DeepSeek-AI

HLE17.70

ARC-AGI-21.30

FrontierMath - Tier 4—

SWE-bench Verified57.60

τ²-Bench—

Free commercial

GPT OSS 20B OpenAI

HLE17.30

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified34.00

τ²-Bench47.70

Free commercial

DeepSeek-V3.1 DeepSeek-AI

HLE15.90

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified66.00

τ²-Bench—

Free commercial

GLM-4.5 智谱AI

HLE14.40

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified64.20

τ²-Bench—

Free commercial

MiniMax M2 MiniMaxAI

HLE12.50

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified69.40

τ²-Bench77.20

Free commercial

GLM-4.5-Air 智谱AI

HLE10.60

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified57.60

τ²-Bench—

Free commercial

Qwen3-32B 阿里巴巴

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Llama 4 Scout Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

MiniMax M3 MiniMaxAI

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Non-commercial

Llama3.1-405B Instruct Facebook AI研究实验室

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

QwQ-32B 阿里巴巴

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Pangu Embedded 华为

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

ERNIE-4.5-300B-A47B 百度

HLE—

ARC-AGI-2—

FrontierMath - Tier 4—

SWE-bench Verified—

τ²-Bench—

Free commercial

Sort by:

Showing 50 of 105 modelsView FrontierMath benchmark page

Leaderboard FAQ

Which open-source models appear on this leaderboard?

The leaderboard tracks open-weight or publicly available models — including Llama, Qwen, DeepSeek, Mistral, GLM, and other releases whose weights or code are available under tracked licenses. It may include permissive, non-commercial, or otherwise restricted licenses; closed-weight API-only models such as GPT or Claude are excluded here.

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

Can I run these open-source models locally?

Most listed models publish weights on Hugging Face or GitHub and can be served via vLLM, Ollama, llama.cpp, or similar runtimes. Hardware requirements scale with parameter count — a 7B model fits on a single consumer GPU, while 65B+ models typically need multi-GPU or quantized deployment.

Explore more

The leaderboard covers benchmarked models. Browse the full catalog by model, organization, or benchmark.

All AI Models

Browse every tracked model — filter by organization, type, and release date, not just benchmark scores.

Browse

All Organizations

Explore the labs and companies behind these models and their full model lineups.

Browse

All Benchmarks

Dive into each benchmark — what it measures, how it scores, and the full ranking.

Browse