DataLearner logo

Open Source LLM Leaderboard

Track benchmark rankings for open-weight and open-source AI models, then compare score, size, and license signals in one place.

View benchmark detailsUpdated on 2026-09-10 15:15:22

Per-Benchmark Rankings

Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.

Recommended models

Ranked by Aider-Polyglot

LLM Performance Results

Data source: DataLearnerAI

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.

Aider-Polyglot74.20
HLE20.30
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified67.80
τ²-Bench66.70
Free commercial
Aider-Polyglot71.40
HLE17.70
ARC-AGI-21.30
FrontierMath - Tier 4
SWE-bench Verified57.60
τ²-Bench
Free commercial
Aider-Polyglot59.60
HLE7.60
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified34.40
τ²-Bench34.40
Free commercial
Aider-Polyglot59.10
HLE4.70
ARC-AGI-2
FrontierMath - Tier 40.01
SWE-bench Verified51.80
τ²-Bench64.30
Free commercial
Aider-Polyglot56.90
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified49.20
τ²-Bench
Free commercial
Aider-Polyglot55.10
HLE5.20
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified38.80
τ²-Bench38.80
Free commercial
Aider-Polyglot48.40
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot41.80
HLE19.00
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified60.10
τ²-Bench
Free commercial
Aider-Polyglot40.00
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot20.90
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot17.80
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot16.40
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot15.60
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot12.00
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Non-commercial
Aider-Polyglot4.90
HLE
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE9.80
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified22.00
τ²-Bench49.00
Free commercial
Aider-Polyglot
HLE8.90
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified46.40
τ²-Bench
Free commercial
Aider-Polyglot
HLE17.20
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench68.20
Free commercial
Aider-Polyglot
HLE8.40
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified56.00
τ²-Bench
Free commercial
Aider-Polyglot
HLE51.50
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified79.00
τ²-Bench
Free commercial
Aider-Polyglot
HLE48.20
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified80.60
τ²-Bench
Free commercial
Aider-Polyglot
HLE7.20
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified55.60
τ²-Bench
Free commercial
Aider-Polyglot
HLE63.90
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE62.50
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Conditional
Aider-Polyglot
HLE14.40
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified59.20
τ²-Bench79.50
Free commercial
Aider-Polyglot
HLE59.80
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Conditional
Aider-Polyglot
HLE56.20
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Conditional
Aider-Polyglot
HLE55.40
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE55.30
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE54.70
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE54.00
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified80.20
τ²-Bench
Free commercial
Aider-Polyglot
HLE53.20
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified78.00
τ²-Bench
Free commercial
Aider-Polyglot
HLE52.30
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE51.00
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified71.30
τ²-Bench
Free commercial
Aider-Polyglot
HLE50.40
ARC-AGI-24.90
FrontierMath - Tier 42.10
SWE-bench Verified77.80
τ²-Bench89.70
Free commercial
Aider-Polyglot
HLE50.20
ARC-AGI-211.80
FrontierMath - Tier 44.20
SWE-bench Verified76.80
τ²-Bench
Free commercial
Aider-Polyglot
HLE30.40
ARC-AGI-2
FrontierMath - Tier 42.10
SWE-bench Verified68.00
τ²-Bench75.90
Free commercial
Aider-Polyglot
HLE48.50
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified72.40
τ²-Bench79.00
Free commercial
Aider-Polyglot
HLE48.30
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified76.40
τ²-Bench86.70
Free commercial
Aider-Polyglot
HLE47.20
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE46.00
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified77.60
τ²-Bench
Free commercial
Aider-Polyglot
HLE42.80
ARC-AGI-2
FrontierMath - Tier 42.10
SWE-bench Verified73.80
τ²-Bench87.40
Free commercial
Aider-Polyglot
HLE37.40
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified70.70
τ²-Bench
Free commercial
Aider-Polyglot
HLE35.90
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Conditional
Aider-Polyglot
HLE30.80
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE30.60
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Free commercial
Aider-Polyglot
HLE28.00
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench
Non-commercial
Aider-Polyglot
HLE26.50
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified
τ²-Bench76.90
Free commercial
Aider-Polyglot
HLE25.10
ARC-AGI-24.00
FrontierMath - Tier 42.10
SWE-bench Verified73.10
τ²-Bench80.30
Free commercial
Aider-Polyglot
HLE24.00
ARC-AGI-2
FrontierMath - Tier 4
SWE-bench Verified77.20
τ²-Bench
Free commercial
Sort by:
Showing 50 of 127 modelsView Aider-Polyglot benchmark page

Leaderboard FAQ

01

Which open-source models appear on this leaderboard?

The leaderboard tracks open-weight or publicly available models — including Llama, Qwen, DeepSeek, Mistral, GLM, and other releases whose weights or code are available under tracked licenses. It may include permissive, non-commercial, or otherwise restricted licenses; closed-weight API-only models such as GPT or Claude are excluded here.

02

Why do scores for the same model differ across benchmarks?

Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.

03

How often is the leaderboard updated?

Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.

04

How should I read the composite ranking?

The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.

05

Can I run these open-source models locally?

Most listed models publish weights on Hugging Face or GitHub and can be served via vLLM, Ollama, llama.cpp, or similar runtimes. Hardware requirements scale with parameter count — a 7B model fits on a single consumer GPU, while 65B+ models typically need multi-GPU or quantized deployment.

Composite Rankings

There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.

AA Intelligence Index

Full ranking

Composite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.

Updated 2026-09-08

#ModelScore
1
53
2
53
5
51
8
50

LMArena Text Generation

Full ranking

Elo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.

Updated 2026-09-02

#ModelElo
1
1507
4
1502
5
F
Muse Spark 1.2 (xhigh)
1499
6
1498
7
Anthropic
Opus 4.7
1494
8
Google
gemini-3.8-flash-high
1494
9
1493
10
F
Muse Spark 1.1
1492
Source: LMArena

Leading model developers

View all 101 organizations

Jump to a developer to explore its full model lineup, series, and product lines.

Today's picksRotates daily · discover more labs

Model comparisons

Head-to-head write-ups: what the benchmark gap actually means, and which model fits which job.

All comparisons

Explore more

The leaderboard covers benchmarked models. Browse the full catalog by model, organization, or benchmark.