DataLearner logoDataLearnerAI
Latest AI Insights
Model Evaluations
Model Directory
Model Comparison
Resource Center
Tools

加载中...

DataLearner logoDataLearner AI

A knowledge platform focused on LLM benchmarking, datasets, and practical instruction with continuously updated capability maps.

产品

  • Leaderboards
  • 模型对比
  • Datasets

资源

  • Tutorials
  • Editorial
  • Tool directory

关于

  • 关于我们
  • 隐私政策
  • 数据收集方法
  • 联系我们

© 2026 DataLearner AI. DataLearner curates industry data and case studies so researchers, enterprises, and developers can rely on trustworthy intelligence.

隐私政策服务条款

LLM Benchmark Performance Comparison

Compare model performance across MMLU Pro, HLE, SWE-Bench and more. Select benchmarks to view rankings.

Detailed benchmark descriptions available at:LLM Benchmark List & Guide

Updated on: 2025/11/08 22:10:24

Benchmark switcher

Pick the leaderboard to sync both chart and table

MMLU ProGPQA DiamondSWE-bench VerifiedMATH-500AIME 2024LiveCodeBench

More benchmark coverage

Browse the benchmark catalog by category and language

More Benchmarks

Filters

Active
All3B and below7B13B34B65B100B and above
AllReasoning ModelsFoundation ModelsInstruction/Chat ModelsCoding Models

LLM Performance Results

Data source: DataLearnerAI
RankModelMMLU ProGPQA DiamondSWE-bench VerifiedMATH-500AIME 2024LiveCodeBenchParams (B)License
1GPT-5.1-Codex-Max0.000.0076.800.000.000.00—不开源
2GPT-5 Codex0.000.0074.500.000.000.00—不开源
3Grok 4 Code0.000.0072.000.000.000.00—不开源
4Grok Code Fast 10.000.0070.800.000.000.00—不开源
5Qwen3-Coder-Next0.000.0070.600.000.000.0080BFree commercial
6GPT-5.1 Codex0.000.0070.400.000.0085.50—不开源
7Qwen3-Coder-480B-A35B0.000.0067.000.000.000.004800BFree commercial
8Devstral Medium0.000.0061.600.000.000.00—不开源
9Devstral Small 1.10.000.0053.600.000.000.00240BFree commercial
10Qwen3-Coder-Flash0.000.0051.600.000.000.00305BFree commercial
11Devstral Small 1.00.000.0046.800.000.000.00240BFree commercial
12Codestral 25.010.000.000.000.000.0037.90—不开源
13Codestral0.000.000.000.000.0031.50220BNon-commercial
1
GPT-5.1-Codex-Max
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified76.80
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
不开源
2
GPT-5 Codex
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified74.50
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
不开源
3
Grok 4 Code
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified72.00
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
不开源
4
Grok Code Fast 1
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified70.80
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
不开源
5
Qwen3-Coder-Next
80B
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified70.60
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
Free commercial
6
GPT-5.1 Codex
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified70.40
MATH-5000.00
AIME 20240.00
LiveCodeBench85.50
不开源
7
Qwen3-Coder-480B-A35B
4800B
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified67.00
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
Free commercial
8
Devstral Medium
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified61.60
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
不开源
9
Devstral Small 1.1
240B
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified53.60
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
Free commercial
10
Qwen3-Coder-Flash
305B
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified51.60
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
Free commercial
11
Devstral Small 1.0
240B
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified46.80
MATH-5000.00
AIME 20240.00
LiveCodeBench0.00
Free commercial
12
Codestral 25.01
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified0.00
MATH-5000.00
AIME 20240.00
LiveCodeBench37.90
不开源
13
Codestral
220B
MMLU Pro0.00
GPQA Diamond0.00
SWE-bench Verified0.00
MATH-5000.00
AIME 20240.00
LiveCodeBench31.50
Non-commercial