DataLearner logo
Back to Main Leaderboard

LLM Math Reasoning Benchmark Leaderboard

This page provides the most comprehensive LLM math reasoning benchmark leaderboard. We evaluate models including GPT, Claude, Qwen, and DeepSeek using authoritative math benchmarks such as AIME 2025, FrontierMath-Tier4, MATH-500, and GSM8K.

Updated on 2026-07-18 08:01:51

As of 2026-07, this page covers AIME2025, FrontierMath - Tier 4, MATH-500, GSM8K and related benchmarks for LLM Math Reasoning Benchmark Leaderboard, making it straightforward to compare within the same task family.

Click any model name to check context length, licensing, and pricing on its detail page. See Data Methodology for scoring details.

Top picks

Ranked by AIME2025

LLM Performance Results

Data source: DataLearnerAI

Click any row to open the model page. Tick the checkboxes to compare up to 4 models side by side.

AIME202581.30
FrontierMath - Tier 4
MATH-500
GSM8K
Free commercial
AIME202575.30
FrontierMath - Tier 4
MATH-50093.70
GSM8K
Free commercial
AIME202567.30
FrontierMath - Tier 4
MATH-50097.40
GSM8K
Free commercial
AIME202547.40
FrontierMath - Tier 4
MATH-500
GSM8K
Free commercial
AIME2025
FrontierMath - Tier 4
MATH-50092.40
GSM8K95.98
Free commercial
AIME2025
FrontierMath - Tier 4
MATH-500
GSM8K85.40
Free commercial
AIME2025
FrontierMath - Tier 4
MATH-500
GSM8K82.40
Free commercial
AIME2025
FrontierMath - Tier 4
MATH-500
GSM8K70.70
Free commercial
AIME2025
FrontierMath - Tier 4
MATH-500
GSM8K55.30
Free commercial
AIME2025
FrontierMath - Tier 4
MATH-500
GSM8K36.20
Free commercial
AIME2025
FrontierMath - Tier 4
MATH-50091.40
GSM8K
Free commercial
Sort by: