LLM Math Reasoning Benchmark Leaderboard
This page provides the most comprehensive LLM math reasoning benchmark leaderboard. We evaluate models including GPT, Claude, Qwen, and DeepSeek using authoritative math benchmarks such as AIME 2025, FrontierMath-Tier4, MATH-500, and GSM8K.
Updated on 2026-07-18 08:01:51
As of 2026-07, this page covers AIME2025, FrontierMath - Tier 4, MATH-500, GSM8K and related benchmarks for LLM Math Reasoning Benchmark Leaderboard, making it straightforward to compare within the same task family.
Click any model name to check context length, licensing, and pricing on its detail page. See Data Methodology for scoring details.
Model release cutoff:
Top picks
Ranked by FrontierMath - Tier 4LLM Performance Results
Data source: DataLearnerAIClick any row to open the model page. Tick the checkboxes to compare up to 4 models side by side.
AIME2025—
FrontierMath - Tier 410.40
MATH-500—
GSM8K—
Proprietary
AIME202583.00
FrontierMath - Tier 42.10
MATH-50098.80
GSM8K—
Proprietary
AIME202592.30
FrontierMath - Tier 4—
MATH-500—
GSM8K—
Free commercial
AIME202590.00
FrontierMath - Tier 4—
MATH-500—
GSM8K—
Free commercial
AIME202563.10
FrontierMath - Tier 4—
MATH-500—
GSM8K—
Proprietary
AIME202547.70
FrontierMath - Tier 4—
MATH-50094.00
GSM8K96.30
Free commercial
AIME202529.70
FrontierMath - Tier 4—
MATH-500—
GSM8K—
Proprietary
AIME2025—
FrontierMath - Tier 4—
MATH-500—
GSM8K—
Free commercial
Sort by:
Showing 50 of 55 modelsView FrontierMath - Tier 4 benchmark page














