Grok 4 Heavy is an AI model published by xAI, released on 2025-07-10, for Chat model, and 128K context length, with a 100.00 score on AIME2025.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
Model basics
Open source & experience
Official resources
API details
Benchmark Results
Grok 4 Heavy currently shows benchmark results led by AIME2025 (1 / 106, score 100), GPQA Diamond (57 / 270, score 88.90), HLE (49 / 185, score 44.40). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.
Coding and Software Engineer
1 evaluationsMath and Reasoning
2 evaluationsPublisher
Model Overview
Grok 4 Heavy is an AI model published by xAI, released on 2025-07-10, for Chat model, and 128K context length, with a 100.00 score on AIME2025.
DataLearner on WeChat
Follow DataLearner on WeChat for AI model updates and research notes.
