Gemini 3.1 Pro Preview Benchmark Details
Gemini 3.1 Pro Preview currently shows benchmark results led by GPQA Diamond (4 / 226, score 94.30), LiveCodeBench (3 / 126, score 91.70), LiveBench (3 / 115, score 79.93). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.
Benchmark Results
Benchmark Results
General Knowledge
6 evaluationsCoding and Software Engineer
8 evaluationsMultimodal Understanding
3 evaluationsAgent Level Benchmark
4 evaluationsMath and Reasoning
3 evaluationsAI Agent - Tool Usage
5 evaluationsOther
2 evaluationsCompetitor Comparison
Benchmark scores for Gemini 3.1 Pro Preview compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Gemini 3.1 Pro PreviewCurrent | Claude Opus 4.6 |
|---|---|---|
ARC-AGI-2 综合评估 | 77.10Thinking Level · High | 66.30Extended Thinking |
HLE 综合评估 | 51.40Thinking Level · High | Tools | 53.00Extended Thinking | Tools |
LiveBench 综合评估 | 79.93Thinking Level · High | 76.33Thinking Level · High |
MMLU 综合评估 | 92.60Thinking Level · High | 91.05Extended Thinking |
GPQA Diamond 科学与综合推理 | 94.30Thinking Level · High | 91.31Extended Thinking |
GSO 编程与软件工程 | 22.55Standard Mode | Tools | 41.20Thinking Level · High | Tools |
LiveCodeBench 编程与软件工程 | 91.70Thinking Level · High | Tools | 76.00Extended Thinking |
SWE-Bench Pro - Commercial 编程与软件工程 | 32.20Thinking Enabled | Tools | 47.10Thinking Enabled | Tools |
SWE-bench Verified 编程与软件工程 | 80.60Thinking Level · High | Tools | 80.84Extended Thinking | Tools |
Text Arena (Coding) 编程与软件工程 | 1461.49Standard Mode | 1555.35Standard Mode |
WeirdML v2 编程与软件工程 | 72.10Standard Mode | Tools | 77.95Thinking Level · High | Tools |
MMMU 多模态理解 | 80.50Thinking Level · High | 77.30Extended Thinking | Tools |
Standard API Pricing: Gemini 3.1 Pro Preview vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Gemini 3.1 Pro Preview | Google Deep Mind | $2 / 1M tokens | $12 / 1M tokens | <= 200K |
Claude Opus 4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | <= 200K |
Version History
How each version of the Gemini 3.1 Pro Preview series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Gemini 3.1 Pro PreviewCurrent | Gemini 3.0 Pro (Preview 11-2025) | Gemini 2.5-Pro | Gemini 2.5 Pro Experimental 03-25 |
|---|---|---|---|---|
ARC-AGI-2 综合评估 | 77.10Thinking Level · High | 45.10Thinking Enabled | 4.90Thinking Enabled | -- |
HLE 综合评估 | 51.40Thinking Level · High | Tools | 45.80Thinking Level · High | Tools | 21.60Thinking Enabled | 18.80Standard Mode |
LiveBench 综合评估 | 79.93Thinking Level · High | 73.39Thinking Level · High | 58.33Thinking Level · High | -- |
GPQA Diamond 科学与综合推理 | 94.30Thinking Level · High | 93.80Thinking Enabled | 86.40Thinking Enabled | 84.00Standard Mode |
GSO 编程与软件工程 | 22.55Standard Mode | Tools | -- | 3.92Standard Mode | Tools | -- |
LiveCodeBench 编程与软件工程 | 91.70Thinking Level · High | Tools | 92.00Thinking Enabled | 77.10Standard Mode | 70.40Standard Mode |
SWE-bench Verified 编程与软件工程 | 80.60Thinking Level · High | Tools | 76.20Thinking Enabled | 67.20Thinking Enabled | 63.80Standard Mode |
MMMU 多模态理解 | 80.50Thinking Level · High | -- | 82.00Thinking Enabled | -- |
SimpleBench 常识推理 | 79.60Standard Mode | 76.40Thinking Enabled | 62.40Thinking Enabled | 51.60Standard Mode |
BALROG Agent能力评测 | 57.00Standard Mode | Tools | -- | -- | 43.30Standard Mode | Tools |
METR Time Horizons v1.1 Agent能力评测 | 384.15Standard Mode | Tools | -- | 38.73Standard Mode | Tools | -- |
τ²-Bench Agent能力评测 | 90.80Thinking Level · High | Tools | 85.40Thinking Enabled | Tools | -- | -- |
Single-Benchmark Version Trend
Viewing: ARC-AGI-2 · 综合评估
Standard API Pricing Across the Gemini 3.1 Pro Preview Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Gemini 3.1 Pro Preview | Google Deep Mind | $2 / 1M tokens | $12 / 1M tokens | <= 200K |
Gemini 3.0 Pro (Preview 11-2025) | Google Deep Mind | $2 / 1M tokens | $12 / 1M tokens | <= 200000 |
Gemini 2.5-Pro | Google Deep Mind | $1.25 / 1M tokens | $10 / 1M tokens | <= 200000 |