Gemini 3.1 Pro Preview Benchmark Details
Gemini 3.1 Pro Preview currently shows benchmark results led by GPQA Diamond (3 / 187, score 94.30), LiveCodeBench (3 / 123, score 91.70), LiveBench (3 / 115, score 79.93). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.
Benchmark Results
Benchmark Results
General Knowledge
7 evaluationsCoding and Software Engineer
4 evaluationsMultimodal Understanding
3 evaluationsAgent Level Benchmark
2 evaluationsMath and Reasoning
3 evaluationsAI Agent - Tool Usage
5 evaluationsOther
2 evaluationsCompetitor Comparison
Benchmark scores for Gemini 3.1 Pro Preview compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Gemini 3.1 Pro PreviewCurrent | Claude Opus 4.6 |
|---|---|---|
ARC-AGI-2 综合评估 | 77.10Thinking Level · High | 66.30Extended Thinking |
GPQA Diamond 综合评估 | 94.30Thinking Level · High | 91.31Extended Thinking |
HLE 综合评估 | 51.40Thinking Level · High | Tools | 53.00Extended Thinking | Tools |
LiveBench 综合评估 | 79.93Thinking Level · High | 76.33Thinking Level · High |
MMLU 综合评估 | 92.60Thinking Level · High | 91.05Extended Thinking |
LiveCodeBench 编程与软件工程 | 91.70Thinking Level · High | Tools | 76.00Extended Thinking |
SWE-bench Verified 编程与软件工程 | 80.60Thinking Level · High | Tools | 80.84Extended Thinking | Tools |
MMMU 多模态理解 | 80.50Thinking Level · High | 77.30Extended Thinking | Tools |
Simple Bench 常识推理 | 79.60Standard Mode | 67.60Standard Mode |
τ²-Bench Agent能力评测 | 90.80Thinking Level · High | Tools | 91.89Extended Thinking | Tools |
τ²-Bench - Telecom Agent能力评测 | 99.30Thinking Level · High | Tools | 99.25Extended Thinking | Tools |
FrontierMath 数学推理 | 36.90Thinking Level · High | 40.70Thinking Level · High |
Standard API Pricing: Gemini 3.1 Pro Preview vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Gemini 3.1 Pro Preview | Google Deep Mind | $2 / 1M tokens | $12 / 1M tokens | <= 200K |
Claude Opus 4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | <= 200K |
Version History
How each version of the Gemini 3.1 Pro Preview series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Gemini 3.1 Pro PreviewCurrent | Gemini 3.0 Pro (Preview 11-2025) | Gemini 2.5-Pro | Gemini 2.5 Pro Experimental 03-25 |
|---|---|---|---|---|
ARC-AGI-2 综合评估 | 77.10Thinking Level · High | 45.10Thinking Enabled | 4.90Thinking Enabled | -- |
GPQA Diamond 综合评估 | 94.30Thinking Level · High | 93.80Thinking Enabled | 86.40Thinking Enabled | 84.00Standard Mode |
HLE 综合评估 | 51.40Thinking Level · High | Tools | 45.80Thinking Level · High | Tools | 21.60Thinking Enabled | 18.80Standard Mode |
LiveBench 综合评估 | 79.93Thinking Level · High | 73.39Thinking Level · High | 58.33Thinking Level · High | -- |
LiveCodeBench 编程与软件工程 | 91.70Thinking Level · High | Tools | 92.00Thinking Enabled | 77.10Standard Mode | 70.40Standard Mode |
SWE-bench Verified 编程与软件工程 | 80.60Thinking Level · High | Tools | 76.20Thinking Enabled | 67.20Thinking Enabled | 63.80Standard Mode |
MMMU 多模态理解 | 80.50Thinking Level · High | -- | 82.00Thinking Enabled | -- |
Simple Bench 常识推理 | 79.60Standard Mode | 76.40Thinking Enabled | 62.40Thinking Enabled | 51.60Standard Mode |
τ²-Bench Agent能力评测 | 90.80Thinking Level · High | Tools | 85.40Thinking Enabled | Tools | -- | -- |
τ²-Bench - Telecom Agent能力评测 | 99.30Thinking Level · High | Tools | 98.00Thinking Level · High | Tools | 54.00Thinking Enabled | Tools | -- |
FrontierMath 数学推理 | 36.90Thinking Level · High | 38.00Thinking Enabled | 11.00Standard Mode | -- |
16.70Standard Mode | 18.80Thinking Enabled | 2.10Standard Mode | 4.20Standard Mode |
Single-Benchmark Version Trend
Viewing: ARC-AGI-2 · 综合评估
Standard API Pricing Across the Gemini 3.1 Pro Preview Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Gemini 3.1 Pro Preview | Google Deep Mind | $2 / 1M tokens | $12 / 1M tokens | <= 200K |
Gemini 3.0 Pro (Preview 11-2025) | Google Deep Mind | $2 / 1M tokens | $12 / 1M tokens | <= 200000 |
Gemini 2.5-Pro | Google Deep Mind | $1.25 / 1M tokens | $10 / 1M tokens | <= 200000 |