DataLearner logo

Gemini 3.1 Pro Preview Benchmark Details

Gemini 3.1 Pro Preview currently shows benchmark results led by GPQA Diamond (3 / 187, score 94.30), LiveCodeBench (3 / 123, score 91.70), LiveBench (3 / 115, score 79.93). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Gemini 3.1 Pro Preview

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

7 evaluations
Benchmark / mode
Score
Rank/total
94.30
3 / 187
MMLU
High
92.60
3 / 66
79.93
3 / 115
77.10
9 / 62
HLE
High
44.40
45 / 172
HLE
HighTools
51.40
22 / 172
0
6 / 9

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
HighTools
91.70
3 / 123
80.60
11 / 112
54.20
32 / 54
DeepSWE
HighTools
12
20 / 20

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
CharXiv RQ
Thinking Enabled
83.30
7 / 8
CharXiv RQ
Thinking EnabledTools
83.20
8 / 8
MMMU
High
80.50
12 / 29

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
79.60
2 / 63

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
99.30
1 / 35
τ²-Bench
HighTools
90.80
2 / 43

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
36.90
11 / 60
16.70
20 / 80
16.70
20 / 80

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighToolsInternet
85.90
5 / 53

AI Agent - Tool Usage

5 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
78.20
9 / 27
OSWorld-Verified
Thinking EnabledTools
76.20
11 / 24
73.80
17 / 28
68.50
8 / 47
MLE-Bench
Thinking EnabledTools
42.60
3 / 3

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
86.70
10 / 37

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Enabled
965
5 / 5

Other

2 evaluations
Benchmark / mode
Score
Rank/total
84.90
2 / 3
26.30
3 / 3

Competitor Comparison

Benchmark scores for Gemini 3.1 Pro Preview compared against top models in its class

Gemini 3.1 Pro PreviewClaude Opus 4.6GPT-5.3
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.1 Pro PreviewCurrentClaude Opus 4.6
ARC-AGI-2
综合评估
77.10Thinking Level · High
66.30Extended Thinking
GPQA Diamond
综合评估
94.30Thinking Level · High
91.31Extended Thinking
HLE
综合评估
51.40Thinking Level · High | Tools
53.00Extended Thinking | Tools
LiveBench
综合评估
79.93Thinking Level · High
76.33Thinking Level · High
MMLU
综合评估
92.60Thinking Level · High
91.05Extended Thinking
LiveCodeBench
编程与软件工程
91.70Thinking Level · High | Tools
76.00Extended Thinking
SWE-bench Verified
编程与软件工程
80.60Thinking Level · High | Tools
80.84Extended Thinking | Tools
MMMU
多模态理解
80.50Thinking Level · High
77.30Extended Thinking | Tools
Simple Bench
常识推理
79.60Standard Mode
67.60Standard Mode
τ²-Bench
Agent能力评测
90.80Thinking Level · High | Tools
91.89Extended Thinking | Tools
τ²-Bench - Telecom
Agent能力评测
99.30Thinking Level · High | Tools
99.25Extended Thinking | Tools
FrontierMath
数学推理
36.90Thinking Level · High
40.70Thinking Level · High
6 additional benchmarks remain in the chart above.

Standard API Pricing: Gemini 3.1 Pro Preview vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
Claude Opus 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K

Version History

How each version of the Gemini 3.1 Pro Preview series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.1 Pro PreviewCurrentGemini 3.0 Pro (Preview 11-2025)Gemini 2.5-ProGemini 2.5 Pro Experimental 03-25
ARC-AGI-2
综合评估
77.10Thinking Level · High
45.10Thinking Enabled
4.90Thinking Enabled
--
GPQA Diamond
综合评估
94.30Thinking Level · High
93.80Thinking Enabled
86.40Thinking Enabled
84.00Standard Mode
HLE
综合评估
51.40Thinking Level · High | Tools
45.80Thinking Level · High | Tools
21.60Thinking Enabled
18.80Standard Mode
LiveBench
综合评估
79.93Thinking Level · High
73.39Thinking Level · High
58.33Thinking Level · High
--
LiveCodeBench
编程与软件工程
91.70Thinking Level · High | Tools
92.00Thinking Enabled
77.10Standard Mode
70.40Standard Mode
SWE-bench Verified
编程与软件工程
80.60Thinking Level · High | Tools
76.20Thinking Enabled
67.20Thinking Enabled
63.80Standard Mode
MMMU
多模态理解
80.50Thinking Level · High
--
82.00Thinking Enabled
--
Simple Bench
常识推理
79.60Standard Mode
76.40Thinking Enabled
62.40Thinking Enabled
51.60Standard Mode
τ²-Bench
Agent能力评测
90.80Thinking Level · High | Tools
85.40Thinking Enabled | Tools
--
--
τ²-Bench - Telecom
Agent能力评测
99.30Thinking Level · High | Tools
98.00Thinking Level · High | Tools
54.00Thinking Enabled | Tools
--
FrontierMath
数学推理
36.90Thinking Level · High
38.00Thinking Enabled
11.00Standard Mode
--
16.70Standard Mode
18.80Thinking Enabled
2.10Standard Mode
4.20Standard Mode
4 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-2 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.1 Pro Preview Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
Gemini 2.5-Pro: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K
Gemini 3.0 Pro (Preview 11-2025)
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200000
Gemini 2.5-Pro
Google Deep Mind$1.25 / 1M tokens$10 / 1M tokens<= 200000

Sources