DataLearner logo

Gemini 3.1 Pro Preview Benchmark Details

Gemini 3.1 Pro Preview currently shows benchmark results led by GPQA Diamond (4 / 226, score 94.30), LiveCodeBench (3 / 126, score 91.70), LiveBench (3 / 115, score 79.93). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Gemini 3.1 Pro Preview

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
MMLU
High
92.60
3 / 66
79.93
3 / 115
77.10
9 / 62
HLE
High
44.40
47 / 181
HLE
HighTools
51.40
24 / 181
0
6 / 9

Other

1 evaluations
Benchmark / mode
Score
Rank/total
94.30
4 / 226

Coding and Software Engineer

8 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1461.49
23 / 35
LiveCodeBench
HighTools
91.70
3 / 126
80.60
11 / 114
WeirdML v2
Standard ModeTools
72.10
17 / 52
54.20
34 / 57
SWE-Bench Pro - Commercial
Thinking EnabledTools
32.20
3 / 3
GSO
Standard ModeTools
22.55
10 / 21
DeepSWE
HighTools
12
27 / 27

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
CharXiv RQ
Thinking Enabled
83.30
13 / 15
CharXiv RQ
Thinking EnabledTools
83.20
14 / 15
MMMU
High
80.50
12 / 29

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
79.60
2 / 67

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Standard ModeTools
384.15
2 / 22
99.30
1 / 35
τ²-Bench
HighTools
90.80
2 / 43
BALROG
Standard ModeTools
57
2 / 12

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
36.90
11 / 60
16.70
20 / 80
16.70
20 / 80

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighToolsInternet
85.90
5 / 54

AI Agent - Tool Usage

5 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
78.20
12 / 38
OSWorld-Verified
Thinking EnabledTools
76.20
12 / 26
73.80
28 / 44
68.50
8 / 48
MLE-Bench
Thinking EnabledTools
42.60
3 / 3

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
86.70
11 / 38

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Enabled
965
12 / 13

Other

2 evaluations
Benchmark / mode
Score
Rank/total
84.90
3 / 8
26.30
3 / 3

Competitor Comparison

Benchmark scores for Gemini 3.1 Pro Preview compared against top models in its class

Gemini 3.1 Pro PreviewClaude Opus 4.6GPT-5.3
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.1 Pro PreviewCurrentClaude Opus 4.6
ARC-AGI-2
综合评估
77.10Thinking Level · High
66.30Extended Thinking
HLE
综合评估
51.40Thinking Level · High | Tools
53.00Extended Thinking | Tools
LiveBench
综合评估
79.93Thinking Level · High
76.33Thinking Level · High
MMLU
综合评估
92.60Thinking Level · High
91.05Extended Thinking
GPQA Diamond
科学与综合推理
94.30Thinking Level · High
91.31Extended Thinking
GSO
编程与软件工程
22.55Standard Mode | Tools
41.20Thinking Level · High | Tools
LiveCodeBench
编程与软件工程
91.70Thinking Level · High | Tools
76.00Extended Thinking
SWE-Bench Pro - Commercial
编程与软件工程
32.20Thinking Enabled | Tools
47.10Thinking Enabled | Tools
SWE-bench Verified
编程与软件工程
80.60Thinking Level · High | Tools
80.84Extended Thinking | Tools
Text Arena (Coding)
编程与软件工程
1461.49Standard Mode
1555.35Standard Mode
WeirdML v2
编程与软件工程
72.10Standard Mode | Tools
77.95Thinking Level · High | Tools
MMMU
多模态理解
80.50Thinking Level · High
77.30Extended Thinking | Tools
10 additional benchmarks remain in the chart above.

Standard API Pricing: Gemini 3.1 Pro Preview vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
Claude Opus 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K

Version History

How each version of the Gemini 3.1 Pro Preview series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.1 Pro PreviewCurrentGemini 3.0 Pro (Preview 11-2025)Gemini 2.5-ProGemini 2.5 Pro Experimental 03-25
ARC-AGI-2
综合评估
77.10Thinking Level · High
45.10Thinking Enabled
4.90Thinking Enabled
--
HLE
综合评估
51.40Thinking Level · High | Tools
45.80Thinking Level · High | Tools
21.60Thinking Enabled
18.80Standard Mode
LiveBench
综合评估
79.93Thinking Level · High
73.39Thinking Level · High
58.33Thinking Level · High
--
GPQA Diamond
科学与综合推理
94.30Thinking Level · High
93.80Thinking Enabled
86.40Thinking Enabled
84.00Standard Mode
GSO
编程与软件工程
22.55Standard Mode | Tools
--
3.92Standard Mode | Tools
--
LiveCodeBench
编程与软件工程
91.70Thinking Level · High | Tools
92.00Thinking Enabled
77.10Standard Mode
70.40Standard Mode
SWE-bench Verified
编程与软件工程
80.60Thinking Level · High | Tools
76.20Thinking Enabled
67.20Thinking Enabled
63.80Standard Mode
MMMU
多模态理解
80.50Thinking Level · High
--
82.00Thinking Enabled
--
SimpleBench
常识推理
79.60Standard Mode
76.40Thinking Enabled
62.40Thinking Enabled
51.60Standard Mode
BALROG
Agent能力评测
57.00Standard Mode | Tools
--
--
43.30Standard Mode | Tools
METR Time Horizons v1.1
Agent能力评测
384.15Standard Mode | Tools
--
38.73Standard Mode | Tools
--
τ²-Bench
Agent能力评测
90.80Thinking Level · High | Tools
85.40Thinking Enabled | Tools
--
--
7 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-2 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.1 Pro Preview Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
Gemini 3.0 Pro (Preview 11-2025): Base price applies to <= 200000
Gemini 2.5-Pro: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K
Gemini 3.0 Pro (Preview 11-2025)
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200000
Gemini 2.5-Pro
Google Deep Mind$1.25 / 1M tokens$10 / 1M tokens<= 200000

Sources