DataLearner logo

Gemini 3.0 Pro (Preview 11-2025) Benchmark Details

Gemini 3.0 Pro (Preview 11-2025) currently shows benchmark results led by MMLU Pro (2 / 133, score 90), LiveCodeBench (2 / 127, score 92), GPQA Diamond (13 / 271, score 93.80). 1 source link is attached for reference.

Benchmark Results

Gemini 3.0 Pro (Preview 11-2025)

Benchmark Results

Thinking
Tool usage
Parallel

General Knowledge

11 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Thinking Mode
90
2 / 133
ARC-AGI-1
Thinking Mode
75
48 / 91
ARC-AGI-1
Thinking Level · High
87.50
37 / 91
LiveBench
Thinking Level · Low
63.90
54 / 115
LiveBench
Thinking Level · High
73.39
24 / 115
HLE
Thinking Mode
37.50
79 / 190
HLE
Thinking Level · High
37.20
81 / 190
HLE
Thinking Level · HighTools
45.80
48 / 190
HLE
Thinking Level · High
41
70 / 190
ARC-AGI-2
Thinking Mode
31.10
54 / 85
ARC-AGI-2
Thinking Level · High
45.10
47 / 85

Other

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
91.90
30 / 271
GPQA Diamond
Thinking Level · High
91
37 / 271
GPQA Diamond
Thinking Level · High
93.80
13 / 271

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Thinking Mode
72.10
6 / 47

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
92
2 / 127
SWE-bench Verified
Thinking Mode
76.20
36 / 115

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
95
25 / 106
AIME 2026
Thinking Mode
90.60
16 / 20
FrontierMath
Thinking Mode
38
10 / 60
18.80
16 / 80
18.80
16 / 80

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Mode
76.40
9 / 92

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking Level · HighTools
98
5 / 35
τ²-Bench
Thinking ModeTools
85.40
8 / 44
Terminal Bench Hard
Thinking ModeTools
39
5 / 13
Terminal Bench Hard
Thinking Level · HighTools
42
4 / 13

Instruction Following

2 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Mode
70
17 / 35
IF Bench
Thinking Level · HighTools
70
17 / 35

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking Level · HighTools
59.20
39 / 56

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Standard ModeTools
70.30
27 / 41
Terminal Bench 2.0
Thinking ModeTools
54.20
30 / 48
Terminal Bench 2.0
Thinking Level · HighTools
56.90
26 / 48

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Level · High
35
18 / 21

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Level · High
71
11 / 28

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
70.70
32 / 38

Sources