DataLearner logo

Gemini 3.0 Pro (Preview 11-2025) Benchmark Details

Gemini 3.0 Pro (Preview 11-2025) currently shows benchmark results led by MMLU Pro (2 / 133, score 90), LiveCodeBench (2 / 127, score 92), GPQA Diamond (12 / 270, score 93.80). 1 source link is attached for reference.

Benchmark Results

Gemini 3.0 Pro (Preview 11-2025)

Benchmark Results

Thinking
Tool usage
Parallel

General Knowledge

11 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Thinking Mode
90
2 / 133
ARC-AGI
Thinking Mode
75
27 / 68
ARC-AGI
Thinking Level · High
87.50
20 / 68
LiveBench
Thinking Level · Low
63.90
54 / 115
LiveBench
Thinking Level · High
73.39
24 / 115
HLE
Thinking Mode
37.50
75 / 185
HLE
Thinking Level · High
37.20
77 / 185
HLE
Thinking Level · HighTools
45.80
44 / 185
HLE
Thinking Level · High
41
66 / 185
ARC-AGI-2
Thinking Mode
31.10
32 / 62
ARC-AGI-2
Thinking Level · High
45.10
26 / 62

Other

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
91.90
29 / 270
GPQA Diamond
Thinking Level · High
91
36 / 270
GPQA Diamond
Thinking Level · High
93.80
12 / 270

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Thinking Mode
72.10
6 / 47

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
92
2 / 127
SWE-bench Verified
Thinking Mode
76.20
36 / 114

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
95
25 / 106
AIME 2026
Thinking Mode
90.60
15 / 19
FrontierMath
Thinking Mode
38
10 / 60
18.80
16 / 80
18.80
16 / 80

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Mode
76.40
5 / 67

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking Level · HighTools
98
5 / 35
τ²-Bench
Thinking ModeTools
85.40
8 / 43
Terminal Bench Hard
Thinking ModeTools
39
5 / 13
Terminal Bench Hard
Thinking Level · HighTools
42
4 / 13

Instruction Following

2 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Mode
70
16 / 34
IF Bench
Thinking Level · HighTools
70
16 / 34

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking Level · HighTools
59.20
38 / 54

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Standard ModeTools
70.30
27 / 40
Terminal Bench 2.0
Thinking ModeTools
54.20
30 / 48
Terminal Bench 2.0
Thinking Level · HighTools
56.90
26 / 48

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Level · High
35
18 / 21

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Level · High
71
11 / 27

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
70.70
32 / 38

Sources