DataLearner logo

Gemini 3.0 Flash Benchmark Details

Gemini 3.0 Flash currently shows benchmark results led by τ²-Bench (3 / 43, score 90.20), AIME2025 (8 / 106, score 99.70), GeoBench ACW (2 / 20, score 88). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Gemini 3.0 Flash

Benchmark Results

Thinking
Tool usage

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
56.35
79 / 115
LiveBench
Thinking Level · High
72.40
26 / 115
HLE
Thinking Mode
33.70
84 / 181
HLE
Thinking ModeTools
43.50
50 / 181
ARC-AGI-2
Thinking Mode
33.60
30 / 62

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
90.40
36 / 224

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Thinking Mode
68.70
8 / 47

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1454
25 / 35
SWE-bench Verified
Thinking Mode
68.70
68 / 114
WeirdML v2
Standard ModeTools
61.60
26 / 52
SWE-Bench Pro - Public
Thinking Level · HighTools
49.60
45 / 57
GSO
Standard ModeTools
9.80
14 / 21

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
95.20
24 / 106
AIME2025
Thinking ModeTools
99.70
8 / 106
4.20
40 / 80

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
61.10
17 / 67

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
90.20
3 / 43
BALROG
Standard ModeTools
48.10
3 / 12

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Standard ModeTools
62
32 / 39
Terminal-Bench 2.1
Thinking Level · HighTools
58
41 / 45
Terminal Bench 2.0
Thinking ModeTools
47.60
39 / 48

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking ModeTools
85.70
15 / 29
Pinch Bench
Thinking ModeTools
85.20
17 / 38

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
GeoBench ACW
Standard Mode
88
2 / 20
VPCT
Standard Mode
72.60
3 / 24

Competitor Comparison

Benchmark scores for Gemini 3.0 Flash compared against top models in its class

Gemini 3.0 FlashClaude Sonnet 4GPT-5.3
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.0 FlashCurrentClaude Sonnet 4
ARC-AGI-2
综合评估
33.60Thinking Enabled
5.90Thinking Enabled
HLE
综合评估
43.50Thinking Enabled | Tools
9.60Thinking Enabled
LiveBench
综合评估
72.40Thinking Level · High
61.2764K
GPQA Diamond
科学与综合推理
90.40Thinking Enabled
83.80Deep Thinking Mode | Tools
SWE-Bench Pro - Public
编程与软件工程
49.60Thinking Level · High | Tools
42.70Thinking Enabled
SWE-bench Verified
编程与软件工程
68.70Thinking Enabled
80.20Thinking Enabled | Tools
AIME2025
数学推理
99.70Thinking Enabled | Tools
85.00Deep Thinking Mode | Tools
SimpleBench
常识推理
61.10Standard Mode
45.50Thinking Enabled
τ²-Bench
Agent能力评测
90.20Thinking Enabled | Tools
52.00Standard Mode | Tools
Claw Bench
OpenClaw智能体能力综合测评
85.70Thinking Enabled | Tools
77.80Thinking Enabled | Tools
Pinch Bench
OpenClaw智能体能力综合测评
85.20Thinking Enabled | Tools
80.50Thinking Enabled | Tools

Standard API Pricing: Gemini 3.0 Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.0 Flash
Google Deep Mind$0.5 / 1M tokens$3 / 1M tokens
Claude Sonnet 4
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200000

Version History

How each version of the Gemini 3.0 Flash series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.0 FlashCurrentGemini 2.5 FlashGemini 2.0 Flash Experimental
HLE
综合评估
43.50Thinking Enabled | Tools
11.00Thinking Enabled
5.10Standard Mode
LiveBench
综合评估
72.40Thinking Level · High
47.74Thinking Level · High
--
GPQA Diamond
科学与综合推理
90.40Thinking Enabled
82.80Thinking Enabled
65.20Standard Mode
SimpleQA
常识问答
68.70Thinking Enabled
26.90Thinking Enabled
29.90Standard Mode
SWE-bench Verified
编程与软件工程
68.70Thinking Enabled
50.00Standard Mode
21.40Standard Mode
WeirdML v2
编程与软件工程
61.60Standard Mode | Tools
40.9516K | Tools
--
AIME2025
数学推理
99.70Thinking Enabled | Tools
72.00Thinking Enabled
29.70Standard Mode
4.20Standard Mode
4.20Standard Mode
--
SimpleBench
常识推理
61.10Standard Mode
41.20Standard Mode
18.90Standard Mode
BALROG
Agent能力评测
48.10Standard Mode | Tools
33.50Standard Mode | Tools
--
Pinch Bench
OpenClaw智能体能力综合测评
85.20Thinking Enabled | Tools
70.70Thinking Enabled | Tools
--
GeoBench ACW
多模态理解
88.00Standard Mode
76.00Standard Mode
--

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.0 Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.0 Flash
Google Deep Mind$0.5 / 1M tokens$3 / 1M tokens
Gemini 2.5 Flash
Google Deep Mind$0.3 / 1M tokens$2.5 / 1M tokens
Gemini 2.0 Flash Experimental
Google Deep Mind$0.1 / 1M tokens$0.4 / 1M tokens

Sources