DataLearner logo

Gemini 3.5 Flash Benchmark Details

Gemini 3.5 Flash currently shows benchmark results led by SimpleBench (4 / 67, score 76.70), GPQA Diamond (19 / 226, score 92.80), MCP-Atlas (4 / 38, score 83.60). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Gemini 3.5 Flash

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
75.02
17 / 115
ARC-AGI-2
HighTools
72.10
13 / 62
HLE
HighTools
40.20
66 / 181

Other

2 evaluations
Benchmark / mode
Score
Rank/total
88.89
50 / 226
92.80
19 / 226

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
76.70
4 / 67

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
55.10
30 / 57
DeepSWE
Thinking Level · MediumTools
37
22 / 26

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
83.60
4 / 38
78.40
10 / 26
76.20
22 / 43
MLE-Bench
Thinking EnabledTools
49.70
2 / 3

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Enabled
1349
11 / 13

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
CharXiv RQ
Thinking Enabled
84.20
11 / 15
CharXiv RQ
Thinking EnabledTools
84.90
8 / 15

Other

2 evaluations
Benchmark / mode
Score
Rank/total
77.30
4 / 8
26.60
2 / 3

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
62.81
18 / 34
26.83
18 / 34

Competitor Comparison

Benchmark scores for Gemini 3.5 Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.5 FlashCurrentClaude Sonnet 4.6Opus 4.7GPT-5.5
ARC-AGI-2
综合评估
72.10Thinking Level · High | Tools
58.30Thinking Enabled
75.80Thinking Level · High
85.00Thinking Level · Extra High
HLE
综合评估
40.20Thinking Level · High | Tools
49.00Thinking Enabled | Tools
54.70Extended Thinking | Tools
52.20Thinking Level · High | Tools
LiveBench
综合评估
75.02Thinking Level · High
75.47Thinking Level · Medium
76.91Deep Thinking Mode
80.71Deep Thinking Mode
GPQA Diamond
科学与综合推理
92.80Thinking Level · High
89.90Thinking Enabled
94.20Extended Thinking
94.00Thinking Level · Extra High
SimpleBench
常识推理
76.70Standard Mode
--
62.90Standard Mode
69.00Standard Mode
DeepSWE
编程与软件工程
37.00Thinking Level · Medium | Tools
30.00Thinking Level · High | Tools
--
67.00Thinking Level · Extra High | Tools
SWE-Bench Pro - Public
编程与软件工程
55.10Thinking Level · High | Tools
--
64.30Extended Thinking | Tools
58.60Thinking Level · High | Tools
MCP-Atlas
AI Agent - 工具使用
83.60Thinking Level · High | Tools
69.50Standard Mode | Tools
79.10Thinking Level · High | Tools
75.30Thinking Level · Extra High | Tools
OSWorld-Verified
AI Agent - 工具使用
78.40Thinking Level · High | Tools
72.50Thinking Enabled | Tools
78.00Extended Thinking | Tools
78.70Thinking Level · High | Tools
Terminal-Bench 2.1
AI Agent - 工具使用
76.20Thinking Level · High | Tools
--
69.70Thinking Level · High | Tools
83.40Thinking Level · High | Tools
26.83Thinking Level · High
--
31.71Thinking Level · High
72.50Thinking Level · Extra High
FrontierMath v2
数学推理
62.81Thinking Level · High
--
70.18Thinking Level · High
85.26Thinking Level · Extra High

Standard API Pricing: Gemini 3.5 Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Sonnet 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.5 Flash
DeepMind$1.5 / 1M tokens$9 / 1M tokens
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens<= 200K
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens

Version History

How each version of the Gemini 3.5 Flash series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.5 FlashCurrentGemini 3.0 FlashGemini 2.5 Flash
ARC-AGI-2
综合评估
72.10Thinking Level · High | Tools
33.60Thinking Enabled
--
HLE
综合评估
40.20Thinking Level · High | Tools
43.50Thinking Enabled | Tools
11.00Thinking Enabled
LiveBench
综合评估
75.02Thinking Level · High
72.40Thinking Level · High
47.74Thinking Level · High
GPQA Diamond
科学与综合推理
92.80Thinking Level · High
90.40Thinking Enabled
82.80Thinking Enabled
SimpleBench
常识推理
76.70Standard Mode
61.10Standard Mode
41.20Standard Mode
SWE-Bench Pro - Public
编程与软件工程
55.10Thinking Level · High | Tools
49.60Thinking Level · High | Tools
--
MCP-Atlas
AI Agent - 工具使用
83.60Thinking Level · High | Tools
62.00Standard Mode | Tools
--
Terminal-Bench 2.1
AI Agent - 工具使用
76.20Thinking Level · High | Tools
58.00Thinking Level · High | Tools
--

Single-Benchmark Version Trend

Viewing: ARC-AGI-2 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.5 Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.5 Flash
DeepMind$1.5 / 1M tokens$9 / 1M tokens
Gemini 3.0 Flash
Google Deep Mind$0.5 / 1M tokens$3 / 1M tokens
Gemini 2.5 Flash
Google Deep Mind$0.3 / 1M tokens$2.5 / 1M tokens

Sources