DataLearner logo

Gemini 3.5 Flash Benchmark Details

Gemini 3.5 Flash currently shows benchmark results led by MMMU-Pro (12 / 229, score 84.30), MedCode (4 / 64, score 55.83), GPQA Diamond (20 / 253, score 92.80). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Gemini 3.5 Flash

Benchmark Results

Thinking
Tool usage
Internet

Abstract Generalization

2 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
HighTools
92.50
52 / 197
ARC-AGI-2
HighTools
72.08
51 / 185

Knowledge Exams

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
HighTools
40.20
85 / 235

Scientific Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
88.89
51 / 253
92.80
20 / 253
CritPt
Standard Mode
1.40
139 / 204
CritPt
Thinking Level · Medium
10.90
72 / 204
CritPt
High
13.10
64 / 204

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
76.70
11 / 96

Repository Engineering

3 evaluations
Benchmark / mode
Score
Rank/total
55.10
33 / 62
DeepSWE
Thinking Level · MediumTools
37
81 / 93
DeepSWE
HighTools
36.06
83 / 93

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
74.64
15 / 117

Agentic Development

4 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Standard ModeTools
46.20
29 / 244
Terminal Bench Hard
Thinking Level · MediumTools
39.40
53 / 244
40.90
50 / 244
6.60
64 / 98

Memory & Persistence

5 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
33.49
109 / 126
50.74
89 / 126
Context Arena
Thinking Level · Medium
73.35
55 / 126
77.19
47 / 126
EBR-bench
HighTools
4.76
21 / 23

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
78.40
12 / 28

Tool Orchestration

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
83.60
7 / 44
74.17
21 / 45
67.30
22 / 33

Capability Indices

3 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
154.55
30 / 167
Vals Index
HighTools
53.08
28 / 42
44.79
31 / 36

Visual Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
80.10
43 / 229
MMMU-Pro
Thinking Level · Medium
83.90
15 / 229
84.30
12 / 229

Documents & Charts

3 evaluations
Benchmark / mode
Score
Rank/total
CharXiv RQ
Thinking Enabled
84.20
15 / 19
CharXiv RQ
Thinking EnabledTools
84.90
11 / 19
19.80
42 / 122

ML Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
MLE-Bench
Thinking EnabledTools
49.70
2 / 3

Long Retrieval

2 evaluations
Benchmark / mode
Score
Rank/total
77.30
4 / 8
26.60
2 / 3

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
53.90
50 / 134

Service Workflows

3 evaluations
Benchmark / mode
Score
Rank/total
59.47
20 / 26
SAGE
HighTools
49.88
11 / 64
τ³-Banking
HighTools
32.20
54 / 167

Legal

3 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
HighTools
82.08
28 / 44
30.77
33 / 45
2.50
29 / 43

Finance

2 evaluations
Benchmark / mode
Score
Rank/total
63.55
24 / 45
57.86
10 / 43

Mathematics

3 evaluations
Benchmark / mode
Score
Rank/total
62.81
34 / 114
31
22 / 28
26.83
40 / 71

Data Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
HighToolsInternet
45
9 / 29

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
26.75
35 / 47

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
48.68
40 / 64

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
HighTools
76.57
50 / 66
MedCode
HighTools
55.83
4 / 64

Competitor Comparison

Benchmark scores for Gemini 3.5 Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.5 FlashCurrentClaude Sonnet 4.6Opus 4.7GPT-5.5
ARC-AGI-1
Score (%)
Abstract Generalization
92.50Thinking Level · High | Tools
86.50Thinking Level · High | Tools
93.50Thinking Level · High
95.00Thinking Level · Extra High
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
72.08Thinking Level · High | Tools
60.42Thinking Level · High | Tools
75.80Thinking Level · High
85.00Thinking Level · Extra High
HLE
Accuracy
Knowledge Exams
40.20Thinking Level · High | Tools
49.00Thinking Enabled | Tools
54.70Extended Thinking | Tools
52.20Thinking Level · High | Tools
CritPt
Score
Scientific Reasoning
13.10Thinking Level · High
3.10Thinking Level · High
12.00Thinking Level · High
27.10Thinking Level · Extra High
GPQA Diamond
Accuracy
Scientific Reasoning
92.80Thinking Level · High
89.90Thinking Enabled
94.20Extended Thinking
94.00Thinking Level · Extra High
SimpleBench
Score (AVG@5)
Commonsense
76.70Standard Mode
--
61.70Standard Mode
69.00Standard Mode
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
37.00Thinking Level · Medium | Tools
29.93Thinking Level · High | Tools
--
67.04Thinking Level · Extra High | Tools
SWE-Bench Pro - Public
Accuracy
Repository Engineering
55.10Thinking Level · High | Tools
--
64.30Extended Thinking | Tools
58.60Thinking Level · High | Tools
LiveBench
Accuracy
Cross-capability Suites
74.64Thinking Level · High
75.32Thinking Level · High
76.53Deep Thinking Mode
79.91Deep Thinking Mode
Terminal Bench Hard
Accuracy
Agentic Development
46.20Standard Mode | Tools
53.00Thinking Level · High | Tools
54.50Standard Mode | Tools
60.60Thinking Level · Extra High | Tools
Terminal-Bench 4.0
Resolution rate (%)
Agentic Development
6.60Thinking Level · High | Tools
3.00Thinking Level · High | Tools
--
14.60Thinking Level · Extra High | Tools
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
77.19Thinking Level · High
82.82Thinking Level · High
46.70Thinking Level · Low
94.18Thinking Level · Extra High
25 additional benchmarks remain in the chart above.

Standard API Pricing: Gemini 3.5 Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.5 Flash
Google DeepMind$1.5 / 1M tokens$9 / 1M tokens—
Claude Sonnet 4.6
Anthropic$3 / 1M tokens$15 / 1M tokens—
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens—
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens—

Version History

How each version of the Gemini 3.5 Flash series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.5 FlashCurrentGemini 3.0 FlashGemini 2.5 Flash
ARC-AGI-1
Score (%)
Abstract Generalization
92.50Thinking Level · High | Tools
--
32.30Standard Mode
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
72.08Thinking Level · High | Tools
33.60Thinking Enabled
--
HLE
Accuracy
Knowledge Exams
40.20Thinking Level · High | Tools
43.50Thinking Enabled | Tools
12.08Thinking Level · High
CritPt
Score
Scientific Reasoning
13.10Thinking Level · High
8.60Thinking Enabled
1.40Standard Mode
GPQA Diamond
Accuracy
Scientific Reasoning
92.80Thinking Level · High
90.40Thinking Enabled
82.80Thinking Enabled
SimpleBench
Score (AVG@5)
Commonsense
76.70Standard Mode
61.10Standard Mode
41.20Standard Mode
SWE-Bench Pro - Public
Accuracy
Repository Engineering
55.10Thinking Level · High | Tools
49.60Thinking Level · High | Tools
--
LiveBench
Accuracy
Cross-capability Suites
74.64Thinking Level · High
72.40Thinking Level · High
47.74Thinking Level · High
Terminal Bench Hard
Accuracy
Agentic Development
46.20Standard Mode | Tools
38.60Thinking Enabled | Tools
13.60Thinking Enabled | Tools
MCP-Atlas
Pass rate / claim coverage
Tool Orchestration
83.60Thinking Level · High | Tools
62.00Standard Mode | Tools
--
PinchBench v2
Average score (%)
Tool Orchestration
74.17Thinking Level · High
72.06Thinking Level · High
--
ECI
ECI score (capability index, higher is better)
Capability Indices
154.55Thinking Level · High
151.83Thinking Level · High
--
5 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.5 Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.5 Flash
Google DeepMind$1.5 / 1M tokens$9 / 1M tokens—
Gemini 3.0 Flash
Google DeepMind$0.5 / 1M tokens$3 / 1M tokens—
Gemini 2.5 Flash
Google DeepMind$0.3 / 1M tokens$2.5 / 1M tokens—

Sources