DataLearner logo

Gemini 3.7 Flash Benchmark Details

Gemini 3.7 Flash currently shows benchmark results led by GPQA Diamond (1 / 253, score 94.82), GDP.pdf (1 / 122, score 34), MMMU-Pro (5 / 229, score 85.50). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Gemini 3.7 Flash

Benchmark Results

Thinking
Tool usage
Internet

Abstract Generalization

6 evaluations
Benchmark / mode
Score
Rank/total
85.17
86 / 176
ARC-AGI-1
Medium
91.17
56 / 176
95.50
29 / 176
52.92
78 / 164
ARC-AGI-2
Medium
63.75
61 / 164
84.58
27 / 164

Scientific Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
94.82
1 / 253
5.70
92 / 204
CritPt
Medium
9.40
77 / 204
CritPt
High
14.30
61 / 204

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1722
24 / 106

Memory & Persistence

3 evaluations
Benchmark / mode
Score
Rank/total
89.57
16 / 126
93.38
8 / 126
95.95
4 / 126

Agentic Development

4 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking EnabledTools
85.80
17 / 57
14.90
10 / 11
Terminal-Bench 4.0
Thinking EnabledTools
11.21
55 / 96
13.60
47 / 96

Repository Engineering

3 evaluations
Benchmark / mode
Score
Rank/total
DeepSWE
LowTools
53.76
59 / 91
DeepSWE
MediumTools
65.49
38 / 91
DeepSWE
HighTools
65.27
40 / 91

Capability Indices

3 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
157.72
11 / 167
Vals Index
HighTools
59.31
16 / 42

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld 2.0
MediumTools
47.90
13 / 13

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
26.30
18 / 24

Office & Business

1 evaluations
Benchmark / mode
Score
Rank/total
AutomationBench
Thinking EnabledTools
30.40
19 / 23

Visual Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
84.90
7 / 229
MMMU-Pro
Medium
84.70
8 / 229
85.50
5 / 229

Documents & Charts

4 evaluations
Benchmark / mode
Score
Rank/total
84.50
14 / 19
CharXiv RQ
MediumTools
88.70
6 / 19
GDP.pdf
Medium
34
1 / 122
23.60
25 / 122

Long Retrieval

1 evaluations
Benchmark / mode
Score
Rank/total

Scientific Computing

3 evaluations
Benchmark / mode
Score
Rank/total
55.70
33 / 134
SciCode
Medium
59.80
5 / 134
57.20
20 / 134

Service Workflows

4 evaluations
Benchmark / mode
Score
Rank/total
SAGE
HighTools
49.23
16 / 64
τ³-Banking
LowTools
29.50
64 / 167
τ³-Banking
MediumTools
35.50
46 / 167
τ³-Banking
HighTools
32.80
52 / 167

Legal

3 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Thinking Enabled
90.66
13 / 44
34.62
28 / 42
8.75
13 / 43

Finance

3 evaluations
Benchmark / mode
Score
Rank/total
71.33
9 / 42
59.04
4 / 42
57.66
19 / 20

Mathematics

3 evaluations
Benchmark / mode
Score
Rank/total
71.58
14 / 58
58
13 / 28
36.59
16 / 42

Code Generation & Editing

2 evaluations
Benchmark / mode
Score
Rank/total
70.39
25 / 60
FrontierCode 1.1 Extended
Thinking EnabledTools
43.60
7 / 7

Preference Arenas

1 evaluations
Benchmark / mode
Score
Rank/total
Code Arena WebDev
Thinking EnabledTools
1588
1 / 1

Video Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
LVBench
Medium
85.40
3 / 4

Knowledge Exams

1 evaluations
Benchmark / mode
Score
Rank/total
53.60
4 / 6

Biology & Genomics

3 evaluations
Benchmark / mode
Score
Rank/total
87.10
2 / 2
LABBench2
MediumToolsInternet
82.10
2 / 2
43.50
2 / 2

Data Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
HighToolsInternet
60
1 / 29

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
34.80
26 / 43

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
IOI (Vals v2)
HighTools
67.83
9 / 26

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
HighTools
83.94
27 / 66
MedCode
HighTools
53.39
7 / 64

Competitor Comparison

Benchmark scores for Gemini 3.7 Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.7 FlashCurrentClaude Sonnet 5Qwen3.8-27BDeepSeek-V4-Flash
CritPt
Score
Scientific Reasoning
14.30Thinking Level · High
16.90Thinking Level · High
5.40Thinking Level · Extra High
7.10Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
94.82Thinking Level · High
90.53Thinking Level · Extra High
89.20Thinking Enabled
71.20Standard Mode
Creative Writing
Elo、大模型评判两两对战
Writing
1722.00Standard Mode
1790.50Standard Mode
1668.40Standard Mode
1555.70Standard Mode
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
95.95Thinking Level · High
79.53Thinking Level · High
93.98Thinking Level · Medium
69.42Thinking Enabled
Terminal-Bench 2.1
Accuracy
Agentic Development
85.80Thinking Enabled | Tools
80.40Thinking Level · Extra High | Tools
73.00Thinking Enabled | Tools
--
Terminal-Bench 3.0
Accuracy
Agentic Development
14.90Thinking Level · High | Tools
--
--
7.60Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
Agentic Development
13.60Thinking Level · High | Tools
12.42Thinking Level · High | Tools
5.60Thinking Level · Extra High | Tools
3.00Thinking Level · High | Tools
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
65.49Thinking Level · Medium | Tools
54.00Deep Thinking Mode | Tools
42.20Thinking Enabled | Tools
53.32Thinking Level · High | Tools
ECI
ECI score (capability index, higher is better)
Capability Indices
157.72Thinking Level · High
156.34Thinking Level · High
149.38Thinking Level · High
146.10Thinking Level · High
Vals Index
跨行业任务准确率综合指数
Capability Indices
59.31Thinking Level · High | Tools
59.61Thinking Level · High | Tools
48.48Thinking Level · Extra High | Tools
--
Agents' Last Exam
Score
Tool Orchestration
26.30Thinking Level · Medium | Tools
--
20.40Thinking Enabled | Tools
25.20Thinking Level · High | Tools
AutomationBench
Pass Rate
Office & Business
30.40Thinking Enabled | Tools
--
--
37.70Thinking Level · High | Tools
22 additional benchmarks remain in the chart above.

Standard API Pricing: Gemini 3.7 Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.7 Flash
Google DeepMind$0.75 / 1M tokens$3.75 / 1M tokens—
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens—
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens—

Version History

How each version of the Gemini 3.7 Flash series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.7 FlashCurrentGemini 3.6 FlashGemini 3.5 FlashGemini 3.0 Flash
ARC-AGI-1
Score (%)
Abstract Generalization
95.50Thinking Level · High
91.17Thinking Level · High | Tools
92.50Thinking Level · High | Tools
--
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
84.58Thinking Level · High
60.42Thinking Level · High | Tools
72.08Thinking Level · High | Tools
33.60Thinking Enabled
CritPt
Score
Scientific Reasoning
14.30Thinking Level · High
10.60Thinking Level · High
13.10Thinking Level · High
8.60Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
94.82Thinking Level · High
94.13Thinking Level · High
92.80Thinking Level · High
90.40Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
Writing
1722.00Standard Mode
1599.80Standard Mode
--
--
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
95.95Thinking Level · High
88.78Thinking Level · High
77.19Thinking Level · High
--
Terminal-Bench 2.1
Accuracy
Agentic Development
85.80Thinking Enabled | Tools
78.00Thinking Enabled | Tools
--
58.00Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
Agentic Development
13.60Thinking Level · High | Tools
7.10Thinking Level · High | Tools
6.60Thinking Level · High | Tools
--
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
65.49Thinking Level · Medium | Tools
49.00Thinking Enabled | Tools
37.00Thinking Level · Medium | Tools
--
ECI
ECI score (capability index, higher is better)
Capability Indices
157.72Thinking Level · High
154.36Thinking Level · High
154.55Thinking Level · High
151.83Thinking Level · High
Vals Index
跨行业任务准确率综合指数
Capability Indices
59.31Thinking Level · High | Tools
55.35Thinking Level · High | Tools
53.08Thinking Level · High | Tools
--
MMMU-Pro
Accuracy
Visual Understanding
85.50Thinking Level · High
83.20Thinking Level · High
84.30Thinking Level · High
79.90Thinking Enabled
20 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.7 Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.7 Flash
Google DeepMind$0.75 / 1M tokens$3.75 / 1M tokens—
Gemini 3.6 Flash
Google DeepMind$1.5 / 1M tokens$7.5 / 1M tokens—
Gemini 3.5 Flash
Google DeepMind$1.5 / 1M tokens$9 / 1M tokens—
Gemini 3.0 Flash
Google DeepMind$0.5 / 1M tokens$3 / 1M tokens—