DataLearner logo

Gemini 3.6 Flash Benchmark Details

Gemini 3.6 Flash currently shows benchmark results led by GPQA Diamond (8 / 271, score 94.13), Context Arena (17 / 126, score 88.78), OSWorld-Verified (5 / 26, score 83). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Gemini 3.6 Flash

Benchmark Results

Thinking
Tool usage

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Low
86.36
86 / 271
94.13
8 / 271

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1599.70
34 / 99

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Thinking EnabledTools
58.70
16 / 60
WeirdML v2
HighTools
56.10
33 / 52
DeepSWE
Thinking EnabledTools
49
28 / 35

Text Embedding

3 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Thinking Level · Low
71.70
60 / 126
86.56
21 / 126
88.78
17 / 126

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking EnabledTools
83
5 / 26
Terminal-Bench 2.1
Thinking EnabledTools
78
27 / 49
MLE-Bench
Thinking EnabledTools
63.90
1 / 3

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Enabled
1421
18 / 25

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
CharXiv RQ
Thinking Enabled
85.20
10 / 19
CharXiv RQ
Thinking EnabledTools
89.40
4 / 19

Other

2 evaluations
Benchmark / mode
Score
Rank/total
91.80
2 / 8
54
1 / 3

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
58.95
23 / 58
21.95
25 / 41

Competitor Comparison

Benchmark scores for Gemini 3.6 Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.6 FlashCurrentGPT-5.6 TerraClaude Sonnet 5Grok 4.5
GPQA Diamond
科学与综合推理
94.13Thinking Level · High
93.31Thinking Level · High
90.53Thinking Level · Extra High
93.43Thinking Level · High
Creative Writing
写作和创作
1599.70Standard Mode
1850.00Standard Mode
1787.60Standard Mode
1576.00Standard Mode
DeepSWE
编程与软件工程
49.00Thinking Enabled | Tools
69.60Thinking Level · Extra High | Tools
54.00Deep Thinking Mode | Tools
53.00Thinking Level · High | Tools
SWE-Bench Pro - Public
编程与软件工程
58.70Thinking Enabled | Tools
--
--
64.70Thinking Level · High | Tools
WeirdML v2
编程与软件工程
56.10Thinking Level · High | Tools
78.27Thinking Level · High | Tools
68.78Thinking Level · High | Tools
--
Context Arena
文本向量检索
88.78Thinking Level · High
92.17Thinking Level · High
79.53Thinking Level · High
--
OSWorld-Verified
AI Agent - 工具使用
83.00Thinking Enabled | Tools
--
81.20Thinking Level · Extra High | Tools
--
Terminal-Bench 2.1
AI Agent - 工具使用
78.00Thinking Enabled | Tools
87.40Thinking Level · High
80.40Thinking Level · Extra High | Tools
83.30Thinking Level · High | Tools
GDPval-AA v2
生产力知识
1421.00Thinking Enabled
1565.62Thinking Level · High | Tools
--
1526.00Thinking Level · High | Tools
21.95Thinking Level · High
70.73Thinking Level · High
29.27Thinking Level · High
24.39Thinking Level · High
FrontierMath v2
数学推理
58.95Thinking Level · High
85.96Thinking Level · High
65.61Thinking Level · High
57.19Thinking Level · High

Standard API Pricing: Gemini 3.6 Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.6 Flash
DeepMind$1.5 / 1M tokens$7.5 / 1M tokens
GPT-5.6 Terra
OpenAI$2 / 1M tokens$12 / 1M tokens
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens
Grok 4.5
xAI$2 / 1M tokens$6 / 1M tokens

Version History

How each version of the Gemini 3.6 Flash series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.6 FlashCurrentGemini 3.5 FlashGemini 3.0 Flash
GPQA Diamond
科学与综合推理
94.13Thinking Level · High
92.80Thinking Level · High
90.40Thinking Enabled
DeepSWE
编程与软件工程
49.00Thinking Enabled | Tools
37.00Thinking Level · Medium | Tools
--
SWE-Bench Pro - Public
编程与软件工程
58.70Thinking Enabled | Tools
55.10Thinking Level · High | Tools
49.60Thinking Level · High | Tools
WeirdML v2
编程与软件工程
56.10Thinking Level · High | Tools
--
61.60Standard Mode | Tools
Context Arena
文本向量检索
88.78Thinking Level · High
77.19Thinking Level · High
--
MLE-Bench
AI Agent - 工具使用
63.90Thinking Enabled | Tools
49.70Thinking Enabled | Tools
--
OSWorld-Verified
AI Agent - 工具使用
83.00Thinking Enabled | Tools
78.40Thinking Level · High | Tools
--
Terminal-Bench 2.1
AI Agent - 工具使用
78.00Thinking Enabled | Tools
76.20Thinking Level · High | Tools
58.00Thinking Level · High | Tools
GDPval-AA v2
生产力知识
1421.00Thinking Enabled
1349.00Thinking Enabled
--
CharXiv RQ
多模态理解
89.40Thinking Enabled | Tools
84.90Thinking Enabled | Tools
--
91.80Thinking Enabled
77.30Thinking Enabled
--
54.00Thinking Enabled
26.60Thinking Enabled
--
2 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 科学与综合推理

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.6 Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.6 Flash
DeepMind$1.5 / 1M tokens$7.5 / 1M tokens
Gemini 3.5 Flash
DeepMind$1.5 / 1M tokens$9 / 1M tokens
Gemini 3.0 Flash
Google Deep Mind$0.5 / 1M tokens$3 / 1M tokens