DataLearner logo

Gemini 3.8 Flash Benchmark Details

Gemini 3.8 Flash currently shows benchmark results led by Terminal-Bench 2.1 (1 / 48, score 89.40), DeepSWE (1 / 33, score 73.70), LVBench (1 / 4, score 87.80). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Gemini 3.8 Flash

Benchmark Results

Thinking
Tool usage
Internet

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking EnabledTools
89.40
1 / 48
88.80
1 / 2
LABBench2
MediumToolsInternet
86.20
1 / 2
OSWorld 2.0
Thinking EnabledTools
59
4 / 7
56.50
1 / 2
Terminal-Bench 4.0
Thinking EnabledTools
19.10
9 / 12

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
DeepSWE
Thinking EnabledTools
73.70
1 / 33

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Enabled
1545
13 / 24
Finance Agent v2
Thinking EnabledTools
61.40
1 / 2

Multimodal Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
LVBench
Medium
87.10
2 / 4
LVBench
MediumTools
87.80
1 / 4
86.20
8 / 19
GDP.pdf
Medium
35
1 / 2

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
54.90
1 / 2

Competitor Comparison

Benchmark scores for Gemini 3.8 Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

5 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.8 FlashCurrentClaude Sonnet 5GPT-5.6 TerraGLM-5.3-Flash
Terminal-Bench 2.1
AI Agent - 工具使用
89.40Thinking Enabled | Tools
80.40Thinking Level · Extra High | Tools
87.40Thinking Level · High
84.30Thinking Level · High | Tools
Terminal-Bench 4.0
AI Agent - 工具使用
19.10Thinking Enabled | Tools
12.42Thinking Level · High | Tools
21.52Thinking Level · High | Tools
--
DeepSWE
编程与软件工程
73.70Thinking Enabled | Tools
54.00Deep Thinking Mode | Tools
69.60Thinking Level · Extra High | Tools
63.40Thinking Level · High | Tools
GDPval-AA v2
生产力知识
1545.00Thinking Enabled
--
1565.62Thinking Level · High | Tools
1773.00Thinking Level · High | Tools
CharXiv RQ
多模态理解
86.20Thinking Level · Medium
--
--
89.40Thinking Level · High | Tools

Standard API Pricing: Gemini 3.8 Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.8 Flash
DeepMind$0.75 / 1M tokens$3.75 / 1M tokens
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens
GPT-5.6 Terra
OpenAI$2 / 1M tokens$12 / 1M tokens
GLM-5.3-Flash
智谱AI$0.075 / 1M tokens$0.25 / 1M tokens

Version History

How each version of the Gemini 3.8 Flash series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.8 FlashCurrentGemini 3.7 FlashGemini 3.6 FlashGemini 3.5 Flash
BioMysteryBench (Human-difficult)
AI Agent - 工具使用
56.50Thinking Level · Medium | Tools
43.50Thinking Level · Medium | Tools
--
--
BioMysteryBench (Human-solvable)
AI Agent - 工具使用
88.80Thinking Level · Medium | Tools
87.10Thinking Level · Medium | Tools
--
--
LABBench2
AI Agent - 工具使用
86.20Thinking Level · Medium | Tools
82.10Thinking Level · Medium | Tools
--
--
OSWorld 2.0
AI Agent - 工具使用
59.00Thinking Enabled | Tools
47.90Thinking Level · Medium | Tools
--
--
Terminal-Bench 2.1
AI Agent - 工具使用
89.40Thinking Enabled | Tools
85.80Thinking Enabled | Tools
78.00Thinking Enabled | Tools
76.20Thinking Level · High | Tools
DeepSWE
编程与软件工程
73.70Thinking Enabled | Tools
65.30Thinking Level · High | Tools
49.00Thinking Enabled | Tools
37.00Thinking Level · Medium | Tools
GDPval-AA v2
生产力知识
1545.00Thinking Enabled
1525.00Thinking Enabled
1421.00Thinking Enabled
1349.00Thinking Enabled
CharXiv RQ
多模态理解
86.20Thinking Level · Medium
88.70Thinking Level · Medium | Tools
89.40Thinking Enabled | Tools
84.90Thinking Enabled | Tools
GDP.pdf
多模态理解
35.00Thinking Level · Medium
34.00Thinking Level · Medium
--
--
LVBench
多模态理解
87.80Thinking Level · Medium | Tools
85.40Thinking Level · Medium
--
--
HLE-Verified
综合评估
54.90Thinking Level · Medium
53.60Thinking Level · Medium
--
--

Single-Benchmark Version Trend

Viewing: BioMysteryBench (Human-difficult) · AI Agent - 工具使用

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.8 Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.8 Flash
DeepMind$0.75 / 1M tokens$3.75 / 1M tokens
Gemini 3.7 Flash
DeepMind$0.75 / 1M tokens$3.75 / 1M tokens
Gemini 3.6 Flash
DeepMind$1.5 / 1M tokens$7.5 / 1M tokens
Gemini 3.5 Flash
DeepMind$1.5 / 1M tokens$9 / 1M tokens