DataLearner logo
GE

Gemini 3.8 Flash

Multimodal modelCoding modelGemini 3.7

Gemini 3.8 Flash

Also known as: gemini-3.8-flash / Gemini 3.8 / Skimaki / Gemini 3.8 Flash Preview / gemini-3.8-flash-preview

Release date: 2026-09-02Updated: 2026-09-02Views: 50
Live demoGitHubHugging FaceCompare
Parameters
No data
Context length
1M
Multilingual
No data
Reasoning ability
4/5

Google's September 2, 2026 GA multimodal Flash model, with the stable gemini-3.8-flash API ID, a 1M-token context window, 64K maximum output, and low/medium/high thinking levels for long-horizon software engineering, autonomous agents, and enterprise workflows.

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

Gemini 3.8 Flash

Model basics

Reasoning traces
Supported
Thinking modes
Thinking Level · Medium (Default)Thinking Level · LowThinking Level · High
Context length
1M tokens
Max output length
64K tokens
Model type
Multimodal model
Modality (in / out)
Text, Image, Audio, Video → Text
Release date
2026-09-02
Model file size
No data
MoE architecture
No
Total params / Active params
No data / Not applicable
Knowledge cutoff
No data
Gemini 3.8 Flash

Open source & experience

Code license
Proprietary
Weights license
Proprietary
GitHub repo
N/A
Hugging Face
N/A
Gemini 3.8 Flash

Official resources

Gemini 3.8 Flash

API details

API speed
4/5
💡Default unit: $/1M tokens. If vendors use other units, follow their published pricing.
Standard
TypeConditionInputOutput
Text-$0.750/ 1M$3.75/ 1M
Batch
TypeConditionInputOutput
Text-$0.375/ 1M$1.88/ 1M
flex
TypeConditionInputOutput
Text-$0.375/ 1M$1.88/ 1M
priority
TypeConditionInputOutput
Text-$1.35/ 1M$6.75/ 1M
Cache PricingPrompt Cache
TypeTTLWriteRead
Text-$0.037/ 1M

“—” means the modality is not billed in that direction, or the vendor has not published a price for it.

Gemini 3.8 Flash

Benchmark Results

Gemini 3.8 Flash currently shows benchmark results led by Terminal-Bench 2.1 (1 / 48, score 89.40), DeepSWE (1 / 33, score 73.70), LVBench (1 / 4, score 87.80). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking
Tool usage
Internet

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking ModeTools
89.40
1 / 48
88.80
1 / 2
LABBench2
MediumToolsInternet
86.20
1 / 2
OSWorld 2.0
Thinking ModeTools
59
4 / 7
56.50
1 / 2
Terminal-Bench 4.0
Thinking ModeTools
19.10
9 / 12

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
DeepSWE
Thinking ModeTools
73.70
1 / 33

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Mode
1545
13 / 24
Finance Agent v2
Thinking ModeTools
61.40
1 / 2

Multimodal Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
LVBench
Medium
87.10
2 / 4
LVBench
MediumTools
87.80
1 / 4
86.20
8 / 19
GDP.pdf
Medium
35
1 / 2

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
54.90
1 / 2

Compare with other models

Gemini 3.8 Flash

Publisher

Google Deep Mind
Google Deep Mind
View publisher details
Gemini 3.8 Flash

Model Overview

Official release

Gemini 3.8 Flash became generally available on September 2, 2026 under the stable API ID gemini-3.8-flash. Google calls it its most intelligent Flash model and positions it for long-horizon software engineering, autonomous agents, and complex enterprise workflows. Pre-release reporting identified the internal codename as Skimaki; that codename is not a public API identifier.

Specifications and I/O

Google documents a 1,048,576-token input limit and a 65,536-token output limit. Inputs can be text, images, video, audio, and PDF; output is text. Supported capabilities include context caching, code execution, Computer Use (Preview), File Search, function calling, Google Search and Maps grounding, structured output, and URL Context. Batch, Flex, and Priority inference are available.

Thinking and agents

Gemini 3.8 Flash supports low, medium, and high thinking levels, with medium as the default. The minimal level is unsupported and returns an error. Google says the model deliberately takes smaller reasoning steps, invokes tools iteratively, and verifies work on long and complex tasks, which can increase token use. Gemini Managed Agents' Antigravity agent and SDK now use 3.8 Flash by default.

Launch evaluations

Google's launch material reports 73.7% on DeepSWE v1.1, 1545 Elo on GDPval-AA v2, 61.4% on Vals Finance Agent v2, 10.0% on Harvey's Legal Agent Benchmark, 89.4% on Terminal-Bench 2.1, 19.1% on Terminal-Bench 4.0, 35.0% on GDP.pdf, 86.2% on CharXiv Reasoning without tools, 87.8% agentic / 87.1% static on LVBench, 54.9% on HLE-Verified, 59.0% partial on OSWorld 2.0, 88.8% / 56.5% on the human-solvable / human-difficult BioMysteryBench subsets, and 86.2% on LABBench2.

Google's methodology states that Gemini scores are pass@1 unless noted and use the default sampling settings for gemini-3.8-flash; smaller benchmarks are averaged over multiple trials. DeepSWE uses the mini-swe agent harness with high thinking, Terminal-Bench 2.1 uses Terminus 2, LVBench runs without tools with 1,024 frames for Gemini, and OSWorld 2.0 reports the best partial score across three single-attempt runs. Scores should only be compared within the same benchmark and test setup.

Pricing

Through December 31, 2026, Gemini Developer API Standard introductory pricing per million tokens is $0.75 input, $3.75 output including thinking tokens, and $0.075 cached input. Batch and Flex are $0.375 / $1.875 / $0.0375; Priority is $1.35 / $6.75 / $0.135. These rates double on January 1, 2027. Cache storage and grounding requests are billed separately.

Known unknowns

Google has not disclosed parameter count, training compute, weights, or a knowledge cutoff. This record does not inherit those values from Gemini 3.7 Flash and does not treat pre-release employee preference testing as a separate benchmark.

Official sources

Gemini 3.8 Flash

FAQ

Has Gemini 3.8 Flash been officially released?

Yes. Google's Gemini API release notes record general availability on September 2, 2026 under the stable API ID gemini-3.8-flash.

What are Gemini 3.8 Flash's context and thinking levels?

Google documents a 1,048,576-token input limit, a 65,536-token output limit, and low, medium, and high thinking levels. Medium is the default; minimal is unsupported.

How much does the Gemini 3.8 Flash API cost?

Through December 31, 2026, Standard introductory pricing per million tokens is $0.75 input, $3.75 output, and $0.075 cached input. The respective rates become $1.50, $7.50, and $0.15 on January 1, 2027.

What are Gemini 3.8 Flash's parameter count and knowledge cutoff?

Google has not disclosed its parameter count, weight architecture, training compute, or knowledge cutoff. This record does not infer those fields from Gemini 3.7 Flash.

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code