Has Gemini 3.8 Flash been officially released?
Yes. Google's Gemini API release notes record general availability on September 2, 2026 under the stable API ID gemini-3.8-flash.
Gemini 3.8 Flash
Also known as: gemini-3.8-flash / Gemini 3.8 / Skimaki / Gemini 3.8 Flash Preview / gemini-3.8-flash-preview
Google's September 2, 2026 GA multimodal Flash model, with the stable gemini-3.8-flash API ID, a 1M-token context window, 64K maximum output, and low/medium/high thinking levels for long-horizon software engineering, autonomous agents, and enterprise workflows.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $0.750/ 1M | $3.75/ 1M |
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $0.375/ 1M | $1.88/ 1M |
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $0.375/ 1M | $1.88/ 1M |
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $1.35/ 1M | $6.75/ 1M |
| Type | TTL | Write | Read |
|---|---|---|---|
| Text | - | — | $0.037/ 1M |
“—” means the modality is not billed in that direction, or the vendor has not published a price for it.
Gemini 3.8 Flash currently shows benchmark results led by Terminal-Bench 2.1 (1 / 48, score 89.40), DeepSWE (1 / 33, score 73.70), LVBench (1 / 4, score 87.80). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.
Want a custom combination? Open the compare tool
Gemini 3.8 Flash became generally available on September 2, 2026 under the stable API ID gemini-3.8-flash. Google calls it its most intelligent Flash model and positions it for long-horizon software engineering, autonomous agents, and complex enterprise workflows. Pre-release reporting identified the internal codename as Skimaki; that codename is not a public API identifier.
Google documents a 1,048,576-token input limit and a 65,536-token output limit. Inputs can be text, images, video, audio, and PDF; output is text. Supported capabilities include context caching, code execution, Computer Use (Preview), File Search, function calling, Google Search and Maps grounding, structured output, and URL Context. Batch, Flex, and Priority inference are available.
Gemini 3.8 Flash supports low, medium, and high thinking levels, with medium as the default. The minimal level is unsupported and returns an error. Google says the model deliberately takes smaller reasoning steps, invokes tools iteratively, and verifies work on long and complex tasks, which can increase token use. Gemini Managed Agents' Antigravity agent and SDK now use 3.8 Flash by default.
Google's launch material reports 73.7% on DeepSWE v1.1, 1545 Elo on GDPval-AA v2, 61.4% on Vals Finance Agent v2, 10.0% on Harvey's Legal Agent Benchmark, 89.4% on Terminal-Bench 2.1, 19.1% on Terminal-Bench 4.0, 35.0% on GDP.pdf, 86.2% on CharXiv Reasoning without tools, 87.8% agentic / 87.1% static on LVBench, 54.9% on HLE-Verified, 59.0% partial on OSWorld 2.0, 88.8% / 56.5% on the human-solvable / human-difficult BioMysteryBench subsets, and 86.2% on LABBench2.
Google's methodology states that Gemini scores are pass@1 unless noted and use the default sampling settings for gemini-3.8-flash; smaller benchmarks are averaged over multiple trials. DeepSWE uses the mini-swe agent harness with high thinking, Terminal-Bench 2.1 uses Terminus 2, LVBench runs without tools with 1,024 frames for Gemini, and OSWorld 2.0 reports the best partial score across three single-attempt runs. Scores should only be compared within the same benchmark and test setup.
Through December 31, 2026, Gemini Developer API Standard introductory pricing per million tokens is $0.75 input, $3.75 output including thinking tokens, and $0.075 cached input. Batch and Flex are $0.375 / $1.875 / $0.0375; Priority is $1.35 / $6.75 / $0.135. These rates double on January 1, 2027. Cache storage and grounding requests are billed separately.
Google has not disclosed parameter count, training compute, weights, or a knowledge cutoff. This record does not inherit those values from Gemini 3.7 Flash and does not treat pre-release employee preference testing as a separate benchmark.
Yes. Google's Gemini API release notes record general availability on September 2, 2026 under the stable API ID gemini-3.8-flash.
Google documents a 1,048,576-token input limit, a 65,536-token output limit, and low, medium, and high thinking levels. Medium is the default; minimal is unsupported.
Through December 31, 2026, Standard introductory pricing per million tokens is $0.75 input, $3.75 output, and $0.075 cached input. The respective rates become $1.50, $7.50, and $0.15 on January 1, 2027.
Google has not disclosed its parameter count, weight architecture, training compute, or knowledge cutoff. This record does not infer those fields from Gemini 3.7 Flash.
Follow DataLearner on WeChat for AI model updates and research notes.
