Is Gemini 3.7 Flash stable or in preview?
It is generally available and ready for production. The stable API model ID is gemini-3.7-flash.
Gemini 3.7 Flash
Google's generally available Flash model released on August 13, 2026, with a 1M-token context window, 64K maximum output, low/medium/high thinking levels, and a focus on coding, UI generation, and multi-step agents. Introductory Standard pricing through 2026 is $0.75/M input and $3.75/M output.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $0.750/ 1M | $3.75/ 1M |
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $0.375/ 1M | $1.88/ 1M |
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $0.375/ 1M | $1.88/ 1M |
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $1.35/ 1M | $6.75/ 1M |
| Type | TTL | Write | Read |
|---|---|---|---|
| Text | - | - | $0.037/ 1M |
Gemini 3.7 Flash currently shows benchmark results led by Terminal-Bench 2.1 (8 / 40, score 85.80), CharXiv RQ (3 / 13, score 88.70), GDM-MRCR v2 (8-needle, 128K) (1 / 4, score 97). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.
Want a custom combination? Open the compare tool
Gemini 3.7 Flash is a generally available Gemini 3 multimodal reasoning model released by Google on August 13, 2026. Its stable API identifier is gemini-3.7-flash. Google positions it as a workhorse model for coding and agents, with algorithmic improvements over Gemini 3.6 Flash aimed at real-world software engineering, web and UI generation, complex knowledge work, and reliable multi-step execution.
The official developer documentation lists a 1,048,576-token input context window and 65,536-token maximum output. It accepts text, image, video, audio, and PDF inputs and produces text. Supported capabilities include context caching, code execution, function calling, File Search, Google Search and Maps grounding, structured outputs, URL Context, and Computer Use in Preview. Batch, Flex, and Priority inference are also supported.
Gemini 3.7 Flash supports low, medium, and high thinking levels; minimal is not supported. Medium is the default and is recommended for most complex code and agent tasks. Low favors latency, while high allows longer reasoning and more tool use for difficult math, coding, and agentic work. Google's migration guide says clients upgrading from older models should remove deprecated temperature, top_p, top_k, and prefilled model turns.
Google DeepMind reports an Artificial Analysis Intelligence Index of 56, FrontierCode 1.1 Main at 43.6%, DeepSWE v1.1 at 65.3%, Code Arena at 1588 Elo, Terminal-Bench 2.1 at 85.8%, Terminal-Bench 3.0 at 14.9%, AutomationBench at 30.4%, GDPval-AA v2 at 1525 Elo, Harvey LAB-AA at 90.7%, GDP.pdf at 34.0%, CharXiv Reasoning at 84.5% without tools and 88.7% with tools, LVBench at 85.4%, GDM-MRCR v2 128K at 97.0%, OSWorld 2.0 at 47.9%, Agent's Last Exam at 26.3%, HLE-Verified at 53.6%, BioMysteryBench at 87.1% on human-solvable tasks and 43.5% on human-difficult tasks, and LABBench2 at 82.1%.
Test conditions differ by benchmark. Google says default model and sampling settings were used unless noted; DeepSWE used high thinking with a mini-SWE agent harness, Terminal-Bench 2.1 used Terminus 2, the tool-enabled CharXiv run used search and code execution, LVBench used 1,024 frames without tools, and the biology evaluations provided Linux, Python, R, bioinformatics tools, and restricted internet access. Scores should therefore be interpreted together with their harness and tool conditions.
The model card gives a March 2026 knowledge cutoff, while warning that some domains may remain closer to the Gemini 3 family's January 2025 baseline. The model can still hallucinate and may occasionally be slow or time out. Google has not disclosed weights, total parameter count, or active parameter count, so DataLearner leaves those fields unspecified.
Through December 31, 2026, Gemini Developer API Standard pricing is $0.75 per million input tokens, $3.75 per million output tokens including thinking tokens, and $0.075 per million cached input tokens. Batch and Flex are $0.375 / $1.875 / $0.0375, while Priority is $1.35 / $6.75 / $0.135. Starting January 1, 2027, Standard pricing becomes $1.50 / $7.50 / $0.15, with corresponding increases for other tiers. Cache storage and grounding requests are billed separately.
Gemini 3.7 Flash is available through the Gemini API, Google AI Studio, Google Antigravity, Gemini Enterprise Agent Platform, the Gemini Enterprise app, and Gemini Spark. It is a proprietary API model governed by Google's Gemini API terms.
It is generally available and ready for production. The stable API model ID is gemini-3.7-flash.
The input context limit is 1,048,576 tokens and the maximum output is 65,536 tokens. Inputs can include text, images, video, audio, and PDFs; output is text.
It supports low, medium, and high, with medium as the default. Minimal is not supported. High is intended for the hardest reasoning, math, coding, and agent tasks and may consume more thinking tokens.
Through December 31, 2026, Standard pricing is $0.75 per million input tokens, $3.75 per million output tokens, and $0.075 per million cached input tokens. Starting January 1, 2027, those rates become $1.50, $7.50, and $0.15.
Google reports 65.3% on DeepSWE, 85.8% on Terminal-Bench 2.1, 1588 Elo on Code Arena, 34.0% on GDP.pdf, 97.0% on GDM-MRCR v2 at 128K, 47.9% on OSWorld 2.0, and 53.6% on HLE-Verified.
No. Google has not released the model weights or disclosed total and active parameter counts; access is provided through the Gemini API and Google products.
Follow DataLearner on WeChat for AI model updates and research notes.
