Gemini 3.8 Flash leads on the shared-benchmark average
Ahead on 2 of 3 shared benchmarks, averaging 1.2 points higher
Summarised only from the 3 percentage-scale benchmarks scored by every selected model; details are below. A further 1 Elo/rating-scale benchmarks are left out of the average — their scale cannot be added to percentages.
“Best available” takes each model’s highest recorded non-parallel mode per benchmark, so it may combine modes into a virtual configuration that does not exist. Read it with the mode breakdown.

Gemini 3.8 Flash
Google Deep Mind
Benchmark-by-benchmark comparison. Changing the thinking mode or tool filters updates the chart and table below.
“Best available” picks the highest non-parallel mode separately for each benchmark. The resulting series can combine several reasoning levels and is not one reproducible runtime configuration. Choose a mode filter to compare like-for-like runs.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
Every model and runtime mode, benchmark by benchmark. Values are comparable along a row, not between different benchmarks.
4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Gemini 3.8 Flash | GPT-5.6 Terra |
|---|---|---|
Terminal-Bench 2.1 AI Agent - 工具使用 | 89.40Thinking Enabled | Tools | 87.40Thinking Level · High |
Terminal-Bench 4.0 AI Agent - 工具使用 | 19.10Thinking Enabled | Tools | 21.52Thinking Level · High | Tools |
DeepSWE 编程与软件工程 | 73.70Thinking Enabled | Tools | 69.60Thinking Level · Extra High | Tools |
GDPval-AA v2 生产力知识 | 1545.00Thinking Enabled | 1565.62Thinking Level · High | Tools |
Official list prices per model API, split by input and output. Unit: USD per 1M tokens.
Architecture, licensing and API modalities. "Not provided" means the field is missing from our database.
| Features & specs | Gemini 3.8 FlashGoogle Deep Mind | GPT-5.6 TerraOpenAI |
|---|---|---|
Core specsRelease | 2026-09-02 | 2026-06-26 |
Context length | 1M | 1.05M |
Max output length | 65,536 tokens | 128,000 tokens |
Architecture | Undisclosed | Undisclosed |
Runtime modes | 低中高 | 关闭低中高 |
Availability & licensingCode availability | Not public | Not public |
Weight availability | Not public | Not public |
Use & commercial terms | Official service only; subject to provider terms | Official service only; subject to provider terms |
API modality supportText Input/Output | Input:YesOutput:Yes | Input:YesOutput:Yes |
Image Input/Output | Input:YesOutput:No | Input:YesOutput:No |
Audio Input/Output | Input:YesOutput:No | Input:NoOutput:No |
Video Input/Output | Input:YesOutput:No | Input:NoOutput:No |
ResourcesPaper / report | Gemini 3.8 Flash evaluation methodology | Previewing GPT-5.6 Sol: a next-generation model |

GPT-5.6 Terra
OpenAI