DataLearner logo

Gemini 3.1 Flash-Lite Benchmark Details

Gemini 3.1 Flash-Lite currently shows benchmark results led by SAGE (13 / 64, score 49.54), PinchBench v2 (10 / 45, score 80.50), MedCode (23 / 64, score 47.60). This page also compares it with 2 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Gemini 3.1 Flash-Lite

Benchmark Results

Thinking
Tool usage
Internet

Knowledge Exams

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
unknown
8.64
202 / 235

Scientific Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
74.24
160 / 253
81.82
115 / 253
CritPt
Thinking Enabled
1.10
151 / 204

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
61.68
63 / 117

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Thinking EnabledTools
24.20
134 / 244
Terminal-Bench 4.0
Thinking EnabledTools
0.50
93 / 98

Tool Orchestration

2 evaluations
Benchmark / mode
Score
Rank/total
80.50
10 / 45
MCP-Atlas
HighTools
57.10
39 / 44

Capability Indices

2 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
144.45
79 / 167
Vals Index
HighTools
15.46
42 / 42

Visual Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Thinking Enabled
75.50
83 / 229

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Enabled
43.40
99 / 134

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
SAGE
HighTools
49.54
13 / 64
τ³-Banking
Thinking EnabledTools
9.70
134 / 167

Legal

3 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Thinking EnabledTools
31.15
41 / 44
3.37
42 / 42

Finance

2 evaluations
Benchmark / mode
Score
Rank/total
29.99
42 / 43
8.63
42 / 42

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Enabled
7.80
92 / 122

Data Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
Thinking EnabledToolsInternet
8.75
26 / 29

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
4.61
42 / 43

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
0
59 / 61

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
HighTools
63.90
65 / 66
MedCode
HighTools
47.60
23 / 64

Competitor Comparison

Benchmark scores for Gemini 3.1 Flash-Lite compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.1 Flash-LiteCurrentDeepSeek-V4-FlashQwen3.6-35B-A3B
HLE
Accuracy
Knowledge Exams
8.64Thinking Level · High
51.50Thinking Level · High | Tools
21.40Thinking Enabled
CritPt
Score
Scientific Reasoning
1.10Thinking Enabled
7.10Thinking Level · High
0.30Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
81.82Thinking Level · High
71.20Standard Mode
84.85Standard Mode
LiveBench
Accuracy
Cross-capability Suites
61.68Thinking Level · High
65.48Standard Mode
--
Terminal Bench Hard
Accuracy
Agentic Development
24.20Thinking Enabled | Tools
38.60Thinking Level · High | Tools
34.80Thinking Enabled | Tools
Terminal-Bench 4.0
Resolution rate (%)
Agentic Development
0.50Thinking Enabled | Tools
3.00Thinking Level · High | Tools
--
PinchBench v2
Average score (%)
Tool Orchestration
80.50Thinking Level · High
81.74Thinking Level · High
--
ECI
ECI score (capability index, higher is better)
Capability Indices
144.45Thinking Level · High
146.10Thinking Level · High
143.92Thinking Level · High
MMMU-Pro
Accuracy
Visual Understanding
75.50Thinking Enabled
--
75.00Thinking Enabled
SciCode
Score
Scientific Computing
43.40Thinking Enabled
45.30Thinking Level · High
36.60Thinking Enabled
τ³-Banking
Score
Service Workflows
9.70Thinking Enabled | Tools
30.90Thinking Level · High | Tools
9.30Thinking Enabled | Tools
Harvey Lab-AA
Criterion pass rate (some entries report all-pass rate)
Legal
31.15Thinking Enabled | Tools
81.33Thinking Level · High | Tools
--
2 additional benchmarks remain in the chart above.

Standard API Pricing: Gemini 3.1 Flash-Lite vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.1 Flash-Lite
Google DeepMind$0.25 / 1M tokens$1.5 / 1M tokens—
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens—

Version History

How each version of the Gemini 3.1 Flash-Lite series stacks up on benchmark tests

Gemini 3.1 Flash-LiteGemini 2.5 Flash-Lite
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

5 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.1 Flash-LiteCurrentGemini 2.5 Flash-Lite
HLE
Accuracy
Knowledge Exams
8.64Thinking Level · High
6.90Standard Mode
GPQA Diamond
Accuracy
Scientific Reasoning
81.82Thinking Level · High
66.70Standard Mode
LiveBench
Accuracy
Cross-capability Suites
61.68Thinking Level · High
42.56Thinking Level · High
Terminal Bench Hard
Accuracy
Agentic Development
24.20Thinking Enabled | Tools
4.50Thinking Enabled | Tools
MMMU-Pro
Accuracy
Visual Understanding
75.50Thinking Enabled
58.20Thinking Enabled

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.1 Flash-Lite Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.1 Flash-Lite
Google DeepMind$0.25 / 1M tokens$1.5 / 1M tokens—
Gemini 2.5 Flash-Lite
Google DeepMind$0.1 / 1M tokens$0.4 / 1M tokens—