DataLearner logo

Gemini 3.5 Flash-Lite Benchmark Details

Gemini 3.5 Flash-Lite currently shows benchmark results led by Creative Writing (41 / 99, score 1556), GPQA Diamond (114 / 270, score 83.33), Context Arena (57 / 126, score 72.57). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Gemini 3.5 Flash-Lite

Benchmark Results

Thinking
Tool usage

Other

2 evaluations
Benchmark / mode
Score
Rank/total
75.76
166 / 270
83.33
114 / 270

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1556
41 / 99

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Thinking EnabledTools
54.20
36 / 60

Text Embedding

3 evaluations
Benchmark / mode
Score
Rank/total
61.90
75 / 126
65.50
72 / 126
72.57
57 / 126

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking EnabledTools
74
14 / 26
Terminal-Bench 2.1
Thinking EnabledTools
54
47 / 49

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
AA-Briefcase
Thinking EnabledTools
636.95
17 / 19

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
25.96
46 / 58

Competitor Comparison

Benchmark scores for Gemini 3.5 Flash-Lite compared against top models in its class

Gemini 3.5 Flash-LiteGPT-5.6 LunaDeepSeek-V4-Flash
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

7 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGemini 3.5 Flash-LiteCurrentGPT-5.6 LunaDeepSeek-V4-Flash
GPQA Diamond
Accuracy
科学与综合推理
83.33Thinking Level · High
91.60Thinking Level · High
88.10Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1556.00Standard Mode
1826.60Standard Mode
1555.70Standard Mode
SWE-Bench Pro - Public
Accuracy
编程与软件工程
54.20Thinking Enabled | Tools
--
52.60Thinking Level · Extra High | Tools
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
72.57Thinking Level · High
81.80Thinking Level · High
69.42Thinking Enabled
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
54.00Thinking Enabled | Tools
84.70Thinking Level · High
82.70Thinking Level · High | Tools
AA-Briefcase
Elo score
生产力知识
636.95Thinking Enabled | Tools
1359.76Thinking Level · High | Tools
--
FrontierMath v2
Accuracy (verification_code)
数学推理
25.96Thinking Level · High
82.11Thinking Level · High
--

Standard API Pricing: Gemini 3.5 Flash-Lite vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.5 Flash-Lite
DeepMind$0.3 / 1M tokens$2.5 / 1M tokens
GPT-5.6 Luna
OpenAI$0.2 / 1M tokens$1.2 / 1M tokens
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens

Version History

How each version of the Gemini 3.5 Flash-Lite series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

1 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGemini 3.5 Flash-LiteCurrentGemini 3.1 Flash-LiteGemini 2.5 Flash-Lite
GPQA Diamond
Accuracy
科学与综合推理
83.33Thinking Level · High
81.82Thinking Level · High
66.70Standard Mode

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 科学与综合推理

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Gemini 3.5 Flash-Lite Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Gemini 3.5 Flash-Lite
DeepMind$0.3 / 1M tokens$2.5 / 1M tokens
Gemini 2.5 Flash-Lite
Google Deep Mind$0.1 / 1M tokens$0.4 / 1M tokens