DataLearner logo

GLM-5.3-FlashvsGemini 3.7 Flash

Across 6 shared benchmarks, GLM-5.3-Flash leads overall: GLM-5.3-Flash wins 3, Gemini 3.7 Flash wins 2, with 1 ties and an average score difference of +43.95.

智谱AI
GLM-5.3-Flash

智谱AI · 2026-08-26 · Multimodal model

Google Deep Mind
Gemini 3.7 Flash

Google Deep Mind · 2026-08-13 · Multimodal model

GLM-5.3-Flash3 wins(50%)Ties1(33%)2 winsGemini 3.7 Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

AI Agent - Tool Usage

Even 2/2
BenchmarkGLM-5.3-FlashGemini 3.7 FlashDiff
AutomationBench48.801 / 9Max (With Tools)30.405 / 9Thinking (With Tools)+18.40
Terminal-Bench 2.184.3011 / 46Max (With Tools)85.809 / 46Thinking (With Tools)-1.50

Agent Level Benchmark

Even 1/1
BenchmarkGLM-5.3-FlashGemini 3.7 FlashDiff
Agents' Last Exam26.308 / 13Max (With Tools)26.308 / 13Thinking Medium (With Tools)

Coding and Software Engineer

Gemini 3.7 Flash 1/1
BenchmarkGLM-5.3-FlashGemini 3.7 FlashDiff
DeepSWE63.4011 / 30Max (With Tools)65.3010 / 30Thinking High (With Tools)-1.90

Multimodal Understanding

GLM-5.3-Flash 1/1
BenchmarkGLM-5.3-FlashGemini 3.7 FlashDiff
CharXiv RQ89.404 / 18Max (With Tools)88.706 / 18Thinking Medium (With Tools)+0.70

Productivity Knowledge

GLM-5.3-Flash 1/1
BenchmarkGLM-5.3-FlashGemini 3.7 FlashDiff
GDPval-AA v21,7732 / 15Max (With Tools)1,52510 / 15Thinking (No Tools)+248

Specs

FieldGLM-5.3-FlashGemini 3.7 Flash
Publisher智谱AIGoogle Deep Mind
Release date2026-08-262026-08-13
Model typeMultimodal modelMultimodal model
ArchitectureMoEDense
Parameters320BNot available
Context length1M1M
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5.3-FlashGemini 3.7 Flash
Text input$0.075 / 1M tokens$0.75 / 1M tokens
Text output$0.25 / 1M tokens$3.75 / 1M tokens
Cache read$0.015 / 1M tokens$0.075 / 1M tokens

Summary

  • GLM-5.3-Flashleads in:Multimodal Understanding (1/1), Productivity Knowledge (1/1)
  • Gemini 3.7 Flashleads in:Coding and Software Engineer (1/1)
  • Tied in:AI Agent - Tool Usage, Agent Level Benchmark

On average across the 6 shared benchmarks, GLM-5.3-Flash scores 43.95 higher.

Largest single-benchmark gap: GDPval-AA v2 — GLM-5.3-Flash 1,773 vs Gemini 3.7 Flash 1,525 (+248).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.