GLM-5.3-FlashvsQwen3.8-Flash-Next
Across 6 shared benchmarks, GLM-5.3-Flash leads overall: GLM-5.3-Flash wins 5, Qwen3.8-Flash-Next wins 1, with 0 ties and an average score difference of +6.33.
GLM-5.3-Flash
智谱AI · 2026-08-26 · Multimodal model
Qwen3.8-Flash-Next
阿里巴巴 · 2026-08-26 · Reasoning model
GLM-5.3-Flash5 wins(83%)(17%)1 winQwen3.8-Flash-Next
Benchmark scores
Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.
Coding and Software Engineer
GLM-5.3-Flash 2/2| Benchmark | GLM-5.3-Flash | Qwen3.8-Flash-Next | Diff |
|---|---|---|---|
| NL2Repo-Bench | 56.304 / 10Max (With Tools) | 48.108 / 10极高强度思考(工具) | +8.20 |
| DeepSWE | 63.4011 / 30Max (With Tools) | 58.7016 / 30极高强度思考(工具) | +4.70 |
Agent Level Benchmark
GLM-5.3-Flash 1/1| Benchmark | GLM-5.3-Flash | Qwen3.8-Flash-Next | Diff |
|---|---|---|---|
| Agents' Last Exam | 26.308 / 13Max (With Tools) | 24.3012 / 13极高强度思考(工具) | +2 |
AI Agent - Tool Usage
GLM-5.3-Flash 1/1| Benchmark | GLM-5.3-Flash | Qwen3.8-Flash-Next | Diff |
|---|---|---|---|
| Toolathlon-Verified | 78.401 / 7Max (With Tools) | 73.504 / 7极高强度思考(工具) | +4.90 |
General Knowledge
GLM-5.3-Flash 1/1| Benchmark | GLM-5.3-Flash | Qwen3.8-Flash-Next | Diff |
|---|---|---|---|
| HLE | 55.3015 / 183Max (With Tools) | 35.9078 / 183极高强度思考(无工具) | +19.40 |
Multimodal Understanding
Qwen3.8-Flash-Next 1/1| Benchmark | GLM-5.3-Flash | Qwen3.8-Flash-Next | Diff |
|---|---|---|---|
| CharXiv RQ | 89.404 / 18Max (With Tools) | 90.602 / 18极高强度思考(工具) | -1.20 |
Specs
| Field | GLM-5.3-Flash | Qwen3.8-Flash-Next |
|---|---|---|
| Publisher | 智谱AI | 阿里巴巴 |
| Release date | 2026-08-26 | 2026-08-26 |
| Model type | Multimodal model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 320B | 125B |
| Context length | 1M | 262144 |
| Max output | 128K | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | GLM-5.3-Flash | Qwen3.8-Flash-Next |
|---|---|---|
| Text input | $0.075 / 1M tokens | Not public |
| Text output | $0.25 / 1M tokens | Not public |
| Cache read | $0.015 / 1M tokens | Not public |
One or both models have incomplete public pricing.
Summary
- GLM-5.3-Flashleads in:Coding and Software Engineer (2/2), Agent Level Benchmark (1/1), AI Agent - Tool Usage (1/1), General Knowledge (1/1)
- Qwen3.8-Flash-Nextleads in:Multimodal Understanding (1/1)
On average across the 6 shared benchmarks, GLM-5.3-Flash scores 6.33 higher.
Largest single-benchmark gap: HLE — GLM-5.3-Flash 55.30 vs Qwen3.8-Flash-Next 35.90 (+19.40).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.