DataLearner logo

Qwen3.8-MaxvsKimi K3

Across 11 shared benchmarks, Kimi K3 leads overall: Qwen3.8-Max wins 3, Kimi K3 wins 8, with 0 ties and an average score difference of -2.48.

阿里巴巴
Qwen3.8-Max

阿里巴巴 · 2026-08-03 · Reasoning model

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Qwen3.8-Max3 wins(27%)(73%)8 winsKimi K3

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

AI Agent - Tool Usage

Kimi K3 3/3
BenchmarkQwen3.8-MaxKimi K3Diff
Toolathlon-Verified72.504 / 5极高强度思考(工具)76.501 / 5Max (With Tools)-4
AutomationBench27.305 / 8极高强度思考(工具)30.803 / 8Max (With Tools)-3.50
Terminal-Bench 2.186.608 / 44极高强度思考(工具)88.302 / 44Max (With Tools)-1.70

Coding and Software Engineer

Kimi K3 3/3
BenchmarkQwen3.8-MaxKimi K3Diff
DeepSWE56.6014 / 27极高强度思考(工具)67.505 / 27Max (With Tools)-10.90
FrontierSWE73.504 / 4极高强度思考(工具)81.201 / 4Max (With Tools)-7.70
MLS Bench412 / 4极高强度思考(工具)48.301 / 4Max (With Tools)-7.30

Math and Reasoning

Qwen3.8-Max 2/2
BenchmarkQwen3.8-MaxKimi K3Diff
FrontierMath Tier 4 v246.3411 / 34极高强度思考(无工具)39.0213 / 34最高(无工具)+7.32
FrontierMath v274.7411 / 34极高强度思考(无工具)72.1813 / 34最高(无工具)+2.55

Agent Level Benchmark

Kimi K3 1/1
BenchmarkQwen3.8-MaxKimi K3Diff
Agents' Last Exam277 / 11极高强度思考(工具)28.305 / 11Max (With Tools)-1.30

General Evaluation

Kimi K3 1/1
BenchmarkQwen3.8-MaxKimi K3Diff
GPQA Diamond92.6021 / 226极高强度思考(无工具)93.5013 / 226最高(无工具)-0.90

General Knowledge

Qwen3.8-Max 1/1
BenchmarkQwen3.8-MaxKimi K3Diff
HLE56.2013 / 181极高强度思考(工具)5614 / 181Max (With Tools)+0.20

Specs

FieldQwen3.8-MaxKimi K3
Publisher阿里巴巴Moonshot AI
Release date2026-08-032026-07-16
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters2.4T2.8T
Context length1M1M
Max output128K1M

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-MaxKimi K3
Text input¥12 / 1M tokens¥20 / 1M tokens
Text output¥36 / 1M tokens¥100 / 1M tokens
Cache read¥1.5 / 1M tokens¥2 / 1M tokens

Summary

  • Qwen3.8-Maxleads in:Math and Reasoning (2/2), General Knowledge (1/1)
  • Kimi K3leads in:AI Agent - Tool Usage (3/3), Coding and Software Engineer (3/3), Agent Level Benchmark (1/1), General Evaluation (1/1)

On average across the 11 shared benchmarks, Kimi K3 scores 2.48 higher.

Largest single-benchmark gap: DeepSWE — Qwen3.8-Max 56.60 vs Kimi K3 67.50 (-10.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.