DataLearner logo

Qwen3.8-Max-0902vsKimi K3

Across 9 shared benchmarks, Qwen3.8-Max-0902 leads overall: Qwen3.8-Max-0902 wins 6, Kimi K3 wins 3, with 0 ties and an average score difference of -1.06.

阿里巴巴
Qwen3.8-Max-0902

阿里巴巴 · 2026-09-02 · Reasoning model

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Qwen3.8-Max-09026 wins(67%)(33%)3 winsKimi K3

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Coding and Software Engineer

Qwen3.8-Max-0902 3/4
BenchmarkQwen3.8-Max-0902Kimi K3Diff
Program Bench285 / 7极高强度思考(工具)77.801 / 7Max (With Tools)-49.80
SWE-Marathon44.801 / 6极高强度思考(工具)423 / 6Max (With Tools)+2.80
DeepSWE69.304 / 32极高强度思考(工具)67.506 / 32Max (With Tools)+1.80
MLS Bench50.101 / 5极高强度思考(工具)48.302 / 5Max (With Tools)+1.80

AI Agent - Tool Usage

Even 2/2
BenchmarkQwen3.8-Max-0902Kimi K3Diff
AutomationBench50.801 / 12极高强度思考(工具)30.807 / 12Max (With Tools)+20
Toolathlon-Verified73.307 / 10极高强度思考(工具)76.502 / 10Max (With Tools)-3.20

Multimodal Understanding

Even 2/2
BenchmarkQwen3.8-Max-0902Kimi K3Diff
BabyVision93.801 / 6极高强度思考(工具)85.702 / 6Max (With Tools)+8.10
MMMU-Pro82.702 / 8极高强度思考(无工具)83.401 / 8Max (With Tools)-0.70

Agent Level Benchmark

Qwen3.8-Max-0902 1/1
BenchmarkQwen3.8-Max-0902Kimi K3Diff
Job Bench641 / 4极高强度思考(工具)54.304 / 4Max (With Tools)+9.70

Specs

FieldQwen3.8-Max-0902Kimi K3
Publisher阿里巴巴Moonshot AI
Release date2026-09-022026-07-16
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters2.4T2.8T
Context length1M1M
Max output131K1M

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.8-Max-0902Kimi K3
Text input$2 / 1M tokens¥20 / 1M tokens
Text output$6 / 1M tokens¥100 / 1M tokens
Cache read$0.25 / 1M tokens¥2 / 1M tokens
Cache write$2.5 / 1M tokensNot public

Summary

  • Qwen3.8-Max-0902leads in:Coding and Software Engineer (3/4), Agent Level Benchmark (1/1)
  • Tied in:AI Agent - Tool Usage, Multimodal Understanding

On average across the 9 shared benchmarks, Kimi K3 scores 1.06 higher.

Largest single-benchmark gap: Program Bench — Qwen3.8-Max-0902 28 vs Kimi K3 77.80 (-49.80).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.