DataLearner logo

Qwen3.7-PlusvsKimi K2.6

Across 7 shared benchmarks, Kimi K2.6 leads overall: Qwen3.7-Plus wins 2, Kimi K2.6 wins 5, with 0 ties and an average score difference of -2.16.

阿里巴巴
Qwen3.7-Plus

阿里巴巴 · 2026-05-31 · Reasoning model

Moonshot AI
Kimi K2.6

Moonshot AI · 2026-04-20 · Reasoning model

Qwen3.7-Plus2 wins(29%)(71%)5 winsKimi K2.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

Multimodal Understanding

Even 2/2
BenchmarkQwen3.7-PlusKimi K2.6Diff
MMMU-Pro80.5037 / 227Thinking (No Tools)79.4049 / 227Thinking (No Tools)+1.10
GDP.pdf12.2070 / 118Thinking (No Tools)1365 / 118Thinking (No Tools)-0.80

Agent Level Benchmark

Kimi K2.6 1/1
BenchmarkQwen3.7-PlusKimi K2.6Diff
τ³-Banking17.5098 / 164Thinking (With Tools)23.3082 / 164Thinking (With Tools)-5.80

AI Agent - Tool Usage

Kimi K2.6 1/1
BenchmarkQwen3.7-PlusKimi K2.6Diff
Terminal-Bench 2.161108 / 192Thinking (With Tools)65.9094 / 192Thinking (With Tools)-4.90

Coding and Software Engineer

Kimi K2.6 1/1
BenchmarkQwen3.7-PlusKimi K2.6Diff
SciCode46.1084 / 130Thinking (No Tools)51.5061 / 130Thinking (No Tools)-5.40

General Evaluation

Qwen3.7-Plus 1/1
BenchmarkQwen3.7-PlusKimi K2.6Diff
GPQA Diamond81.82207 / 462Normal (No Tools)78.80249 / 462Normal (No Tools)+3.02

Productivity Knowledge

Kimi K2.6 1/1
BenchmarkQwen3.7-PlusKimi K2.6Diff
Harvey Lab-AA81.7929 / 43Thinking (With Tools)84.1225 / 43Thinking (With Tools)-2.33

Specs

FieldQwen3.7-PlusKimi K2.6
Publisher阿里巴巴Moonshot AI
Release date2026-05-312026-04-20
Model typeReasoning modelReasoning model
ArchitectureDenseMoE
ParametersNot available1T
Context length1M256K
Max output64KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen3.7-PlusKimi K2.6
Text input¥2 / 1M tokens$0.95 / 1M tokens
Text output¥8 / 1M tokens$4 / 1M tokens
Cache readNot public$0.16 / 1M tokens
Cache writeNot public$0.95 / 1M tokens

Summary

  • Qwen3.7-Plusleads in:General Evaluation (1/1)
  • Kimi K2.6leads in:Agent Level Benchmark (1/1), AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1), Productivity Knowledge (1/1)
  • Tied in:Multimodal Understanding

On average across the 7 shared benchmarks, Kimi K2.6 scores 2.16 higher.

Largest single-benchmark gap: τ³-Banking — Qwen3.7-Plus 17.50 vs Kimi K2.6 23.30 (-5.80).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.