DataLearner logo

Qwen 3.6 Plus PreviewvsKimi K2.5

Across 11 shared benchmarks, Qwen 3.6 Plus Preview leads overall: Qwen 3.6 Plus Preview wins 11, Kimi K2.5 wins 0, with 0 ties and an average score difference of +6.52.

阿里巴巴
Qwen 3.6 Plus Preview

阿里巴巴 · 2026-03-31 · Chat model

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · Multimodal model

Qwen 3.6 Plus Preview11 wins(100%)(0%)0 winsKimi K2.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

Coding and Software Engineer

Qwen 3.6 Plus Preview 3/3
BenchmarkQwen 3.6 Plus PreviewKimi K2.5Diff
LiveCodeBench87.1019 / 250Thinking (No Tools)8529 / 250Thinking (No Tools)+2.10
SWE-bench Verified78.8021 / 116Thinking (With Tools)76.8030 / 116Thinking (With Tools)+2
SWE-bench Multilingual73.8014 / 30Thinking (No Tools)7318 / 30Thinking (No Tools)+0.80

AI Agent - Tool Usage

Qwen 3.6 Plus Preview 2/2
BenchmarkQwen 3.6 Plus PreviewKimi K2.5Diff
Terminal-Bench 2.161.40107 / 192Thinking (With Tools)45.70137 / 192Thinking (With Tools)+15.70
Terminal Bench 2.061.6016 / 48Thinking (With Tools)50.8035 / 48Thinking (With Tools)+10.80

Math and Reasoning

Qwen 3.6 Plus Preview 2/2
BenchmarkQwen 3.6 Plus PreviewKimi K2.5Diff
AIME 202695.3012 / 29Thinking (No Tools)92.5021 / 29Thinking (No Tools)+2.80
IMO-AnswerBench83.8014 / 24Thinking (No Tools)81.8018 / 24Thinking (No Tools)+2

Agent Level Benchmark

Qwen 3.6 Plus Preview 1/1
BenchmarkQwen 3.6 Plus PreviewKimi K2.5Diff
τ³-Banking20.8088 / 164Thinking (With Tools)14.20113 / 164Thinking (With Tools)+6.60

Claw-style Agent Evaluation

Qwen 3.6 Plus Preview 1/1
BenchmarkQwen 3.6 Plus PreviewKimi K2.5Diff
PinchBench v272.5023 / 45Reported best (effort unspecified)54.6036 / 45Reported best (effort unspecified)+17.90

General Knowledge

Qwen 3.6 Plus Preview 1/1
BenchmarkQwen 3.6 Plus PreviewKimi K2.5Diff
MMLU Pro88.506 / 176Thinking (No Tools)78.5071 / 176Thinking (No Tools)+10

Long Context

Qwen 3.6 Plus Preview 1/1
BenchmarkQwen 3.6 Plus PreviewKimi K2.5Diff
LongBench v2624 / 14Normal (No Tools)616 / 14Normal (No Tools)+1

Specs

FieldQwen 3.6 Plus PreviewKimi K2.5
Publisher阿里巴巴Moonshot AI
Release date2026-03-312026-01-27
Model typeChat modelMultimodal model
ArchitectureDenseMoE
ParametersNot available1T
Context length1M256K
Max output64K16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemQwen 3.6 Plus PreviewKimi K2.5
Text input$0.5 / 1M tokens$0.6 / 1M tokens
Text output$3 / 1M tokens$3 / 1M tokens
Cache read$0.05 / 1M tokens$0.1 / 1M tokens
Cache write$0.625 / 1M tokensNot public

Summary

  • Qwen 3.6 Plus Previewleads in:Coding and Software Engineer (3/3), AI Agent - Tool Usage (2/2), Math and Reasoning (2/2), Agent Level Benchmark (1/1), Claw-style Agent Evaluation (1/1), General Knowledge (1/1), Long Context (1/1)

On average across the 11 shared benchmarks, Qwen 3.6 Plus Preview scores 6.52 higher.

Largest single-benchmark gap: PinchBench v2 — Qwen 3.6 Plus Preview 72.50 vs Kimi K2.5 54.60 (+17.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.