DataLearner logo

Kimi K3vsKimi K2.7 Code

Across 7 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 7, Kimi K2.7 Code wins 0, with 0 ties and an average score difference of +18.24.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Moonshot AI
Kimi K2.7 Code

Moonshot AI · 2026-06-12 · Coding model

Kimi K37 wins(100%)(0%)0 winsKimi K2.7 Code

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

Coding and Software Engineer

Kimi K3 4/4
BenchmarkKimi K3Kimi K2.7 CodeDiff
DeepSWE67.505 / 26Max (With Tools)3123 / 26Normal (With Tools)+36.50
Program Bench77.801 / 5Max (With Tools)53.603 / 5Thinking (With Tools)+24.20
MLS Bench48.301 / 4Max (With Tools)35.103 / 4Thinking (With Tools)+13.20
Kimi Code Bench 2.072.901 / 3Max (With Tools)622 / 3Thinking (With Tools)+10.90

AI Agent - Tool Usage

Kimi K3 3/3
BenchmarkKimi K3Kimi K2.7 CodeDiff
Terminal-Bench 2.188.302 / 43Max (With Tools)67.0433 / 43Thinking (With Tools)+21.26
MCPMark-Verified94.501 / 3Max (With Tools)81.102 / 3Thinking (With Tools)+13.40
MCP-Atlas84.203 / 38Max (With Tools)7617 / 38Thinking (With Tools)+8.20

Specs

FieldKimi K3Kimi K2.7 Code
PublisherMoonshot AIMoonshot AI
Release date2026-07-162026-06-12
Model typeReasoning modelCoding model
ArchitectureMoEMoE
Parameters2.8T1T
Context length1M256K
Max output1MNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3Kimi K2.7 Code
Text input¥20 / 1M tokens$0.95 / 1M tokens
Text output¥100 / 1M tokens$4 / 1M tokens
Cache read¥2 / 1M tokens$0.19 / 1M tokens

Summary

  • Kimi K3leads in:Coding and Software Engineer (4/4), AI Agent - Tool Usage (3/3)

On average across the 7 shared benchmarks, Kimi K3 scores 18.24 higher.

Largest single-benchmark gap: DeepSWE — Kimi K3 67.50 vs Kimi K2.7 Code 31 (+36.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.