DataLearner logo

Kimi K3vsKimi K2.6

Across 12 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 12, Kimi K2.6 wins 0, with 0 ties and an average score difference of +44.55.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Moonshot AI
Kimi K2.6

Moonshot AI · 2026-04-20 · Reasoning model

Kimi K312 wins(100%)(0%)0 winsKimi K2.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 12 shared benchmarks.

AI Agent - Tool Usage

Kimi K3 4/4
BenchmarkKimi K3Kimi K2.6Diff
Terminal-Bench 2.188.304 / 49Max (With Tools)53.5648 / 49Thinking (No Tools)+34.74
MCPMark-Verified94.501 / 3Max (With Tools)72.803 / 3Thinking (With Tools)+21.70
MCP-Atlas84.204 / 41Max (With Tools)69.4030 / 41Thinking (With Tools)+14.80
OSWorld-Verified84.802 / 26Max (With Tools)73.1015 / 26Thinking (With Tools)+11.70

Coding and Software Engineer

Kimi K3 3/3
BenchmarkKimi K3Kimi K2.6Diff
Program Bench77.801 / 7Max (With Tools)48.304 / 7Thinking (With Tools)+29.50
Kimi Code Bench 2.072.901 / 3Max (With Tools)50.903 / 3Thinking (With Tools)+22
MLS Bench48.302 / 5Max (With Tools)26.705 / 5Thinking (With Tools)+21.60

AI Agent - Information Search

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.6Diff
BrowseComp91.201 / 55Max (With Tools + Internet)83.2014 / 55Thinking (With Tools + Internet)+8

General Evaluation

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.6Diff
GPQA Diamond93.5015 / 270最高(无工具)90.5041 / 270Thinking (No Tools)+3

General Knowledge

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.6Diff
HLE5616 / 189Max (With Tools)5421 / 189Thinking (With Tools + Internet)+2

Text Embedding

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.6Diff
Context Arena71.7559 / 126最高(无工具)51.8886 / 126Normal (No Tools)+19.87

Writing and Creative Capabilities

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.6Diff
Creative Writing2,0712 / 99Normal (No Tools)1,72521 / 99Normal (No Tools)+345.70

Specs

FieldKimi K3Kimi K2.6
PublisherMoonshot AIMoonshot AI
Release date2026-07-162026-04-20
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters2.8T1T
Context length1M256K
Max output1MNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3Kimi K2.6
Text input¥20 / 1M tokens$0.95 / 1M tokens
Text output¥100 / 1M tokens$4 / 1M tokens
Cache read¥2 / 1M tokens$0.16 / 1M tokens
Cache writeNot public$0.95 / 1M tokens

Summary

  • Kimi K3leads in:AI Agent - Tool Usage (4/4), Coding and Software Engineer (3/3), AI Agent - Information Search (1/1), General Evaluation (1/1), General Knowledge (1/1), Text Embedding (1/1), Writing and Creative Capabilities (1/1)

On average across the 12 shared benchmarks, Kimi K3 scores 44.55 higher.

Largest single-benchmark gap: Creative Writing — Kimi K3 2,071 vs Kimi K2.6 1,725 (+345.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.