DataLearner logo

Kimi K3vsKimi K2.6

Across 4 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 4, Kimi K2.6 wins 0, with 0 ties and an average score difference of +11.94.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Moonshot AI
Kimi K2.6

Moonshot AI · 2026-04-20 · Reasoning model

Kimi K34 wins(100%)(0%)0 winsKimi K2.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.

General Knowledge

Kimi K3 2/2
BenchmarkKimi K3Kimi K2.6Diff
GPQA Diamond93.508 / 187最高(无工具)90.5018 / 187Thinking (No Tools)+3
HLE5610 / 170Max (With Tools)5413 / 170Thinking (With Tools + Internet)+2

AI Agent - Information Search

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.6Diff
BrowseComp91.201 / 52Max (With Tools + Internet)83.2013 / 52Thinking (With Tools + Internet)+8

AI Agent - Tool Usage

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.6Diff
TerminalBench 2.188.302 / 25Max (With Tools)53.5625 / 25Thinking (No Tools)+34.74

Specs

FieldKimi K3Kimi K2.6
PublisherMoonshot AIMoonshot AI
Release date2026-07-162026-04-20
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters2.8T1T
Context length1M256K
Max output1MNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3Kimi K2.6
Text input¥20 / 1M tokens$0.95 / 1M tokens
Text output¥100 / 1M tokens$4 / 1M tokens
Cache read¥2 / 1M tokens$0.16 / 1M tokens
Cache writeNot public$0.95 / 1M tokens

Summary

  • Kimi K3leads in:General Knowledge (2/2), AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1)

On average across the 4 shared benchmarks, Kimi K3 scores 11.94 higher.

Largest single-benchmark gap: TerminalBench 2.1 — Kimi K3 88.30 vs Kimi K2.6 53.56 (+34.74).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.