DataLearner logo

Kimi K3vsKimi K2.5

Across 9 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 9, Kimi K2.5 wins 0, with 0 ties and an average score difference of +94.49.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · Multimodal model

Kimi K39 wins(100%)(0%)0 winsKimi K2.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

AI Agent - Information Search

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
BrowseComp91.202 / 56Max (With Tools + Internet)60.6037 / 56Thinking (With Tools + Internet)+30.60

AI Agent - Tool Usage

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
MCP-Atlas84.204 / 41Max (With Tools)64.4032 / 41Normal (With Tools)+19.80

Coding and Software Engineer

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
Text Arena (Coding)1,6822 / 35Max (No Tools)1,43127 / 35Normal (No Tools)+251.19

Commonsense Reasoning

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
SimpleBench60.7029 / 92Max (No Tools)46.8052 / 92Thinking (No Tools)+13.90

General Evaluation

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
GPQA Diamond93.5016 / 271Max (No Tools)87.6076 / 271Thinking (No Tools)+5.90

General Knowledge

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
HLE5617 / 190Max (With Tools)50.2034 / 190Thinking (With Tools)+5.80

Long Context

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
AA-LCR74.708 / 28Max (No Tools)6522 / 28Thinking (No Tools)+9.70

Text Embedding

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
Context Arena71.7559 / 126Max (No Tools)53.3383 / 126Normal (No Tools)+18.42

Writing and Creative Capabilities

Kimi K3 1/1
BenchmarkKimi K3Kimi K2.5Diff
Creative Writing2,0712 / 99Normal (No Tools)1,57639 / 99Normal (No Tools)+495.10

Specs

FieldKimi K3Kimi K2.5
PublisherMoonshot AIMoonshot AI
Release date2026-07-162026-01-27
Model typeReasoning modelMultimodal model
ArchitectureMoEMoE
Parameters2.8T1T
Context length1M256K
Max output1M16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3Kimi K2.5
Text input¥20 / 1M tokens$0.6 / 1M tokens
Text output¥100 / 1M tokens$3 / 1M tokens
Cache read¥2 / 1M tokens$0.1 / 1M tokens

Summary

  • Kimi K3leads in:AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1), Commonsense Reasoning (1/1), General Evaluation (1/1), General Knowledge (1/1), Long Context (1/1), Text Embedding (1/1), Writing and Creative Capabilities (1/1)

On average across the 9 shared benchmarks, Kimi K3 scores 94.49 higher.

Largest single-benchmark gap: Creative Writing — Kimi K3 2,071 vs Kimi K2.5 1,576 (+495.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.