DataLearner logo

Kimi K3 Benchmark Details

Kimi K3 currently shows benchmark results led by BrowseComp (1 / 54, score 91.20), Creative Writing (2 / 99, score 2070.80), Terminal-Bench 2.1 (2 / 47, score 88.30). This page also compares it with 4 competitor models and 4 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Kimi K3

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
HLE
Max
43.50
52 / 185
HLE
MaxTools
56
14 / 185
23.40
1 / 3

Other

3 evaluations
Benchmark / mode
Score
Rank/total
84.85
100 / 270
91.92
28 / 270
93.50
15 / 270

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
2070.80
2 / 99

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
DeepSearchQA
MaxToolsInternet
95
1 / 2
BrowseComp
MaxToolsInternet
91.20
1 / 54

Text Embedding

3 evaluations
Benchmark / mode
Score
Rank/total
58.20
78 / 126
65.97
71 / 126
71.75
59 / 126

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
74.70
8 / 27

AI Agent - Tool Usage

9 evaluations
Benchmark / mode
Score
Rank/total
94.50
1 / 3
88.30
2 / 47
84.80
2 / 26
MCP-Atlas
MaxTools
84.20
4 / 40
76.50
2 / 8
SaaS-Bench
MaxTools
60.10
1 / 1
OSWorld 2.0
MaxTools
58.30
3 / 5
30.80
5 / 10

Coding and Software Engineer

9 evaluations
Benchmark / mode
Score
Rank/total
1681.75
2 / 35
FrontierSWE
MaxTools
81.20
1 / 4
77.80
1 / 6
72.90
1 / 3
DeepSWE
MaxTools
67.50
5 / 31
SciCode
MaxTools
58.70
3 / 14
MLS Bench
MaxTools
48.30
1 / 4
SWE-Marathon
MaxTools
42
2 / 5
36.60
3 / 5

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total
Job Bench
MaxTools
54.30
3 / 3
APEX-Agents
MaxTools
41
5 / 6
τ³-Banking
MaxTools
33.40
6 / 12
28.30
5 / 14

Productivity Knowledge

10 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
MaxTools
1686
7 / 22
AA-Briefcase
MaxTools
1528.20
7 / 19
94.60
1 / 6
ResearchRubrics
MaxToolsInternet
76.20
1 / 1
CorpFin v2
MaxTools
71.60
1 / 1
63.30
2 / 3
54.40
1 / 1
44.20
1 / 1
AA-AnalystAgent
MaxToolsInternet
38.75
7 / 12
34.80
1 / 1

Multimodal Understanding

14 evaluations
Benchmark / mode
Score
Rank/total
94.30
5 / 12
MathVision
MaxTools
97.80
1 / 12
84.80
11 / 18
CharXiv RQ
MaxTools
91.30
1 / 18
91.10
1 / 3
BabyVision
MaxTools
85.70
1 / 5
81.60
3 / 7
MMMU-Pro
MaxTools
83.40
1 / 7
MMVU
Max
82.10
1 / 2
58.50
1 / 1
23
3 / 3
41
1 / 3

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
72.18
13 / 58
39.02
13 / 40

Competitor Comparison

Benchmark scores for Kimi K3 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkKimi K3CurrentGLM-5.2MiniMax M3Claude Opus 4.8GPT-5.6 Sol
HLE
综合评估
56.00Thinking Level · High | Tools
54.70Thinking Enabled | Tools
--
57.90Extended Thinking | Tools
--
GPQA Diamond
科学与综合推理
93.50Thinking Level · High
91.86Thinking Level · High
81.31Standard Mode
93.60Thinking Level · High
93.50Thinking Level · High
Creative Writing
写作和创作
2070.80Standard Mode
1750.90Standard Mode
--
1835.30Standard Mode
1964.10Standard Mode
BrowseComp
AI Agent - 信息收集
91.20Thinking Level · High | Tools
--
83.50Thinking Enabled | Tools
84.30Thinking Level · High | Tools
--
Context Arena
文本向量检索
71.75Thinking Level · High
72.34Thinking Level · High
51.15Thinking Enabled
90.04Thinking Level · High
97.63Thinking Level · High
AA-LCR
长上下文能力
74.70Thinking Level · High
--
80.33Thinking Enabled
--
77.67Thinking Level · High
MCP-Atlas
AI Agent - 工具使用
84.20Thinking Level · High | Tools
76.80Thinking Enabled | Tools
74.20Thinking Enabled | Tools
82.20Thinking Level · High | Tools
--
OSWorld 2.0
AI Agent - 工具使用
58.30Thinking Level · High | Tools
--
--
--
62.60Thinking Level · Extra High | Tools
OSWorld-Verified
AI Agent - 工具使用
84.80Thinking Level · High | Tools
--
70.00Thinking Enabled | Tools
83.40Extended Thinking | Tools
--
Terminal-Bench 2.1
AI Agent - 工具使用
88.30Thinking Level · High | Tools
81.00Thinking Level · High | Tools
66.00Thinking Enabled | Tools
78.90Thinking Level · High | Tools
88.80Thinking Level · High
Terminal-Bench-Science 0.1
AI Agent - 工具使用
7.10Thinking Level · High | Tools
--
--
10.50Thinking Level · High | Tools
22.40Thinking Level · High | Tools
DeepSWE
编程与软件工程
67.50Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
--
59.00Deep Thinking Mode | Tools
72.70Thinking Level · Extra High | Tools
15 additional benchmarks remain in the chart above.

Standard API Pricing: Kimi K3 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
Claude Opus 4.8
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
GPT-5.6 Sol
Supplier: OpenAI
Standard input: $4 / 1M tokens
Standard output: $20 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens

Version History

How each version of the Kimi K3 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkKimi K3CurrentKimi K2.7 CodeKimi K2.6Kimi K2.5Kimi K2 Thinking
HLE
综合评估
56.00Thinking Level · High | Tools
--
54.00Thinking Enabled | Tools
50.20Thinking Enabled | Tools
51.00Thinking Enabled | Tools
GPQA Diamond
科学与综合推理
93.50Thinking Level · High
--
90.50Thinking Enabled
87.60Thinking Enabled
84.50Thinking Enabled
Creative Writing
写作和创作
2070.80Standard Mode
--
1725.10Standard Mode
1575.70Standard Mode
1627.80Standard Mode
BrowseComp
AI Agent - 信息收集
91.20Thinking Level · High | Tools
--
83.20Thinking Enabled | Tools
60.60Thinking Enabled | Tools
60.20Thinking Enabled | Tools
Context Arena
文本向量检索
71.75Thinking Level · High
--
64.63Thinking Enabled
59.22Thinking Enabled
--
AA-LCR
长上下文能力
74.70Thinking Level · High
--
--
65.00Thinking Enabled
--
MCP-Atlas
AI Agent - 工具使用
84.20Thinking Level · High | Tools
76.00Thinking Enabled | Tools
69.40Thinking Enabled | Tools
64.40Standard Mode | Tools
--
MCPMark-Verified
AI Agent - 工具使用
94.50Thinking Level · High | Tools
81.10Thinking Enabled | Tools
72.80Thinking Enabled | Tools
--
--
OSWorld-Verified
AI Agent - 工具使用
84.80Thinking Level · High | Tools
--
73.10Thinking Enabled | Tools
--
--
Terminal-Bench 2.1
AI Agent - 工具使用
88.30Thinking Level · High | Tools
67.04Thinking Enabled | Tools
53.56Thinking Enabled
--
--
DeepSWE
编程与软件工程
67.50Thinking Level · High | Tools
31.00Standard Mode | Tools
--
--
--
Kimi Code Bench 2.0
编程与软件工程
72.90Thinking Level · High | Tools
62.00Thinking Enabled | Tools
50.90Thinking Enabled | Tools
--
--
3 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Kimi K3 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
Kimi K2.7 Code
Supplier: Moonshot AI
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
Kimi K2.6
Supplier: Facebook AI研究实验室
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
Kimi K2.5
Supplier: Moonshot AI
Standard input: $0.6 / 1M tokens
Standard output: $3 / 1M tokens
Kimi K2 Thinking
Supplier: Fireworks AI
Standard input: $0.6 / 1M tokens
Standard output: $2.5 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
Kimi K2.7 Code
Moonshot AI$0.95 / 1M tokens$4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens
Kimi K2 Thinking
Fireworks AI$0.6 / 1M tokens$2.5 / 1M tokens