DataLearner logo

Kimi K3 Benchmark Details

Kimi K3 currently shows benchmark results led by BrowseComp (1 / 53, score 91.20), GPQA Diamond (8 / 187, score 93.50), AA-LCR (1 / 15, score 74.70). This page also compares it with 4 competitor models and 4 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Kimi K3

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
93.50
8 / 187
HLE
Max
43.50
47 / 172
HLE
MaxTools
56
12 / 172
23.40
1 / 1

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
DeepSearchQA
MaxToolsInternet
95
1 / 1
BrowseComp
MaxToolsInternet
91.20
1 / 53

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
74.70
1 / 15

AI Agent - Tool Usage

8 evaluations
Benchmark / mode
Score
Rank/total
94.50
1 / 1
88.30
2 / 28
84.80
2 / 24
MCP-Atlas
MaxTools
84.20
2 / 27
76.50
1 / 2
SaaS-Bench
MaxTools
60.10
1 / 1
OSWorld 2.0
MaxTools
58.30
3 / 3
30.80
1 / 3

Coding and Software Engineer

8 evaluations
Benchmark / mode
Score
Rank/total
FrontierSWE
MaxTools
81.20
1 / 1
77.80
1 / 1
72.90
1 / 1
DeepSWE
MaxTools
67.50
5 / 20
SciCode
MaxTools
58.70
1 / 1
MLS Bench
MaxTools
48.30
1 / 1
SWE-Marathon
MaxTools
42
1 / 2
36.60
1 / 1

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total
Job Bench
MaxTools
54.30
1 / 1
APEX-Agents
MaxTools
41
1 / 1
τ³-Banking
MaxTools
33.40
1 / 1
28.30
4 / 5

Productivity Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
MaxTools
1686
2 / 5
AA-Briefcase
MaxTools
1548
2 / 2
94.60
1 / 1
ResearchRubrics
MaxToolsInternet
76.20
1 / 1
CorpFin v2
MaxTools
71.60
1 / 1
63.30
1 / 1
54.40
1 / 1
44.20
1 / 1
34.80
1 / 1

Multimodal Understanding

14 evaluations
Benchmark / mode
Score
Rank/total
94.30
3 / 4
MathVision
MaxTools
97.80
1 / 4
84.80
5 / 8
CharXiv RQ
MaxTools
91.30
1 / 8
91.10
1 / 1
BabyVision
MaxTools
85.70
1 / 2
81.60
3 / 6
MMMU-Pro
MaxTools
83.40
1 / 6
MMVU
Max
82.10
1 / 1
58.50
1 / 1
23
2 / 2
41
1 / 2

Competitor Comparison

Benchmark scores for Kimi K3 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

9 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkKimi K3CurrentGLM-5.2MiniMax M3Claude Opus 4.8GPT-5.6 Sol
GPQA Diamond
综合评估
93.50Thinking Level · High
91.20Thinking Enabled
--
93.60Thinking Level · High
--
HLE
综合评估
56.00Thinking Level · High | Tools
54.70Thinking Enabled | Tools
--
57.90Extended Thinking | Tools
--
BrowseComp
AI Agent - 信息收集
91.20Thinking Level · High | Tools
--
83.50Thinking Enabled | Tools
84.30Thinking Level · High | Tools
--
MCP-Atlas
AI Agent - 工具使用
84.20Thinking Level · High | Tools
--
--
82.20Deep Thinking Mode | Tools
--
OSWorld 2.0
AI Agent - 工具使用
58.30Thinking Level · High | Tools
--
--
--
62.60Thinking Level · Extra High | Tools
OSWorld-Verified
AI Agent - 工具使用
84.80Thinking Level · High | Tools
--
70.00Thinking Enabled | Tools
83.40Extended Thinking | Tools
--
TerminalBench 2.1
AI Agent - 工具使用
88.30Thinking Level · High | Tools
81.00Thinking Level · High | Tools
66.00Thinking Enabled | Tools
78.90Thinking Level · High | Tools
88.80Thinking Level · High
DeepSWE
编程与软件工程
67.50Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
--
59.00Deep Thinking Mode | Tools
72.70Thinking Level · Extra High | Tools
Agents' Last Exam
Agent能力评测
28.30Thinking Level · High | Tools
--
--
--
52.70Thinking Level · Extra High | Tools

Standard API Pricing: Kimi K3 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
Claude Opus 4.8
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
GPT-5.6 Sol
Supplier: OpenAI
Standard input: $5 / 1M tokens
Standard output: $30 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.6 Sol
OpenAI$5 / 1M tokens$30 / 1M tokens

Version History

How each version of the Kimi K3 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkKimi K3CurrentKimi K2.7 CodeKimi K2.6Kimi K2.5Kimi K2 Thinking
GPQA Diamond
综合评估
93.50Thinking Level · High
--
90.50Thinking Enabled
87.60Thinking Enabled
84.50Thinking Enabled
HLE
综合评估
56.00Thinking Level · High | Tools
--
54.00Thinking Enabled | Tools
50.20Thinking Enabled | Tools
51.00Thinking Enabled | Tools
BrowseComp
AI Agent - 信息收集
91.20Thinking Level · High | Tools
--
83.20Thinking Enabled | Tools
60.60Thinking Enabled | Tools
60.20Thinking Enabled | Tools
AA-LCR
长上下文能力
74.70Thinking Level · High
--
--
65.00Thinking Enabled
--
MCP-Atlas
AI Agent - 工具使用
84.20Thinking Level · High | Tools
--
--
64.40Standard Mode | Tools
--
OSWorld-Verified
AI Agent - 工具使用
84.80Thinking Level · High | Tools
--
73.10Thinking Enabled | Tools
--
--
TerminalBench 2.1
AI Agent - 工具使用
88.30Thinking Level · High | Tools
67.04Thinking Enabled | Tools
53.56Thinking Enabled
--
--
DeepSWE
编程与软件工程
67.50Thinking Level · High | Tools
31.00Standard Mode | Tools
--
--
--

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Kimi K3 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
Kimi K2.7 Code
Supplier: Moonshot AI
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
Kimi K2.6
Supplier: Facebook AI研究实验室
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
Kimi K2.5
Supplier: Moonshot AI
Standard input: $0.6 / 1M tokens
Standard output: $3 / 1M tokens
Kimi K2 Thinking
Supplier: Fireworks AI
Standard input: $0.6 / 1M tokens
Standard output: $2.5 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
Kimi K2.7 Code
Moonshot AI$0.95 / 1M tokens$4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens
Kimi K2 Thinking
Fireworks AI$0.6 / 1M tokens$2.5 / 1M tokens