DataLearner logo

Kimi K2.6 Benchmark Details

Kimi K2.6 currently shows benchmark results led by LiveCodeBench (9 / 128, score 89.60), HLE (23 / 197, score 54), SWE-bench Verified (14 / 116, score 80.20). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Kimi K2.6

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Thinking Mode
70.54
34 / 117
HLE
Thinking Mode
34.70
94 / 197
HLE
Thinking ModeToolsInternet
54
23 / 197

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
90.50
43 / 274

Coding and Software Engineer

7 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
89.60
9 / 128
SWE-bench Verified
Thinking ModeTools
80.20
14 / 116
SWE-bench Multilingual
Thinking ModeTools
76.70
8 / 29
SWE-Bench Pro - Public
Thinking ModeTools
58.60
18 / 62
Kimi Code Bench 2.0
Thinking ModeTools
50.90
3 / 3
Program Bench
Thinking ModeTools
48.30
4 / 11
MLS Bench
Thinking ModeTools
26.70
5 / 5

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1721.30
25 / 106

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeToolsInternet
83.20
16 / 57

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking ModeTools
73.10
15 / 26
MCPMark-Verified
Thinking ModeTools
72.80
3 / 3
MCP-Atlas
Thinking ModeTools
69.40
30 / 41
Terminal Bench 2.0
Thinking ModeTools
66.70
10 / 48
Terminal-Bench 2.1
Thinking Mode
53.56
51 / 53
Tool Decathlon
Thinking ModeTools
50
2 / 10

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
51.88
86 / 126
Context Arena
Thinking Mode
64.63
73 / 126

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
96.40
3 / 21
IMO-AnswerBench
Thinking Mode
86
10 / 24

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking ModeTools
80.90
19 / 29

Competitor Comparison

Benchmark scores for Kimi K2.6 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkKimi K2.6CurrentQwen3.6-Max-PreviewMiniMax-M2.7GLM 5.1
HLE
Accuracy
综合评估
54.00Thinking Enabled | Tools
50.20Thinking Enabled | Tools
28.00Thinking Enabled
52.30Thinking Enabled | Tools
LiveBench
Accuracy
综合评估
70.54Thinking Enabled
--
63.49Deep Thinking Mode
70.18Standard Mode
GPQA Diamond
Accuracy
科学与综合推理
90.50Thinking Enabled
90.40Thinking Level · High
87.00Thinking Enabled
86.20Thinking Enabled
LiveCodeBench
Pass @K
编程与软件工程
89.60Thinking Enabled
87.10Thinking Level · High
--
--
SWE-bench Multilingual
Accuracy
编程与软件工程
76.70Thinking Enabled | Tools
73.80Thinking Enabled | Tools
76.50Thinking Enabled | Tools
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
58.60Thinking Enabled | Tools
57.30Deep Thinking Mode | Tools
56.20Thinking Enabled | Tools
58.40Thinking Enabled | Tools
SWE-bench Verified
Accuracy
编程与软件工程
80.20Thinking Enabled | Tools
78.80Thinking Enabled | Tools
--
--
Creative Writing
Elo、大模型评判两两对战
写作和创作
1721.30Standard Mode
--
--
1589.20Standard Mode
BrowseComp
Accuracy
AI Agent - 信息收集
83.20Thinking Enabled | Tools
--
--
79.30Thinking Enabled | Tools
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
69.40Thinking Enabled | Tools
--
--
75.60Standard Mode | Tools
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
66.70Thinking Enabled | Tools
65.40Deep Thinking Mode | Tools
57.00Thinking Enabled | Tools
63.50Thinking Enabled | Tools
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
53.56Thinking Enabled
--
--
58.70Thinking Level · High | Tools
5 additional benchmarks remain in the chart above.

Standard API Pricing: Kimi K2.6 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Qwen3.6-Max-Preview: Base price applies to <= 128
ModelSupplierStandard inputStandard outputBase price applies to
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
Qwen3.6-Max-Preview
阿里巴巴$1.3 / 1M tokens$7.8 / 1M tokens<= 128
MiniMax-M2.7
MiniMaxAI$0.3 / 1M tokens$1.2 / 1M tokens
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens

Version History

How each version of the Kimi K2.6 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkKimi K2.6CurrentKimi K2.5Kimi K2 ThinkingKimi K2
HLE
Accuracy
综合评估
54.00Thinking Enabled | Tools
50.20Thinking Enabled | Tools
51.00Thinking Enabled | Tools
4.70Standard Mode
LiveBench
Accuracy
综合评估
70.54Thinking Enabled
69.07Thinking Enabled
61.59Thinking Enabled
48.10Standard Mode
GPQA Diamond
Accuracy
科学与综合推理
90.50Thinking Enabled
87.60Thinking Enabled
84.50Thinking Enabled
75.10Standard Mode
LiveCodeBench
Pass @K
编程与软件工程
89.60Thinking Enabled
85.00Thinking Enabled
83.10Thinking Enabled
53.70Standard Mode
SWE-bench Multilingual
Accuracy
编程与软件工程
76.70Thinking Enabled | Tools
73.00Thinking Enabled
--
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
58.60Thinking Enabled | Tools
50.70Thinking Enabled | Tools
--
--
SWE-bench Verified
Accuracy
编程与软件工程
80.20Thinking Enabled | Tools
76.80Thinking Enabled | Tools
71.30Thinking Enabled | Tools
51.80Standard Mode
Creative Writing
Elo、大模型评判两两对战
写作和创作
1721.30Standard Mode
1575.80Standard Mode
1627.90Standard Mode
1662.70Standard Mode
BrowseComp
Accuracy
AI Agent - 信息收集
83.20Thinking Enabled | Tools
60.60Thinking Enabled | Tools
60.20Thinking Enabled | Tools
--
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
69.40Thinking Enabled | Tools
64.40Standard Mode | Tools
--
--
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
66.70Thinking Enabled | Tools
50.80Thinking Enabled | Tools
--
--
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
64.63Thinking Enabled
59.22Thinking Enabled
--
--
3 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Kimi K2.6 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens
Kimi K2 Thinking
Fireworks AI$0.6 / 1M tokens$2.5 / 1M tokens
Kimi K2
Moonshot AI$0.6 / 1M tokens$2.5 / 1M tokens

Sources

Kimi K2.6 Benchmark Results Analysis & Model Comparisons | DataLearnerAI