DataLearner logo

Kimi K2.7 Code Benchmark Details

Kimi K2.7 Code currently shows benchmark results led by Terminal Bench Hard (35 / 244, score 44.70), GPQA Diamond (93 / 463, score 89.60), τ²-Bench - Telecom (66 / 264, score 90.10). This page also compares it with 5 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Kimi K2.7 Code

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
68.41
46 / 117
HLE
Thinking Mode
35
172 / 565
CritPt
Thinking Mode
10
71 / 201

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
89.60
93 / 463

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Mode
57.90
38 / 93

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking ModeTools
90.10
66 / 264
Terminal Bench Hard
Thinking ModeTools
44.70
35 / 244
τ³-Banking
Thinking ModeTools
20.20
93 / 167

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Mode
63.10
110 / 282

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
MCPMark-Verified
Thinking ModeTools
81.10
5 / 8
MCP-Atlas
Thinking ModeTools
76
22 / 44
Terminal-Bench 2.1
Thinking ModeTools
67.04
94 / 194
Terminal-Bench 4.0
Thinking ModeTools
1
76 / 88

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
Kimi Code Bench 2.0
Thinking ModeTools
62
2 / 3
Program Bench
Thinking ModeTools
53.60
4 / 12
SciCode
Thinking Mode
47.80
78 / 131
MLS Bench
Thinking ModeTools
35.10
5 / 6
DeepSWE
Standard ModeTools
31
78 / 86

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Thinking ModeTools
85.02
22 / 43

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Mode
11.20
74 / 119

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
76.11
14 / 45

Competitor Comparison

Benchmark scores for Kimi K2.7 Code compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkKimi K2.7 CodeCurrentGLM-5.2Claude Opus 4.8MiniMax M3Qwen3.7 Max
CritPt
Score
综合评估
10.00Thinking Enabled
20.90Thinking Level · High
20.90Thinking Level · High
3.70Thinking Enabled
13.40Thinking Enabled
HLE
Accuracy
综合评估
35.00Thinking Enabled
54.70Thinking Enabled | Tools
57.90Extended Thinking | Tools
39.00Thinking Enabled
53.50Thinking Enabled | Tools
LiveBench
Accuracy
综合评估
68.41Standard Mode
73.18Standard Mode
78.93Deep Thinking Mode
67.26Deep Thinking Mode
73.14Deep Thinking Mode
GPQA Diamond
Accuracy
科学与综合推理
89.60Thinking Enabled
91.86Thinking Level · High
93.60Thinking Level · High
92.90Thinking Enabled
92.40Thinking Level · High
SimpleBench
Score (AVG@5)
常识推理
57.90Thinking Enabled
58.80Standard Mode
64.80Standard Mode
45.80Thinking Enabled
70.40Standard Mode
Terminal Bench Hard
Accuracy
Agent能力评测
44.70Thinking Enabled | Tools
50.80Thinking Level · High | Tools
58.30Thinking Level · High | Tools
42.40Thinking Enabled | Tools
50.80Thinking Enabled | Tools
τ²-Bench - Telecom
Accuracy
Agent能力评测
90.10Thinking Enabled | Tools
99.10Thinking Level · High | Tools
94.40Thinking Level · High | Tools
88.90Thinking Enabled | Tools
94.70Thinking Enabled | Tools
τ³-Banking
Score
Agent能力评测
20.20Thinking Enabled | Tools
37.11Thinking Level · Extra High | Tools
39.69Thinking Level · High | Tools
15.30Thinking Enabled | Tools
11.80Thinking Enabled | Tools
IF Bench
Accuracy
指令跟随
63.10Thinking Enabled
73.30Thinking Level · High
62.20Thinking Level · High
82.90Thinking Enabled
80.50Thinking Enabled
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
76.00Thinking Enabled | Tools
77.80Thinking Enabled | Tools
82.20Thinking Level · High | Tools
74.20Thinking Enabled | Tools
76.40Thinking Enabled | Tools
MCPMark-Verified
Score
AI Agent - 工具使用
81.10Thinking Enabled | Tools
--
76.38Thinking Level · High | Tools
--
--
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
67.04Thinking Enabled | Tools
81.00Thinking Level · High | Tools
84.60Thinking Level · High | Tools
66.00Thinking Enabled | Tools
74.50Thinking Enabled | Tools
7 additional benchmarks remain in the chart above.

Standard API Pricing: Kimi K2.7 Code vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Kimi K2.7 Code
Supplier: Moonshot AI
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
Composer 2.5
Supplier: Cursor
Standard input: $0.5 / 1M tokens
Standard output: $2.5 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Claude Opus 4.8
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
Qwen3.7 Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Kimi K2.7 Code
Moonshot AI$0.95 / 1M tokens$4 / 1M tokens
Composer 2.5
Cursor$0.5 / 1M tokens$2.5 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens
Qwen3.7 Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens

Version History

How each version of the Kimi K2.7 Code series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkKimi K2.7 CodeCurrentKimi K2.6Kimi K2.5Kimi K2 Thinking
CritPt
Score
综合评估
10.00Thinking Enabled
8.00Thinking Enabled
3.10Thinking Enabled
2.60Thinking Enabled
HLE
Accuracy
综合评估
35.00Thinking Enabled
54.00Thinking Enabled | Tools
50.20Thinking Enabled | Tools
51.00Thinking Enabled | Tools
LiveBench
Accuracy
综合评估
68.41Standard Mode
70.54Thinking Enabled
69.07Thinking Enabled
61.59Thinking Enabled
GPQA Diamond
Accuracy
科学与综合推理
89.60Thinking Enabled
90.50Thinking Enabled
87.60Thinking Enabled
84.50Thinking Enabled
SimpleBench
Score (AVG@5)
常识推理
57.90Thinking Enabled
--
46.80Thinking Enabled
39.60Standard Mode
Terminal Bench Hard
Accuracy
Agent能力评测
44.70Thinking Enabled | Tools
43.90Thinking Enabled | Tools
34.80Thinking Enabled | Tools
31.10Thinking Enabled | Tools
τ²-Bench - Telecom
Accuracy
Agent能力评测
90.10Thinking Enabled | Tools
95.90Thinking Enabled | Tools
95.90Thinking Enabled | Tools
93.00Thinking Enabled | Tools
τ³-Banking
Score
Agent能力评测
20.20Thinking Enabled | Tools
23.30Thinking Enabled | Tools
14.20Thinking Enabled | Tools
--
IF Bench
Accuracy
指令跟随
63.10Thinking Enabled
76.00Thinking Enabled
70.20Thinking Enabled
68.10Thinking Enabled
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
76.00Thinking Enabled | Tools
69.40Thinking Enabled | Tools
64.40Standard Mode | Tools
--
MCPMark-Verified
Score
AI Agent - 工具使用
81.10Thinking Enabled | Tools
72.80Thinking Enabled | Tools
--
--
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
67.04Thinking Enabled | Tools
65.90Thinking Enabled | Tools
45.70Thinking Enabled | Tools
--
7 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: CritPt · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Kimi K2.7 Code Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Kimi K2.7 Code
Moonshot AI$0.95 / 1M tokens$4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens
Kimi K2 Thinking
Fireworks AI$0.6 / 1M tokens$2.5 / 1M tokens