DataLearner logo

Kimi K2.5 Benchmark Details

Kimi K2.5 currently shows benchmark results led by LiveCodeBench (19 / 128, score 85), HLE (36 / 197, score 50.20), AIME2025 (21 / 107, score 96.10). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Kimi K2.5

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Thinking Mode
78.50
70 / 134
LiveBench
Thinking Mode
69.07
41 / 117
ARC-AGI-1
Thinking Mode
65.30
58 / 92
HLE
Thinking Mode
30.10
109 / 197
HLE
Thinking ModeTools
50.20
36 / 197
ARC-AGI-2
Thinking Mode
11.80
62 / 85

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
87.60
78 / 274

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1430.56
27 / 35
LiveCodeBench
Thinking Mode
85
19 / 128
SWE-bench Verified
Thinking ModeTools
76.80
30 / 116
73
17 / 29
SWE-Bench Pro - Public
Thinking ModeTools
50.70
47 / 62
WeirdML v2
Standard ModeTools
45.60
42 / 52

Math and Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
96.10
21 / 107
AIME 2026
Thinking Mode
92.50
13 / 21
IMO-AnswerBench
Thinking Mode
81.80
18 / 24
4.20
40 / 80

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1575.80
43 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Mode
46.80
54 / 94

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeToolsInternet
60.60
38 / 57

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Standard ModeTools
64.40
32 / 41
Terminal Bench 2.0
Thinking ModeTools
50.80
35 / 48

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
53.33
83 / 126
Context Arena
Thinking Mode
59.22
77 / 126

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Mode
40
15 / 21

Long Context

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Mode
65
22 / 29
LongBench v2
Standard Mode
61
6 / 14

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
84.80
18 / 38
Claw Bench
Thinking ModeTools
81.70
18 / 29

Other

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Standard Mode
86.10
7 / 16

Competitor Comparison

Benchmark scores for Kimi K2.5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkKimi K2.5CurrentGLM-5MiniMax M2.5
ARC-AGI-1
Score (%)
综合评估
65.30Thinking Enabled
44.70Thinking Enabled
63.70Thinking Enabled
ARC-AGI-2
Accuracy and cost per task
综合评估
11.80Thinking Enabled
4.90Thinking Enabled
4.90Thinking Enabled
HLE
Accuracy
综合评估
50.20Thinking Enabled | Tools
50.40Thinking Enabled | Tools
19.40Thinking Enabled
LiveBench
Accuracy
综合评估
69.07Thinking Enabled
68.85Standard Mode
60.14Deep Thinking Mode
GPQA Diamond
Accuracy
科学与综合推理
87.60Thinking Enabled
86.00Thinking Enabled
85.20Thinking Enabled
SWE-Bench Pro - Public
Accuracy
编程与软件工程
50.70Thinking Enabled | Tools
--
55.40Thinking Enabled | Tools
SWE-bench Verified
Accuracy
编程与软件工程
76.80Thinking Enabled | Tools
77.80Thinking Enabled
80.20Thinking Enabled | Tools
AIME 2026
Accuracy
数学推理
92.50Thinking Enabled
92.70Thinking Enabled
--
AIME2025
Accuracy
数学推理
96.10Thinking Enabled
--
86.30Thinking Enabled
FrontierMath - Tier 4
Accuracy
数学推理
4.20Standard Mode
2.10Standard Mode
--
IMO-AnswerBench
Accuracy
数学推理
81.80Thinking Enabled
82.50Thinking Enabled
--
Creative Writing
Elo、大模型评判两两对战
写作和创作
1575.80Standard Mode
1597.50Standard Mode
1358.40Standard Mode
8 additional benchmarks remain in the chart above.

Standard API Pricing: Kimi K2.5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens
MiniMax M2.5
MiniMaxAI$0.3 / 1M tokens$2.4 / 1M tokens

Version History

How each version of the Kimi K2.5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkKimi K2.5CurrentKimi K2 ThinkingKimi K2 0905Kimi K2
ARC-AGI-1
Score (%)
综合评估
65.30Thinking Enabled
--
--
13.30Standard Mode
HLE
Accuracy
综合评估
50.20Thinking Enabled | Tools
51.00Thinking Enabled | Tools
21.70Thinking Enabled | Tools
4.70Standard Mode
LiveBench
Accuracy
综合评估
69.07Thinking Enabled
61.59Thinking Enabled
--
48.10Standard Mode
MMLU Pro
Accuracy
综合评估
78.50Thinking Enabled
84.60Thinking Enabled
--
81.10Standard Mode
GPQA Diamond
Accuracy
科学与综合推理
87.60Thinking Enabled
84.50Thinking Enabled
--
75.10Standard Mode
LiveCodeBench
Pass @K
编程与软件工程
85.00Thinking Enabled
83.10Thinking Enabled
--
53.70Standard Mode
SWE-Bench Pro - Public
Accuracy
编程与软件工程
50.70Thinking Enabled | Tools
--
27.67Standard Mode
--
SWE-bench Verified
Accuracy
编程与软件工程
76.80Thinking Enabled | Tools
71.30Thinking Enabled | Tools
69.20Standard Mode
51.80Standard Mode
AIME2025
Accuracy
数学推理
96.10Thinking Enabled
100.00Thinking Enabled | Tools
75.20Thinking Enabled | Tools
54.00Standard Mode
FrontierMath - Tier 4
Accuracy
数学推理
4.20Standard Mode
0.00Thinking Enabled
--
0.01Standard Mode
Creative Writing
Elo、大模型评判两两对战
写作和创作
1575.80Standard Mode
1627.90Standard Mode
--
1662.70Standard Mode
SimpleBench
Score (AVG@5)
常识推理
46.80Thinking Enabled
39.60Standard Mode
--
26.30Standard Mode
2 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Kimi K2.5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens
Kimi K2 Thinking
Fireworks AI$0.6 / 1M tokens$2.5 / 1M tokens
Kimi K2 0905
Fireworks AI$0.6 / 1M tokens$2.5 / 1M tokens
Kimi K2
Moonshot AI$0.6 / 1M tokens$2.5 / 1M tokens

Sources