See key specs and per-benchmark scores for each model/mode. Scroll horizontally for all columns. 当前对比 2 个模型的评测数据与核心参数。

Kimi K2 Thinking
Moonshot AI
Best overall
Kimi K3 · 80.23
Best single
Kimi K3 · GPQA Diamond 93.50
Modality coverage
Kimi K3 · 3 modalities
Head to head
3
Benchmarks
0
Wins
3
Losses
-17.03
Average diff
Compare benchmark results across thinking modes and tool usage.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
Complete scores for each model/mode across selected benchmarks.
3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Kimi K2 Thinking | Kimi K3 |
|---|---|---|
GPQA Diamond 综合评估 | 84.50Thinking Enabled | 93.50Thinking Level · High |
HLE 综合评估 | 44.90Thinking Enabled | Tools | 56.00Thinking Level · High | Tools |
BrowseComp AI Agent - 信息收集 | 60.20Thinking Enabled | Tools | 91.20Thinking Level · High | Tools |
Side-by-side input/output token pricing
Licensing, MoE architecture, and multi-modality support.
| Features & specs | Kimi K2 ThinkingMoonshot AI | Kimi K3Moonshot AI |
|---|---|---|
Core specsRelease | 2025-11-06 | 2026-07-16 |
Context length | 256K | 1M |
Parameters | 10400 | 28000 |
Active parameters | 320 | 500 |
Max output | Not provided | 1048576 |
MoE | Yes | Yes |
Supported modes | 思考模式(Thinking Mode) | No mode data |
LicenseCode Open Source | Not provided | Not provided |
Weights Open Source | Not provided | Not provided |
Commercial use | 免费商用授权 | 免费商用授权 |
Modality supportText Input/Output | / | / |
Image Input/Output | Not provided | / |
Video Input/Output | Not provided | / |
ResourcesPaper / report | Introducing Kimi K2 Thinking | Not provided |
DataLearner blog | Moonshot AI 发布 Kimi K2 Thinking:连续执行200-300次顺序工具调用,人类最后难题评测得分超过所有模型,全球第一!依然免费开源商用! | Not provided |

Kimi K3
Moonshot AI