Kimi K3vsClaude Opus 4.8
Across 14 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 8, Claude Opus 4.8 wins 6, with 0 ties and an average score difference of +25.43.
Kimi K3
Moonshot AI · 2026-07-16 · Reasoning model
Claude Opus 4.8
Anthropic · 2026-05-28 · Reasoning model
Kimi K38 wins(57%)(43%)6 winsClaude Opus 4.8
Benchmark scores
Grouped by capability, sorted by largest gap within each. 14 shared benchmarks.
AI Agent - Tool Usage
Kimi K3 3/4| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 88.304 / 49Max (With Tools) | 78.9025 / 49Thinking High (With Tools) | +9.40 |
| Terminal-Bench-Science 0.1 | 7.109 / 11Max (With Tools) | 10.506 / 11Max (With Tools) | -3.40 |
| MCP-Atlas | 84.204 / 41Max (With Tools) | 82.208 / 41Deep Thinking (With Tools) | +2 |
| OSWorld-Verified | 84.802 / 26Max (With Tools) | 83.404 / 26Extended (with tools) | +1.40 |
Coding and Software Engineer
Kimi K3 2/2| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| Text Arena (Coding) | 1,6822 / 35Max (No Tools) | 1,5459 / 35Normal (No Tools) | +136.70 |
| DeepSWE | 67.509 / 35Max (With Tools) | 5920 / 35Deep Thinking (With Tools) | +8.50 |
Math and Reasoning
Claude Opus 4.8 2/2| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| FrontierMath Tier 4 v2 | 39.0214 / 41Max (No Tools) | 56.1010 / 41Max (No Tools) | -17.07 |
| FrontierMath v2 | 72.1813 / 58Max (No Tools) | 809 / 58Max (No Tools) | -7.82 |
AI Agent - Information Search
Kimi K3 1/1| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| BrowseComp | 91.202 / 56Max (With Tools + Internet) | 84.3010 / 56Thinking High (With Tools + Internet) | +6.90 |
Commonsense Reasoning
Claude Opus 4.8 1/1| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| SimpleBench | 60.7029 / 92Max (No Tools) | 64.8019 / 92Normal (No Tools) | -4.10 |
General Evaluation
Kimi K3 1/1| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| GPQA Diamond | 93.5016 / 271Max (No Tools) | 85.3597 / 271Normal (No Tools) | +8.15 |
General Knowledge
Claude Opus 4.8 1/1| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| HLE | 5617 / 190Max (With Tools) | 57.9010 / 190Extended (with tools) | -1.90 |
Text Embedding
Claude Opus 4.8 1/1| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| Context Arena | 71.7559 / 126Max (No Tools) | 90.0415 / 126Max (No Tools) | -18.29 |
Writing and Creative Capabilities
Kimi K3 1/1| Benchmark | Kimi K3 | Claude Opus 4.8 | Diff |
|---|---|---|---|
| Creative Writing | 2,0712 / 99Normal (No Tools) | 1,83513 / 99Normal (No Tools) | +235.50 |
Specs
| Field | Kimi K3 | Claude Opus 4.8 |
|---|---|---|
| Publisher | Moonshot AI | Anthropic |
| Release date | 2026-07-16 | 2026-05-28 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | Dense |
| Parameters | 2.8T | Not available |
| Context length | 1M | 1M |
| Max output | 1M | 125K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Kimi K3 | Claude Opus 4.8 |
|---|---|---|
| Text input | ¥20 / 1M tokens | $5 / 1M tokens |
| Text output | ¥100 / 1M tokens | $25 / 1M tokens |
| Cache read | ¥2 / 1M tokens | $0.5 / 1M tokens |
| Cache write | Not public | $6.25 / 1M tokens |
Summary
- Kimi K3leads in:AI Agent - Tool Usage (3/4), Coding and Software Engineer (2/2), AI Agent - Information Search (1/1), General Evaluation (1/1), Writing and Creative Capabilities (1/1)
- Claude Opus 4.8leads in:Math and Reasoning (2/2), Commonsense Reasoning (1/1), General Knowledge (1/1), Text Embedding (1/1)
On average across the 14 shared benchmarks, Kimi K3 scores 25.43 higher.
Largest single-benchmark gap: Creative Writing — Kimi K3 2,071 vs Claude Opus 4.8 1,835 (+235.50).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.