DataLearner logo

Kimi K3vsClaude Opus 4.8

Across 14 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 8, Claude Opus 4.8 wins 6, with 0 ties and an average score difference of +25.43.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Anthropic
Claude Opus 4.8

Anthropic · 2026-05-28 · Reasoning model

Kimi K38 wins(57%)(43%)6 winsClaude Opus 4.8

Benchmark scores

Grouped by capability, sorted by largest gap within each. 14 shared benchmarks.

AI Agent - Tool Usage

Kimi K3 3/4
BenchmarkKimi K3Claude Opus 4.8Diff
Terminal-Bench 2.188.304 / 49Max (With Tools)78.9025 / 49Thinking High (With Tools)+9.40
Terminal-Bench-Science 0.17.109 / 11Max (With Tools)10.506 / 11Max (With Tools)-3.40
MCP-Atlas84.204 / 41Max (With Tools)82.208 / 41Deep Thinking (With Tools)+2
OSWorld-Verified84.802 / 26Max (With Tools)83.404 / 26Extended (with tools)+1.40

Coding and Software Engineer

Kimi K3 2/2
BenchmarkKimi K3Claude Opus 4.8Diff
Text Arena (Coding)1,6822 / 35Max (No Tools)1,5459 / 35Normal (No Tools)+136.70
DeepSWE67.509 / 35Max (With Tools)5920 / 35Deep Thinking (With Tools)+8.50

Math and Reasoning

Claude Opus 4.8 2/2
BenchmarkKimi K3Claude Opus 4.8Diff
FrontierMath Tier 4 v239.0214 / 41Max (No Tools)56.1010 / 41Max (No Tools)-17.07
FrontierMath v272.1813 / 58Max (No Tools)809 / 58Max (No Tools)-7.82

AI Agent - Information Search

Kimi K3 1/1
BenchmarkKimi K3Claude Opus 4.8Diff
BrowseComp91.202 / 56Max (With Tools + Internet)84.3010 / 56Thinking High (With Tools + Internet)+6.90

Commonsense Reasoning

Claude Opus 4.8 1/1
BenchmarkKimi K3Claude Opus 4.8Diff
SimpleBench60.7029 / 92Max (No Tools)64.8019 / 92Normal (No Tools)-4.10

General Evaluation

Kimi K3 1/1
BenchmarkKimi K3Claude Opus 4.8Diff
GPQA Diamond93.5016 / 271Max (No Tools)85.3597 / 271Normal (No Tools)+8.15

General Knowledge

Claude Opus 4.8 1/1
BenchmarkKimi K3Claude Opus 4.8Diff
HLE5617 / 190Max (With Tools)57.9010 / 190Extended (with tools)-1.90

Text Embedding

Claude Opus 4.8 1/1
BenchmarkKimi K3Claude Opus 4.8Diff
Context Arena71.7559 / 126Max (No Tools)90.0415 / 126Max (No Tools)-18.29

Writing and Creative Capabilities

Kimi K3 1/1
BenchmarkKimi K3Claude Opus 4.8Diff
Creative Writing2,0712 / 99Normal (No Tools)1,83513 / 99Normal (No Tools)+235.50

Specs

FieldKimi K3Claude Opus 4.8
PublisherMoonshot AIAnthropic
Release date2026-07-162026-05-28
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters2.8TNot available
Context length1M1M
Max output1M125K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3Claude Opus 4.8
Text input¥20 / 1M tokens$5 / 1M tokens
Text output¥100 / 1M tokens$25 / 1M tokens
Cache read¥2 / 1M tokens$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • Kimi K3leads in:AI Agent - Tool Usage (3/4), Coding and Software Engineer (2/2), AI Agent - Information Search (1/1), General Evaluation (1/1), Writing and Creative Capabilities (1/1)
  • Claude Opus 4.8leads in:Math and Reasoning (2/2), Commonsense Reasoning (1/1), General Knowledge (1/1), Text Embedding (1/1)

On average across the 14 shared benchmarks, Kimi K3 scores 25.43 higher.

Largest single-benchmark gap: Creative Writing — Kimi K3 2,071 vs Claude Opus 4.8 1,835 (+235.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.