DataLearner logo

Kimi K3vsClaude Opus 4.8

Across 6 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 4, Claude Opus 4.8 wins 2, with 0 ties and an average score difference of +4.13.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Anthropic
Claude Opus 4.8

Anthropic · 2026-05-28 · Reasoning model

Kimi K34 wins(67%)(33%)2 winsClaude Opus 4.8

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

AI Agent - Tool Usage

Kimi K3 2/2
BenchmarkKimi K3Claude Opus 4.8Diff
TerminalBench 2.188.302 / 25Max (With Tools)78.9012 / 25Thinking High (With Tools)+9.40
MCP-Atlas84.202 / 27Max (With Tools)82.205 / 27Deep Thinking (With Tools)+2

General Knowledge

Claude Opus 4.8 2/2
BenchmarkKimi K3Claude Opus 4.8Diff
HLE5610 / 170Max (With Tools)57.906 / 170Extended (with tools)-1.90
GPQA Diamond93.508 / 187最高(无工具)93.606 / 187Thinking High (No Tools)-0.10

AI Agent - Information Search

Kimi K3 1/1
BenchmarkKimi K3Claude Opus 4.8Diff
BrowseComp91.201 / 52Max (With Tools + Internet)84.308 / 52Thinking High (With Tools + Internet)+6.90

Coding and Software Engineer

Kimi K3 1/1
BenchmarkKimi K3Claude Opus 4.8Diff
DeepSWE67.504 / 17Max (With Tools)597 / 17Deep Thinking (With Tools)+8.50

Specs

FieldKimi K3Claude Opus 4.8
PublisherMoonshot AIAnthropic
Release date2026-07-162026-05-28
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters2.8TNot available
Context length1M1M
Max output1M125K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3Claude Opus 4.8
Text input¥20 / 1M tokens$5 / 1M tokens
Text output¥100 / 1M tokens$25 / 1M tokens
Cache read¥2 / 1M tokens$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • Kimi K3leads in:AI Agent - Tool Usage (2/2), AI Agent - Information Search (1/1), Coding and Software Engineer (1/1)
  • Claude Opus 4.8leads in:General Knowledge (2/2)

On average across the 6 shared benchmarks, Kimi K3 scores 4.13 higher.

Largest single-benchmark gap: TerminalBench 2.1 — Kimi K3 88.30 vs Claude Opus 4.8 78.90 (+9.40).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.