DataLearner logo

Claude Opus 5vsKimi K3

Across 14 shared benchmarks, Claude Opus 5 leads overall: Claude Opus 5 wins 12, Kimi K3 wins 2, with 0 ties and an average score difference of +27.43.

Anthropic
Claude Opus 5

Anthropic · 2026-07-24 · Reasoning model

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Claude Opus 512 wins(86%)(14%)2 winsKimi K3

Benchmark scores

Grouped by capability, sorted by largest gap within each. 14 shared benchmarks.

Productivity Knowledge

Claude Opus 5 2/3
BenchmarkClaude Opus 5Kimi K3Diff
GDPval-AA v21,8611 / 26Max (With Tools)1,6869 / 26Max (With Tools)+175
AA-AnalystAgent53.752 / 12Max (With Tools + Internet)38.757 / 12Max (With Tools + Internet)+15
AutomationBench2612 / 15Max (With Tools)30.809 / 15Max (With Tools)-4.80

AI Agent - Tool Usage

Claude Opus 5 2/2
BenchmarkClaude Opus 5Kimi K3Diff
Terminal-Bench-Science 0.1303 / 11Max (With Tools)7.109 / 11Max (With Tools)+22.90
OSWorld 2.070.573 / 10Max (With Tools)58.308 / 10Max (With Tools)+12.27

Coding and Software Engineer

Claude Opus 5 2/2
BenchmarkClaude Opus 5Kimi K3Diff
Text Arena (Coding)1,7121 / 35Max (No Tools)1,6822 / 35Max (No Tools)+30.13
DeepSWE68.808 / 35Max (With Tools)67.509 / 35Max (With Tools)+1.30

Math and Reasoning

Claude Opus 5 2/2
BenchmarkClaude Opus 5Kimi K3Diff
FrontierMath Tier 4 v273.176 / 42Max (No Tools)39.0215 / 42Max (No Tools)+34.15
FrontierMath v285.615 / 58Max (No Tools)72.1813 / 58Max (No Tools)+13.43

AI Agent - Information Search

Kimi K3 1/1
BenchmarkClaude Opus 5Kimi K3Diff
BrowseComp90.803 / 57Max (With Tools + Internet)91.202 / 57Max (With Tools + Internet)-0.40

General Evaluation

Claude Opus 5 1/1
BenchmarkClaude Opus 5Kimi K3Diff
GPQA Diamond93.8813 / 273Max (No Tools)93.5017 / 273Max (No Tools)+0.38

General Knowledge

Claude Opus 5 1/1
BenchmarkClaude Opus 5Kimi K3Diff
HLE64.702 / 191Max (With Tools)5617 / 191Max (With Tools)+8.70

Text Embedding

Claude Opus 5 1/1
BenchmarkClaude Opus 5Kimi K3Diff
Context Arena97.721 / 126Max (No Tools)71.7559 / 126Max (No Tools)+25.97

Writing and Creative Capabilities

Claude Opus 5 1/1
BenchmarkClaude Opus 5Kimi K3Diff
Creative Writing2,1213 / 106Normal (No Tools)2,0714 / 106Normal (No Tools)+50

Specs

FieldClaude Opus 5Kimi K3
PublisherAnthropicMoonshot AI
Release date2026-07-242026-07-16
Model typeReasoning modelReasoning model
ArchitectureDenseMoE
ParametersNot available2.8T
Context length1M1M
Max output128K1M

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Opus 5Kimi K3
Text input$5 / 1M tokens¥20 / 1M tokens
Text output$25 / 1M tokens¥100 / 1M tokens
Cache read$0.5 / 1M tokens¥2 / 1M tokens
Cache write$6.25 / 1M tokensNot public

Summary

  • Claude Opus 5leads in:Productivity Knowledge (2/3), AI Agent - Tool Usage (2/2), Coding and Software Engineer (2/2), Math and Reasoning (2/2), General Evaluation (1/1), General Knowledge (1/1), Text Embedding (1/1), Writing and Creative Capabilities (1/1)
  • Kimi K3leads in:AI Agent - Information Search (1/1)

On average across the 14 shared benchmarks, Claude Opus 5 scores 27.43 higher.

Largest single-benchmark gap: GDPval-AA v2 — Claude Opus 5 1,861 vs Kimi K3 1,686 (+175).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.