DataLearner logo

Claude Opus 5vsKimi K3

Across 6 shared benchmarks, Claude Opus 5 leads overall: Claude Opus 5 wins 4, Kimi K3 wins 2, with 0 ties and an average score difference of +61.63.

Anthropic
Claude Opus 5

Anthropic · 2026-07-24 · Reasoning model

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

Claude Opus 54 wins(67%)(33%)2 winsKimi K3

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

Productivity Knowledge

Claude Opus 5 2/2
BenchmarkClaude Opus 5Kimi K3Diff
GDPval-AA v21,8611 / 5Max (With Tools)1,6682 / 5Max (With Tools)+193
AA-Briefcase1,7201 / 2Max (With Tools)1,5482 / 2Max (With Tools)+172

AI Agent - Information Search

Kimi K3 1/1
BenchmarkClaude Opus 5Kimi K3Diff
BrowseComp90.802 / 53Max (With Tools + Internet)91.201 / 53Max (With Tools + Internet)-0.40

AI Agent - Tool Usage

Kimi K3 1/1
BenchmarkClaude Opus 5Kimi K3Diff
Automation Bench262 / 2Max (With Tools)30.801 / 2Max (With Tools)-4.80

Coding and Software Engineer

Claude Opus 5 1/1
BenchmarkClaude Opus 5Kimi K3Diff
DeepSWE68.804 / 19Max (With Tools)67.505 / 19Max (With Tools)+1.30

General Knowledge

Claude Opus 5 1/1
BenchmarkClaude Opus 5Kimi K3Diff
HLE64.701 / 172Max (With Tools)5612 / 172Max (With Tools)+8.70

Specs

FieldClaude Opus 5Kimi K3
PublisherAnthropicMoonshot AI
Release date2026-07-242026-07-16
Model typeReasoning modelReasoning model
ArchitectureDenseMoE
ParametersNot available2.8T
Context length1M1M
Max output128K1M

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Opus 5Kimi K3
Text input$5 / 1M tokens¥20 / 1M tokens
Text output$25 / 1M tokens¥100 / 1M tokens
Cache read$0.5 / 1M tokens¥2 / 1M tokens
Cache write$6.25 / 1M tokensNot public

Summary

  • Claude Opus 5leads in:Productivity Knowledge (2/2), Coding and Software Engineer (1/1), General Knowledge (1/1)
  • Kimi K3leads in:AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1)

On average across the 6 shared benchmarks, Claude Opus 5 scores 61.63 higher.

Largest single-benchmark gap: GDPval-AA v2 — Claude Opus 5 1,861 vs Kimi K3 1,668 (+193).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.