DataLearner logo

Kimi K3vsInkling

Across 5 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 5, Inkling wins 0, with 0 ties and an average score difference of +12.62.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

IN
Inkling

Thinking Machines Lab · 2026-07-15 · Multimodal model

Kimi K35 wins(100%)(0%)0 winsInkling

Benchmark scores

Grouped by capability, sorted by largest gap within each. 5 shared benchmarks.

AI Agent - Tool Usage

Kimi K3 2/2
BenchmarkKimi K3InklingDiff
Terminal-Bench 2.188.302 / 44Max (With Tools)63.8036 / 44Thinking (With Tools)+24.50
MCP-Atlas84.203 / 38Max (With Tools)7617 / 38极高强度思考(工具)+8.20

AI Agent - Information Search

Kimi K3 1/1
BenchmarkKimi K3InklingDiff
BrowseComp91.201 / 54Max (With Tools + Internet)77.1022 / 54Thinking (With Tools + Internet)+14.10

General Evaluation

Kimi K3 1/1
BenchmarkKimi K3InklingDiff
GPQA Diamond93.5013 / 226最高(无工具)87.2070 / 226Thinking (No Tools)+6.30

General Knowledge

Kimi K3 1/1
BenchmarkKimi K3InklingDiff
HLE5614 / 181Max (With Tools)4641 / 181Thinking (With Tools)+10

Specs

FieldKimi K3Inkling
PublisherMoonshot AIThinking Machines Lab
Release date2026-07-162026-07-15
Model typeReasoning modelMultimodal model
ArchitectureMoEMoE
Parameters2.8T975B
Context length1M1M
Max output1MNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3Inkling
Text input¥20 / 1M tokens$3.74 / 1M tokens
Text output¥100 / 1M tokens$9.36 / 1M tokens
Cache read¥2 / 1M tokens$0.748 / 1M tokens

Summary

  • Kimi K3leads in:AI Agent - Tool Usage (2/2), AI Agent - Information Search (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 5 shared benchmarks, Kimi K3 scores 12.62 higher.

Largest single-benchmark gap: Terminal-Bench 2.1 — Kimi K3 88.30 vs Inkling 63.80 (+24.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.