DataLearner logo

Kimi K3vsGPT-5.6 Sol

Across 21 shared benchmarks, GPT-5.6 Sol leads overall: Kimi K3 wins 5, GPT-5.6 Sol wins 16, with 0 ties and an average score difference of +0.51.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

OpenAI
GPT-5.6 Sol

OpenAI · 2026-06-26 · Reasoning model

Kimi K35 wins(24%)(76%)16 winsGPT-5.6 Sol

Benchmark scores

Grouped by capability, sorted by largest gap within each. 21 shared benchmarks.

Productivity Knowledge

Kimi K3 3/5
BenchmarkKimi K3GPT-5.6 SolDiff
Harvey Lab-AA94.601 / 6Max (With Tools)2.506 / 6Max (With Tools)+92.10
GDPval-AA v21,68610 / 27Max (With Tools)1,7289 / 27Max (With Tools)-42
AA-Briefcase1,5288 / 20Max (With Tools)1,4949 / 20Max (With Tools)+34.28
AA-AnalystAgent38.757 / 12Max (With Tools + Internet)47.505 / 12Max (With Tools + Internet)-8.75
AutomationBench46.707 / 17Max (With Tools)45.808 / 17Max (With Tools)+0.90

AI Agent - Tool Usage

GPT-5.6 Sol 4/4
BenchmarkKimi K3GPT-5.6 SolDiff
Terminal-Bench 4.012.6017 / 20Max (With Tools)39.909 / 20Max (With Tools)-27.30
Terminal-Bench 3.017.708 / 11Max (With Tools)34.402 / 11Max (With Tools)-16.70
Terminal-Bench-Science 0.17.109 / 11Max (With Tools)22.404 / 11Max (With Tools)-15.30
CyberGym805 / 8Max (With Tools)84.502 / 8Max (With Tools)-4.50

Coding and Software Engineer

GPT-5.6 Sol 2/3
BenchmarkKimi K3GPT-5.6 SolDiff
DeepSWE67.5011 / 38Max (With Tools)736 / 38Max (With Tools)-5.50
Program Bench17.5010 / 11Max (With Tools)237 / 11Max (With Tools)-5.50
NL2Repo-Bench586 / 16Max (With Tools)56.809 / 16Max (With Tools)+1.20

Multimodal Understanding

GPT-5.6 Sol 3/3
BenchmarkKimi K3GPT-5.6 SolDiff
ZeroBench Main414 / 6Max (With Tools)531 / 6Max (With Tools)-12
Chartography68.106 / 8Max (With Tools)79.903 / 8Max (With Tools)-11.80
BabyVision85.705 / 9Max (With Tools)88.904 / 9Max (With Tools)-3.20

Agent Level Benchmark

GPT-5.6 Sol 2/2
BenchmarkKimi K3GPT-5.6 SolDiff
APEX-Agents415 / 6Max (With Tools)56.703 / 6Max (With Tools)-15.70
τ³-Banking33.406 / 13Max (With Tools)44.331 / 13Max (With Tools)-10.93

Long Context

GPT-5.6 Sol 1/1
BenchmarkKimi K3GPT-5.6 SolDiff
AA-LCR74.708 / 29Max (No Tools)77.674 / 29Max (No Tools)-2.97

Math and Reasoning

GPT-5.6 Sol 1/1
BenchmarkKimi K3GPT-5.6 SolDiff
FrontierMath v272.1813 / 58Max (No Tools)89.121 / 58Max (No Tools)-16.94

Text Embedding

GPT-5.6 Sol 1/1
BenchmarkKimi K3GPT-5.6 SolDiff
Context Arena71.7559 / 126Max (No Tools)97.632 / 126Max (No Tools)-25.88

Writing and Creative Capabilities

Kimi K3 1/1
BenchmarkKimi K3GPT-5.6 SolDiff
Creative Writing2,0714 / 106Normal (No Tools)1,9636 / 106Normal (No Tools)+107.20

Specs

FieldKimi K3GPT-5.6 Sol
PublisherMoonshot AIOpenAI
Release date2026-07-162026-06-26
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters2.8TNot available
Context length1M1.05M
Max output1M128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3GPT-5.6 Sol
Text input¥20 / 1M tokens$4 / 1M tokens
Text output¥100 / 1M tokens$20 / 1M tokens
Cache read¥2 / 1M tokens$0.4 / 1M tokens
Cache writeNot public$5 / 1M tokens

Summary

  • Kimi K3leads in:Productivity Knowledge (3/5), Writing and Creative Capabilities (1/1)
  • GPT-5.6 Solleads in:AI Agent - Tool Usage (4/4), Coding and Software Engineer (2/3), Multimodal Understanding (3/3), Agent Level Benchmark (2/2), Long Context (1/1), Math and Reasoning (1/1), Text Embedding (1/1)

On average across the 21 shared benchmarks, Kimi K3 scores 0.51 higher.

Largest single-benchmark gap: Creative Writing — Kimi K3 2,071 vs GPT-5.6 Sol 1,963 (+107.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.