DataLearner logo

Kimi K3vsGPT-5.6 Sol

Across 4 shared benchmarks, GPT-5.6 Sol leads overall: Kimi K3 wins 0, GPT-5.6 Sol wins 4, with 0 ties and an average score difference of -8.60.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

OpenAI
GPT-5.6 Sol

OpenAI · 2026-06-26 · Reasoning model

Kimi K30 wins(0%)(100%)4 winsGPT-5.6 Sol

Benchmark scores

Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.

AI Agent - Tool Usage

GPT-5.6 Sol 3/3
BenchmarkKimi K3GPT-5.6 SolDiff
Agents' Last Exam28.304 / 4Max (With Tools)52.701 / 4极高强度思考(工具)-24.40
OSWorld 2.058.303 / 3Max (With Tools)62.602 / 3极高强度思考(工具)-4.30
TerminalBench 2.188.302 / 27Max (With Tools)88.801 / 27最高(无工具)-0.50

Coding and Software Engineer

GPT-5.6 Sol 1/1
BenchmarkKimi K3GPT-5.6 SolDiff
DeepSWE67.505 / 19Max (With Tools)72.701 / 19极高强度思考(工具)-5.20

Specs

FieldKimi K3GPT-5.6 Sol
PublisherMoonshot AIOpenAI
Release date2026-07-162026-06-26
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters2.8TNot available
Context length1M1.05M
Max output1M128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3GPT-5.6 Sol
Text input¥20 / 1M tokens$5 / 1M tokens
Text output¥100 / 1M tokens$30 / 1M tokens
Cache read¥2 / 1M tokens$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • GPT-5.6 Solleads in:AI Agent - Tool Usage (3/3), Coding and Software Engineer (1/1)

On average across the 4 shared benchmarks, GPT-5.6 Sol scores 8.60 higher.

Largest single-benchmark gap: Agents' Last Exam — Kimi K3 28.30 vs GPT-5.6 Sol 52.70 (-24.40).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.