DataLearner logo

InklingvsGLM-5.2

Across 6 shared benchmarks, GLM-5.2 leads overall: Inkling wins 0, GLM-5.2 wins 6, with 0 ties and an average score difference of -6.88.

IN
Inkling

Thinking Machines Lab · 2026-07-15 · Multimodal model

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

Inkling0 wins(0%)(100%)6 winsGLM-5.2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

AI Agent - Tool Usage

GLM-5.2 2/2
BenchmarkInklingGLM-5.2Diff
Terminal-Bench 2.163.8036 / 44Thinking (With Tools)8117 / 44Thinking High (With Tools)-17.20
MCP-Atlas7617 / 38极高强度思考(工具)76.8013 / 38Thinking (With Tools)-0.80

Coding and Software Engineer

GLM-5.2 1/1
BenchmarkInklingGLM-5.2Diff
SWE-Bench Pro - Public54.3033 / 57Thinking (With Tools)62.109 / 57Thinking (With Tools)-7.80

General Evaluation

GLM-5.2 1/1
BenchmarkInklingGLM-5.2Diff
GPQA Diamond87.2070 / 226Thinking (No Tools)91.8626 / 226最高(无工具)-4.66

General Knowledge

GLM-5.2 1/1
BenchmarkInklingGLM-5.2Diff
HLE4641 / 181Thinking (With Tools)54.7015 / 181Thinking (With Tools)-8.70

Math and Reasoning

GLM-5.2 1/1
BenchmarkInklingGLM-5.2Diff
AIME 202697.102 / 19Thinking (No Tools)99.201 / 19Thinking (No Tools)-2.10

Specs

FieldInklingGLM-5.2
PublisherThinking Machines Lab智谱AI
Release date2026-07-152026-06-13
Model typeMultimodal modelReasoning model
ArchitectureMoEMoE
Parameters975B753.33B
Context length1M1M
Max outputNot available128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemInklingGLM-5.2
Text input$3.74 / 1M tokens$1.4 / 1M tokens
Text output$9.36 / 1M tokens$4.4 / 1M tokens
Cache read$0.748 / 1M tokens$0.26 / 1M tokens

Summary

  • GLM-5.2leads in:AI Agent - Tool Usage (2/2), Coding and Software Engineer (1/1), General Evaluation (1/1), General Knowledge (1/1), Math and Reasoning (1/1)

On average across the 6 shared benchmarks, GLM-5.2 scores 6.88 higher.

Largest single-benchmark gap: Terminal-Bench 2.1 — Inkling 63.80 vs GLM-5.2 81 (-17.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.