DataLearner logo

MiniMax M2.5vsGLM-5

Across 7 shared benchmarks, MiniMax M2.5 leads overall: MiniMax M2.5 wins 4, GLM-5 wins 3, with 0 ties and an average score difference of -33.90.

MiniMaxAI
MiniMax M2.5

MiniMaxAI · 2026-02-12 · Reasoning model

智谱AI
GLM-5

智谱AI · 2026-02-11 · Chat model

MiniMax M2.54 wins(57%)(43%)3 winsGLM-5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

Claw-style Agent Evaluation

MiniMax M2.5 2/2
BenchmarkMiniMax M2.5GLM-5Diff
Pinch Bench87.807 / 38Thinking (With Tools)86.4013 / 38Thinking (With Tools)+1.40
Claw Bench92.104 / 29Thinking (With Tools)91.705 / 29Thinking (With Tools)+0.40

AI Agent - Information Search

MiniMax M2.5 1/1
BenchmarkMiniMax M2.5GLM-5Diff
BrowseComp76.3025 / 57Thinking (With Tools)75.9026 / 57Thinking (With Tools)+0.40

AI Agent - Tool Usage

GLM-5 1/1
BenchmarkMiniMax M2.5GLM-5Diff
Terminal Bench 2.051.7032 / 48Thinking (With Tools)61.1018 / 48Thinking (With Tools)-9.40

General Knowledge

MiniMax M2.5 1/1
BenchmarkMiniMax M2.5GLM-5Diff
ARC-AGI-163.6791 / 147Thinking (No Tools)44.67111 / 147Thinking (No Tools)+19

Productivity Knowledge

GLM-5 1/1
BenchmarkMiniMax M2.5GLM-5Diff
GDPval-AA3611 / 15Thinking (No Tools)468 / 15Thinking (No Tools)-10

Writing and Creative Capabilities

GLM-5 1/1
BenchmarkMiniMax M2.5GLM-5Diff
Creative Writing1,35868 / 106Normal (No Tools)1,59839 / 106Normal (No Tools)-239.10

Specs

FieldMiniMax M2.5GLM-5
PublisherMiniMaxAI智谱AI
Release date2026-02-122026-02-11
Model typeReasoning modelChat model
ArchitectureMoEMoE
Parameters229B744B
Context length128K200K
Max outputNot available128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemMiniMax M2.5GLM-5
Text input$0.3 / 1M tokens$1 / 1M tokens
Text output$2.4 / 1M tokens$3.2 / 1M tokens
Cache writeNot public$0.2 / 1M tokens

Summary

  • MiniMax M2.5leads in:Claw-style Agent Evaluation (2/2), AI Agent - Information Search (1/1), General Knowledge (1/1)
  • GLM-5leads in:AI Agent - Tool Usage (1/1), Productivity Knowledge (1/1), Writing and Creative Capabilities (1/1)

On average across the 7 shared benchmarks, GLM-5 scores 33.90 higher.

Largest single-benchmark gap: Creative Writing — MiniMax M2.5 1,358 vs GLM-5 1,598 (-239.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.