DataLearner logo

GPT-5.6 Terra Benchmark Details

GPT-5.6 Terra currently shows benchmark results led by Terminal Bench Hard (2 / 244, score 62.90), CritPt (7 / 201, score 30), FrontierMath v2 (4 / 58, score 85.96). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GPT-5.6 Terra

Benchmark Results

Thinking
Tool usage

General Knowledge

31 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
LowTools
60.17
95 / 147
ARC-AGI-1
MediumTools
77
76 / 147
ARC-AGI-1
HighTools
92
40 / 147
ARC-AGI-1
Thinking Level · Max
96.50
15 / 147
ARC-AGI-1
Thinking Level · Extra HighTools
94
30 / 147
ARC-AGI-2
LowTools
18.75
85 / 136
ARC-AGI-2
MediumTools
37.50
74 / 136
ARC-AGI-2
HighTools
67.08
43 / 136
ARC-AGI-2
Thinking Level · Max
83.90
25 / 136
ARC-AGI-2
Thinking Level · Extra HighTools
74.17
34 / 136
Vals Index
Thinking Level · Extra High
65.14
2 / 2
55
19 / 28
51.10
5 / 6
HLE
Standard Mode
11.40
375 / 565
HLE
Low
29.20
214 / 565
HLE
Medium
33.30
186 / 565
HLE
High
38.50
145 / 565
HLE
Thinking Level · Max
42.90
103 / 565
HLE
Thinking Level · Extra High
41.90
119 / 565
AA Intelligence Index v4.3
Thinking Level · MaxTools
42.30
11 / 26
CritPt
Standard Mode
2
128 / 201
9.40
74 / 201
CritPt
Medium
17.40
46 / 201
CritPt
High
22.90
31 / 201
CritPt
Thinking Level · Max
30
7 / 201
CritPt
Thinking Level · Extra High
27.10
19 / 201
ARC-AGI-3 (Standard harness)
Thinking Level · Max
0.80
13 / 32
ARC-AGI-3 (Standard harness)
Thinking Level · Extra High
0.65
14 / 32

Other

6 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
77.27
265 / 463
87.37
128 / 463
87.20
131 / 463
89.60
93 / 463
GPQA Diamond
Thinking Level · Max
93.31
32 / 463
GPQA Diamond
Thinking Level · Extra High
90.80
68 / 463

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1850.30
11 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Level · Extra High
48.90
52 / 93

Agent Level Benchmark

16 evaluations
Benchmark / mode
Score
Rank/total
60.50
158 / 264
72.80
131 / 264
78.40
118 / 264
τ²-Bench - Telecom
Thinking Level · MaxTools
86.30
82 / 264
τ²-Bench - Telecom
Thinking Level · Extra HighTools
80.40
112 / 264
43.90
37 / 244
57.60
11 / 244
Terminal Bench Hard
Thinking Level · MaxTools
57.60
11 / 244
Terminal Bench Hard
Thinking Level · Extra HighTools
62.90
2 / 244
Agents' Last Exam
Thinking Level · Extra HighTools
50.40
4 / 22
τ³-Banking
Standard ModeTools
15.70
107 / 167
τ³-Banking
LowTools
18.80
96 / 167
τ³-Banking
MediumTools
25.60
75 / 167
τ³-Banking
HighTools
28.70
67 / 167
τ³-Banking
Thinking Level · MaxTools
40.20
32 / 167
τ³-Banking
Thinking Level · Extra HighTools
29.70
63 / 167

Instruction Following

5 evaluations
Benchmark / mode
Score
Rank/total
59.70
120 / 282
IF Bench
Medium
62.20
112 / 282
64.40
104 / 282
IF Bench
Thinking Level · Max
71.20
60 / 282
IF Bench
Thinking Level · Extra High
66.30
95 / 282

Text Embedding

1 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Thinking Level · Max
92.17
11 / 126

Long Context

6 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
58.70
130 / 171
71.30
98 / 171
AA-LCR
Medium
74
89 / 171
AA-LCR
High
77.70
70 / 171
AA-LCR
Thinking Level · Max
83
13 / 171
AA-LCR
Thinking Level · Extra High
79
58 / 171

AI Agent - Tool Usage

14 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Standard ModeTools
56.20
123 / 196
62.50
107 / 196
72.30
88 / 196
75.70
78 / 196
Terminal-Bench 2.1
Thinking Level · Max
87.40
22 / 196
Terminal-Bench 2.1
Thinking Level · MaxTools
78.40
68 / 196
Terminal-Bench 2.1
Thinking Level · Extra HighTools
80.10
59 / 196
59.60
8 / 14
1.50
76 / 91
1
79 / 91
1.50
76 / 91
Terminal-Bench 4.0
Thinking Level · MaxTools
21.52
30 / 91
Terminal-Bench 4.0
Thinking Level · Extra HighTools
10.10
51 / 91
Terminal-Bench-Science 0.1
Thinking Level · MaxTools
8.60
7 / 11

Coding and Software Engineer

18 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Thinking Level · Extra High
1521.61
16 / 35
WeirdML v2
HighTools
78.27
10 / 52
AA Coding Agent Index (historical, pre-v1.4)
Thinking Level · Extra HighTools
77.40
2 / 3
DeepSWE
LowTools
24.05
86 / 89
DeepSWE
MediumTools
35.11
80 / 89
DeepSWE
HighTools
53.76
57 / 89
DeepSWE
Thinking Level · MaxTools
69.62
19 / 89
DeepSWE
Thinking Level · Extra HighTools
69.60
20 / 89
49.90
72 / 131
SciCode
Medium
50.50
67 / 131
52.40
55 / 131
SciCode
Thinking Level · Max
55
37 / 131
SciCode
Thinking Level · Extra High
52.30
56 / 131
25.20
39 / 44
CursorBench 4.0
MediumTools
27.60
38 / 44
30.70
33 / 44
CursorBench 4.0
Thinking Level · MaxTools
41.30
14 / 44
CursorBench 4.0
Thinking Level · Extra HighTools
33.60
25 / 44

Productivity Knowledge

13 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Standard ModeTools
1169
77 / 107
GDPval-AA v2
LowTools
1178
76 / 107
GDPval-AA v2
MediumTools
1318
60 / 107
GDPval-AA v2
HighTools
1415
46 / 107
GDPval-AA v2
Thinking Level · MaxTools
1477
35 / 107
GDPval-AA v2
Thinking Level · Extra HighTools
1479
34 / 107
AA-Briefcase
Standard ModeTools
947
67 / 85
AA-Briefcase
LowTools
1014
62 / 85
AA-Briefcase
MediumTools
1041
61 / 85
AA-Briefcase
HighTools
1201
47 / 85
AA-Briefcase
Thinking Level · MaxTools
1330
34 / 85
AA-Briefcase
Thinking Level · Extra HighTools
1337
33 / 85
Harvey Lab-AA
Thinking Level · MaxTools
85.18
20 / 44

Multimodal Understanding

11 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
66.70
148 / 229
76.10
77 / 229
MMMU-Pro
Medium
76.80
72 / 229
79.10
50 / 229
MMMU-Pro
Thinking Level · Max
80.70
36 / 229
MMMU-Pro
Thinking Level · Extra High
79.50
48 / 229
17.60
46 / 119
GDP.pdf
Medium
17
51 / 119
20.80
34 / 119
GDP.pdf
Thinking Level · Max
24
19 / 119
GDP.pdf
Thinking Level · Extra High
24.60
17 / 119

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Thinking Level · Max
85.96
4 / 58
FrontierMath Tier 4 v2
Thinking Level · Max
70.73
8 / 42

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
75.89
16 / 45

Competitor Comparison

Benchmark scores for GPT-5.6 Terra compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.6 TerraCurrentClaude Sonnet 5Gemini 3.5 FlashGLM-5.2
AA Intelligence Index v4.3
Index (0-100)
综合评估
42.30Standard Mode | Tools
38.40Standard Mode | Tools
--
34.00Standard Mode | Tools
ARC-AGI-1
Score (%)
综合评估
96.50Thinking Level · High
--
92.50Thinking Level · High | Tools
--
ARC-AGI-2
Score (% solved); cost per task (USD)
综合评估
83.90Thinking Level · High
--
72.08Thinking Level · High | Tools
--
CritPt
Score
综合评估
30.00Thinking Level · High
16.90Thinking Level · High
13.10Thinking Level · High
20.90Thinking Level · High
HLE
Accuracy
综合评估
42.90Thinking Level · High
57.40Thinking Level · Extra High | Tools
42.70Thinking Level · High
54.70Thinking Enabled | Tools
HLE-Verified
Accuracy (%)
综合评估
51.10Thinking Level · High
31.00Thinking Level · High
--
--
GPQA Diamond
Accuracy
科学与综合推理
93.31Thinking Level · High
90.53Thinking Level · Extra High
92.80Thinking Level · High
91.86Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1850.30Standard Mode
1790.50Standard Mode
--
1752.80Standard Mode
SimpleBench
Score (AVG@5)
常识推理
48.90Thinking Level · Extra High
60.60Standard Mode
76.70Standard Mode
58.80Standard Mode
Terminal Bench Hard
Accuracy
Agent能力评测
62.90Thinking Level · Extra High | Tools
--
46.20Standard Mode | Tools
50.80Thinking Level · High | Tools
τ²-Bench - Telecom
Accuracy
Agent能力评测
86.30Thinking Level · High | Tools
--
95.60Thinking Level · Medium | Tools
99.10Thinking Level · High | Tools
τ³-Banking
Score
Agent能力评测
40.20Thinking Level · High | Tools
37.30Thinking Level · High | Tools
32.20Thinking Level · High | Tools
37.11Thinking Level · Extra High | Tools
18 additional benchmarks remain in the chart above.

Standard API Pricing: GPT-5.6 Terra vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Terra
OpenAI$2 / 1M tokens$12 / 1M tokens
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens
Gemini 3.5 Flash
Google DeepMind$1.5 / 1M tokens$9 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens

Version History

How each version of the GPT-5.6 Terra series stacks up on benchmark tests

GPT-5.6 TerraGPT-5.5GPT-5.4
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.6 TerraCurrentGPT-5.5GPT-5.4
AA Intelligence Index v4.3
Index (0-100)
综合评估
42.30Standard Mode | Tools
38.60Thinking Level · Extra High | Tools
--
ARC-AGI-1
Score (%)
综合评估
96.50Thinking Level · High
95.00Thinking Level · Extra High
93.67Thinking Level · Extra High
ARC-AGI-2
Score (% solved); cost per task (USD)
综合评估
83.90Thinking Level · High
85.00Thinking Level · Extra High
77.10Standard Mode
CritPt
Score
综合评估
30.00Thinking Level · High
27.10Thinking Level · Extra High
23.40Thinking Level · Extra High
HLE
Accuracy
综合评估
42.90Thinking Level · High
52.20Thinking Level · High | Tools
52.10Thinking Level · Extra High | Tools
GPQA Diamond
Accuracy
科学与综合推理
93.31Thinking Level · High
94.00Thinking Level · Extra High
92.00Thinking Level · Extra High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1850.30Standard Mode
1843.50Standard Mode
1835.60Standard Mode
SimpleBench
Score (AVG@5)
常识推理
48.90Thinking Level · Extra High
69.00Standard Mode
--
Terminal Bench Hard
Accuracy
Agent能力评测
62.90Thinking Level · Extra High | Tools
60.60Thinking Level · Extra High | Tools
57.60Thinking Level · Extra High | Tools
τ²-Bench - Telecom
Accuracy
Agent能力评测
86.30Thinking Level · High | Tools
93.90Thinking Level · Extra High | Tools
87.10Thinking Level · Extra High | Tools
τ³-Banking
Score
Agent能力评测
40.20Thinking Level · High | Tools
44.59Thinking Level · Extra High | Tools
39.43Thinking Level · Extra High | Tools
IF Bench
Accuracy
指令跟随
71.20Thinking Level · High
75.90Thinking Level · Extra High
73.90Thinking Level · Extra High
16 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: AA Intelligence Index v4.3 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.6 Terra Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Terra
OpenAI$2 / 1M tokens$12 / 1M tokens
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens