DataLearner logo

Hy4 preview Benchmark Details

Hy4 preview currently shows benchmark results led by HLE (15 / 185, score 55.40), SWE-Bench Pro - Public (6 / 59, score 65.70), GPQA Diamond (24 / 226, score 92.30). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Hy4 preview

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
HLE
High
43.40
54 / 185
HLE
HighTools
55.40
15 / 185
CritPt
High
16.90
2 / 3

Other

1 evaluations
Benchmark / mode
Score
Rank/total
92.30
24 / 226

Coding and Software Engineer

7 evaluations
Benchmark / mode
Score
Rank/total
82.90
3 / 27
65.70
6 / 59
DeepSWE
HighTools
64.30
11 / 31
NL2Repo-Bench
HighTools
58.90
2 / 11
35.60
4 / 5
SWE-Marathon
HighTools
31.90
3 / 5
Program Bench
HighTools
17.50
6 / 6

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
85.40
10 / 47
MCP-Atlas
HighTools
83.70
5 / 40
CyberGym
HighTools
78.40
3 / 4
74.10
3 / 8
71.30
1 / 1
32.10
3 / 10

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
Job Bench
HighTools
61.70
1 / 3
APEX-Agents
HighTools
37.10
6 / 6
22.80
13 / 14

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
HighTools
1678
8 / 16
Office QA Pro
HighTools
66.20
1 / 3

Competitor Comparison

Benchmark scores for Hy4 preview compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkHy4 previewCurrentGLM-5.3DeepSeek-V4-ProKimi K3
CritPt
综合评估
16.90Thinking Level · High
--
--
23.40Thinking Level · High
HLE
综合评估
55.40Thinking Level · High | Tools
62.50Thinking Level · High | Tools
48.20Thinking Level · Extra High | Tools
56.00Thinking Level · High | Tools
GPQA Diamond
科学与综合推理
92.30Thinking Level · High
--
90.10Thinking Level · High
93.50Thinking Level · High
DeepSWE
编程与软件工程
64.30Thinking Level · High | Tools
66.90Thinking Level · High | Tools
62.70Thinking Level · Extra High | Tools
67.50Thinking Level · High | Tools
NL2Repo-Bench
编程与软件工程
58.90Thinking Level · High | Tools
58.00Thinking Level · High | Tools
61.50Thinking Level · Extra High | Tools
--
PostTrain Bench
编程与软件工程
35.60Thinking Level · High | Tools
39.80Thinking Level · High | Tools
--
36.60Thinking Level · High | Tools
Program Bench
编程与软件工程
17.50Thinking Level · High | Tools
19.00Thinking Level · High | Tools
--
77.80Thinking Level · High | Tools
SWE-bench Multilingual
编程与软件工程
82.90Thinking Level · High | Tools
--
76.20Thinking Level · Extra High | Tools
--
SWE-Bench Pro - Public
编程与软件工程
65.70Thinking Level · High | Tools
--
55.40Thinking Level · Extra High | Tools
--
SWE-Marathon
编程与软件工程
31.90Thinking Level · High | Tools
42.50Thinking Level · High | Tools
--
42.00Thinking Level · High | Tools
AutomationBench
AI Agent - 工具使用
32.10Thinking Level · High | Tools
48.20Thinking Level · High | Tools
31.80Thinking Level · Extra High | Tools
30.80Thinking Level · High | Tools
CyberGym
AI Agent - 工具使用
78.40Thinking Level · High | Tools
84.50Thinking Level · High | Tools
83.30Thinking Level · Extra High | Tools
--
8 additional benchmarks remain in the chart above.

Standard API Pricing: Hy4 preview vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Hy4 preview
Supplier: 腾讯AI实验室
Standard input: ¥6 / 1M tokens
Standard output: ¥18 / 1M tokens
DeepSeek-V4-Pro
Supplier: DeepSeek-AI
Standard input: $0.435 / 1M tokens
Standard output: $0.87 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Hy4 preview
腾讯AI实验室¥6 / 1M tokens¥18 / 1M tokens
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens

Version History

How each version of the Hy4 preview series stacks up on benchmark tests

Hy4 previewHy3Hy3 Pre
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

7 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkHy4 previewCurrentHy3
HLE
综合评估
55.40Thinking Level · High | Tools
53.20Thinking Level · High | Tools
GPQA Diamond
科学与综合推理
92.30Thinking Level · High
90.40Thinking Level · High
DeepSWE
编程与软件工程
64.30Thinking Level · High | Tools
28.00Thinking Level · High | Tools
SWE-bench Multilingual
编程与软件工程
82.90Thinking Level · High | Tools
75.80Thinking Level · High | Tools
SWE-Bench Pro - Public
编程与软件工程
65.70Thinking Level · High | Tools
57.90Thinking Level · High | Tools
MCP-Atlas
AI Agent - 工具使用
83.70Thinking Level · High | Tools
79.10Thinking Level · High | Tools
Terminal-Bench 2.1
AI Agent - 工具使用
85.40Thinking Level · High | Tools
71.70Thinking Level · High | Tools

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Hy4 preview Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Hy3: Base price applies to <= 16000
ModelSupplierStandard inputStandard outputBase price applies to
Hy4 preview
腾讯AI实验室¥6 / 1M tokens¥18 / 1M tokens
Hy3
腾讯AI实验室¥1.2 / 1M tokens¥4 / 1M tokens<= 16000

Sources

Hy4 preview Benchmark Results Analysis & Model Comparisons | DataLearnerAI