Hy4 preview Benchmark Details
Hy4 preview currently shows benchmark results led by HLE (15 / 185, score 55.40), SWE-Bench Pro - Public (6 / 59, score 65.70), GPQA Diamond (24 / 226, score 92.30). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.
Benchmark Results
Benchmark Results
General Knowledge
3 evaluationsCoding and Software Engineer
7 evaluationsAI Agent - Tool Usage
6 evaluationsAgent Level Benchmark
3 evaluationsProductivity Knowledge
2 evaluationsCompetitor Comparison
Benchmark scores for Hy4 preview compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Hy4 previewCurrent | GLM-5.3 | DeepSeek-V4-Pro | Kimi K3 |
|---|---|---|---|---|
CritPt 综合评估 | 16.90Thinking Level · High | -- | -- | 23.40Thinking Level · High |
HLE 综合评估 | 55.40Thinking Level · High | Tools | 62.50Thinking Level · High | Tools | 48.20Thinking Level · Extra High | Tools | 56.00Thinking Level · High | Tools |
GPQA Diamond 科学与综合推理 | 92.30Thinking Level · High | -- | 90.10Thinking Level · High | 93.50Thinking Level · High |
DeepSWE 编程与软件工程 | 64.30Thinking Level · High | Tools | 66.90Thinking Level · High | Tools | 62.70Thinking Level · Extra High | Tools | 67.50Thinking Level · High | Tools |
NL2Repo-Bench 编程与软件工程 | 58.90Thinking Level · High | Tools | 58.00Thinking Level · High | Tools | 61.50Thinking Level · Extra High | Tools | -- |
PostTrain Bench 编程与软件工程 | 35.60Thinking Level · High | Tools | 39.80Thinking Level · High | Tools | -- | 36.60Thinking Level · High | Tools |
Program Bench 编程与软件工程 | 17.50Thinking Level · High | Tools | 19.00Thinking Level · High | Tools | -- | 77.80Thinking Level · High | Tools |
SWE-bench Multilingual 编程与软件工程 | 82.90Thinking Level · High | Tools | -- | 76.20Thinking Level · Extra High | Tools | -- |
SWE-Bench Pro - Public 编程与软件工程 | 65.70Thinking Level · High | Tools | -- | 55.40Thinking Level · Extra High | Tools | -- |
SWE-Marathon 编程与软件工程 | 31.90Thinking Level · High | Tools | 42.50Thinking Level · High | Tools | -- | 42.00Thinking Level · High | Tools |
AutomationBench AI Agent - 工具使用 | 32.10Thinking Level · High | Tools | 48.20Thinking Level · High | Tools | 31.80Thinking Level · Extra High | Tools | 30.80Thinking Level · High | Tools |
CyberGym AI Agent - 工具使用 | 78.40Thinking Level · High | Tools | 84.50Thinking Level · High | Tools | 83.30Thinking Level · Extra High | Tools | -- |
Standard API Pricing: Hy4 preview vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.
These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Hy4 preview | 腾讯AI实验室 | ¥6 / 1M tokens | ¥18 / 1M tokens | — |
DeepSeek-V4-Pro | DeepSeek-AI | $0.435 / 1M tokens | $0.87 / 1M tokens | — |
Kimi K3 | Moonshot AI | ¥20 / 1M tokens | ¥100 / 1M tokens | — |
Version History
How each version of the Hy4 preview series stacks up on benchmark tests
7 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Hy4 previewCurrent | Hy3 |
|---|---|---|
HLE 综合评估 | 55.40Thinking Level · High | Tools | 53.20Thinking Level · High | Tools |
GPQA Diamond 科学与综合推理 | 92.30Thinking Level · High | 90.40Thinking Level · High |
DeepSWE 编程与软件工程 | 64.30Thinking Level · High | Tools | 28.00Thinking Level · High | Tools |
SWE-bench Multilingual 编程与软件工程 | 82.90Thinking Level · High | Tools | 75.80Thinking Level · High | Tools |
SWE-Bench Pro - Public 编程与软件工程 | 65.70Thinking Level · High | Tools | 57.90Thinking Level · High | Tools |
MCP-Atlas AI Agent - 工具使用 | 83.70Thinking Level · High | Tools | 79.10Thinking Level · High | Tools |
Terminal-Bench 2.1 AI Agent - 工具使用 | 85.40Thinking Level · High | Tools | 71.70Thinking Level · High | Tools |
Single-Benchmark Version Trend
Viewing: HLE · 综合评估
Standard API Pricing Across the Hy4 preview Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Hy4 preview | 腾讯AI实验室 | ¥6 / 1M tokens | ¥18 / 1M tokens | — |
Hy3 | 腾讯AI实验室 | ¥1.2 / 1M tokens | ¥4 / 1M tokens | <= 16000 |