Hy3 Benchmark Details
Hy3 currently shows benchmark results led by HLE (21 / 185, score 53.20), IMO-AnswerBench (3 / 23, score 90), GPQA Diamond (42 / 270, score 90.40). This page also compares it with 3 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.
Benchmark Results
Benchmark Results
Coding and Software Engineer
4 evaluationsAI Agent - Tool Usage
3 evaluationsText Embedding
3 evaluationsCompetitor Comparison
Benchmark scores for Hy3 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Hy3Current | GLM-5.2 | Qwen3.7 Max | MiniMax M3 |
|---|---|---|---|---|
HLE 综合评估 | 53.20Thinking Level · High | Tools | 54.70Thinking Enabled | Tools | 53.50Thinking Enabled | Tools | -- |
GPQA Diamond 科学与综合推理 | 90.40Thinking Level · High | 91.86Thinking Level · High | 92.40Thinking Level · High | 81.31Standard Mode |
DeepSWE 编程与软件工程 | 28.00Thinking Level · High | Tools | 44.00Deep Thinking Mode | Tools | -- | -- |
SWE-bench Multilingual 编程与软件工程 | 75.80Thinking Level · High | Tools | -- | 78.30Thinking Enabled | Tools | -- |
SWE-Bench Pro - Public 编程与软件工程 | 57.90Thinking Level · High | Tools | 62.10Thinking Enabled | Tools | 60.60Thinking Enabled | Tools | 59.00Thinking Enabled | Tools |
SWE-bench Verified 编程与软件工程 | 78.00Thinking Level · High | Tools | -- | 80.40Thinking Enabled | Tools | -- |
BrowseComp AI Agent - 信息收集 | 84.20Thinking Level · High | Tools | -- | -- | 83.50Thinking Enabled | Tools |
MCP-Atlas AI Agent - 工具使用 | 79.10Thinking Level · High | Tools | 76.80Thinking Enabled | Tools | 76.40Thinking Enabled | Tools | 74.20Thinking Enabled | Tools |
Terminal-Bench 2.1 AI Agent - 工具使用 | 71.70Thinking Level · High | Tools | 81.00Thinking Level · High | Tools | -- | 66.00Thinking Enabled | Tools |
Tool Decathlon AI Agent - 工具使用 | 48.50Thinking Level · High | Tools | 48.20Thinking Enabled | Tools | -- | -- |
Context Arena 文本向量检索 | 66.79Thinking Level · High | 72.34Thinking Level · High | 56.01Standard Mode | 51.15Thinking Enabled |
IMO-AnswerBench 数学推理 | 90.00Thinking Level · High | 91.00Thinking Enabled | 90.00Thinking Level · High | -- |
Standard API Pricing: Hy3 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.
These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Hy3 | 腾讯AI实验室 | ¥1.2 / 1M tokens | ¥4 / 1M tokens | <= 16000 |
GLM-5.2 | 智谱AI | $1.4 / 1M tokens | $4.4 / 1M tokens | — |
Qwen3.7 Max | 阿里巴巴 | ¥12 / 1M tokens | ¥36 / 1M tokens | — |
MiniMax M3 | MiniMaxAI | ¥2.1 / 1M tokens | ¥8.4 / 1M tokens | — |
Version History
How each version of the Hy3 series stacks up on benchmark tests
Standard API Pricing Across the Hy3 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Hy3 | 腾讯AI实验室 | ¥1.2 / 1M tokens | ¥4 / 1M tokens | <= 16000 |