DataLearner logo

Hy3 Benchmark Details

Hy3 currently shows benchmark results led by HLE (21 / 185, score 53.20), IMO-AnswerBench (3 / 23, score 90), GPQA Diamond (42 / 270, score 90.40). This page also compares it with 3 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Hy3

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
High
37
78 / 185
HLE
HighTools
53.20
21 / 185

Other

1 evaluations
Benchmark / mode
Score
Rank/total
90.40
42 / 270

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
78
24 / 114
75.80
10 / 27
57.90
20 / 59
DeepSWE
HighTools
28
30 / 31

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighToolsInternet
84.20
10 / 54

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
79.10
11 / 40
71.70
33 / 47
48.50
3 / 10

Text Embedding

3 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
25.35
118 / 126
66.10
70 / 126
66.79
69 / 126

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
90
3 / 23

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
High
73.40
10 / 27

Competitor Comparison

Benchmark scores for Hy3 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkHy3CurrentGLM-5.2Qwen3.7 MaxMiniMax M3
HLE
综合评估
53.20Thinking Level · High | Tools
54.70Thinking Enabled | Tools
53.50Thinking Enabled | Tools
--
GPQA Diamond
科学与综合推理
90.40Thinking Level · High
91.86Thinking Level · High
92.40Thinking Level · High
81.31Standard Mode
DeepSWE
编程与软件工程
28.00Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
--
--
SWE-bench Multilingual
编程与软件工程
75.80Thinking Level · High | Tools
--
78.30Thinking Enabled | Tools
--
SWE-Bench Pro - Public
编程与软件工程
57.90Thinking Level · High | Tools
62.10Thinking Enabled | Tools
60.60Thinking Enabled | Tools
59.00Thinking Enabled | Tools
SWE-bench Verified
编程与软件工程
78.00Thinking Level · High | Tools
--
80.40Thinking Enabled | Tools
--
BrowseComp
AI Agent - 信息收集
84.20Thinking Level · High | Tools
--
--
83.50Thinking Enabled | Tools
MCP-Atlas
AI Agent - 工具使用
79.10Thinking Level · High | Tools
76.80Thinking Enabled | Tools
76.40Thinking Enabled | Tools
74.20Thinking Enabled | Tools
Terminal-Bench 2.1
AI Agent - 工具使用
71.70Thinking Level · High | Tools
81.00Thinking Level · High | Tools
--
66.00Thinking Enabled | Tools
Tool Decathlon
AI Agent - 工具使用
48.50Thinking Level · High | Tools
48.20Thinking Enabled | Tools
--
--
Context Arena
文本向量检索
66.79Thinking Level · High
72.34Thinking Level · High
56.01Standard Mode
51.15Thinking Enabled
IMO-AnswerBench
数学推理
90.00Thinking Level · High
91.00Thinking Enabled
90.00Thinking Level · High
--
1 additional benchmarks remain in the chart above.

Standard API Pricing: Hy3 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Hy3
Supplier: 腾讯AI实验室
Standard input: ¥1.2 / 1M tokens
Standard output: ¥4 / 1M tokens
Base price applies to <= 16000
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Qwen3.7 Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Hy3
腾讯AI实验室¥1.2 / 1M tokens¥4 / 1M tokens<= 16000
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Qwen3.7 Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens

Version History

How each version of the Hy3 series stacks up on benchmark tests

No benchmark data matches the selected filters.

Standard API Pricing Across the Hy3 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Hy3: Base price applies to <= 16000
ModelSupplierStandard inputStandard outputBase price applies to
Hy3
腾讯AI实验室¥1.2 / 1M tokens¥4 / 1M tokens<= 16000