DataLearner logo

Hy3 Pre Benchmark Details

Hy3 Pre currently shows benchmark results led by τ²-Bench - Telecom (51 / 264, score 92.70), GPQA Diamond (137 / 461, score 86.70), Terminal Bench Hard (82 / 244, score 34.10). This page also compares it with 3 competitor models, including performance and pricing views when available.

Benchmark Results

Hy3 Pre

Benchmark Results

Thinking
Tool usage

Knowledge Exams

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
Standard Mode
7
446 / 568
HLE
Thinking Enabled
27.80
233 / 568

Scientific Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
73.20
300 / 461
GPQA Diamond
Thinking Enabled
86.70
137 / 461
CritPt
Standard Mode
0.30
186 / 204
CritPt
Thinking Enabled
4.60
106 / 204

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
67.50
143 / 264
τ²-Bench - Telecom
Thinking EnabledTools
92.70
51 / 264

Instruction Following

2 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
48
166 / 282
IF Bench
Thinking Enabled
63.10
110 / 282

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Standard ModeTools
31.80
100 / 244
Terminal Bench Hard
Thinking EnabledTools
34.10
82 / 244

Competitor Comparison

Benchmark scores for Hy3 Pre compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkHy3 PreCurrentGLM 5.1Qwen3.6-Max-PreviewDeepSeek-V4-Flash
HLE
Accuracy
Knowledge Exams
27.80Thinking Enabled
52.30Thinking Enabled | Tools
50.20Thinking Enabled | Tools
51.50Thinking Level · High | Tools
CritPt
Score
Scientific Reasoning
4.60Thinking Enabled
4.60Thinking Enabled
3.70Thinking Enabled
7.10Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
86.70Thinking Enabled
86.20Thinking Enabled
90.40Thinking Level · High
89.40Thinking Level · High
τ²-Bench - Telecom
Accuracy
Service Workflows
92.70Thinking Enabled | Tools
97.70Thinking Enabled | Tools
95.90Thinking Enabled | Tools
95.60Thinking Level · High | Tools
IF Bench
Accuracy
Instruction Following
63.10Thinking Enabled
76.30Thinking Enabled
76.60Thinking Enabled
79.20Thinking Level · High
Terminal Bench Hard
Accuracy
Agentic Development
34.10Thinking Enabled | Tools
43.20Thinking Enabled | Tools
43.90Thinking Enabled | Tools
38.60Thinking Level · High | Tools

Standard API Pricing: Hy3 Pre vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Qwen3.6-Max-Preview: Base price applies to <= 128
ModelSupplierStandard inputStandard outputBase price applies to
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—
Qwen3.6-Max-Preview
阿里巴巴$1.3 / 1M tokens$7.8 / 1M tokens<= 128
DeepSeek-V4-Flash
DeepSeek-AI$0.14 / 1M tokens$0.28 / 1M tokens—