DataLearner logo

Muse Spark 1.1 Benchmark Details

Muse Spark 1.1 currently shows benchmark results led by HLE (6 / 565, score 62.10), MCP-Atlas (2 / 44, score 88.10), Creative Writing (8 / 106, score 1916.10). This page also compares it with 3 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Muse Spark 1.1

Benchmark Results

Thinking
Tool usage

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking ModeTools
62.10
6 / 565
HLE
Thinking Level · Extra High
46.20
77 / 565
AA Intelligence Index v4.3
Thinking Level · Extra HighTools
34.30
23 / 26
CritPt
Thinking Level · Extra High
15.10
55 / 201

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Extra High
89.80
91 / 463

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1916.10
8 / 106

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1536.36
12 / 35
SWE-Bench Pro - Public
Thinking ModeTools
61.50
14 / 62
SciCode
Thinking Level · Extra High
58.80
10 / 131
DeepSWE
Thinking ModeTools
53.32
57 / 86

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Thinking ModeTools
88.10
2 / 44
OSWorld-Verified
Thinking ModeTools
80.80
7 / 26
Terminal-Bench 2.1
Thinking ModeTools
80
59 / 194
Terminal-Bench 2.1
Thinking Level · Extra HighTools
76.20
75 / 194
Tool Decathlon
Thinking ModeTools
75.60
1 / 10
Terminal-Bench 4.0
Thinking Level · Extra HighTools
6.10
55 / 88

Text Embedding

4 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Thinking Level · Low
70.32
63 / 126
Context Arena
Thinking Level · Medium
74.38
51 / 126
Context Arena
Thinking Level · High
69.86
65 / 126
Context Arena
Thinking Level · Extra High
70.01
64 / 126

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Level · Extra High
77.70
70 / 171

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Level · Extra HighTools
1294
66 / 106
AA-Briefcase
Thinking Level · Extra HighTools
860
74 / 84

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ³-Banking
Thinking Level · Extra HighTools
40.46
30 / 167

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Level · Extra High
14.40
60 / 119

Other

1 evaluations
Benchmark / mode
Score
Rank/total
28.10
8 / 44

Competitor Comparison

Benchmark scores for Muse Spark 1.1 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkMuse Spark 1.1CurrentGPT-5.6 SolClaude Fable 5GLM-5.2
AA Intelligence Index v4.3
Index (0-100)
综合评估
34.30Thinking Level · Extra High | Tools
47.10Standard Mode | Tools
49.70Thinking Level · High
34.00Standard Mode | Tools
CritPt
Score
综合评估
15.10Thinking Level · Extra High
32.30Thinking Level · High
28.60Thinking Level · High
20.90Thinking Level · High
HLE
Accuracy
综合评估
62.10Thinking Enabled | Tools
49.50Thinking Level · High
59.00Deep Thinking Mode
54.70Thinking Enabled | Tools
GPQA Diamond
Accuracy
科学与综合推理
89.80Thinking Level · Extra High
94.60Thinking Level · High
85.86Thinking Level · High
91.86Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1916.10Standard Mode
1963.40Standard Mode
1934.60Standard Mode
1752.80Standard Mode
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
53.32Thinking Enabled | Tools
72.70Thinking Level · Extra High | Tools
69.91Deep Thinking Mode | Tools
44.00Deep Thinking Mode | Tools
SciCode
Score
编程与软件工程
58.80Thinking Level · Extra High
57.80Thinking Level · High
61.00Thinking Level · High
51.20Thinking Level · High
SWE-Bench Pro - Public
Accuracy
编程与软件工程
61.50Thinking Enabled | Tools
64.60Thinking Level · Extra High | Tools
80.30Deep Thinking Mode | Tools
62.10Thinking Enabled | Tools
Text Arena (Coding)
Arena Score
编程与软件工程
1536.36Standard Mode
1620.27Thinking Level · Extra High
--
1593.25Thinking Level · High
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
88.10Thinking Enabled | Tools
81.80Thinking Level · High | Tools
83.30Standard Mode | Tools
77.80Thinking Enabled | Tools
OSWorld-Verified
Accuracy
AI Agent - 工具使用
80.80Thinking Enabled | Tools
--
85.00Thinking Level · High | Tools
--
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
80.00Thinking Enabled | Tools
89.50Thinking Level · Extra High | Tools
88.00Deep Thinking Mode | Tools
81.00Thinking Level · High | Tools
9 additional benchmarks remain in the chart above.

Standard API Pricing: Muse Spark 1.1 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens
Claude Fable 5
Anthropic$10 / 1M tokens$50 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens

Version History

How each version of the Muse Spark 1.1 series stacks up on benchmark tests

Muse Spark 1.1Muse Spark
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkMuse Spark 1.1CurrentMuse Spark
CritPt
Score
综合评估
15.10Thinking Level · Extra High
11.30Thinking Enabled
HLE
Accuracy
综合评估
62.10Thinking Enabled | Tools
58.00Deep Thinking Mode
GPQA Diamond
Accuracy
科学与综合推理
89.80Thinking Level · Extra High
89.50Thinking Enabled
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
88.10Thinking Enabled | Tools
82.20Standard Mode | Tools
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
80.00Thinking Enabled | Tools
62.20Thinking Enabled | Tools
AA-Omniscience
AA-Omniscience Index; Accuracy; Hallucination Rate
真实性评估
28.10Thinking Level · High
7.20Thinking Level · High

Single-Benchmark Version Trend

Viewing: CritPt · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Muse Spark 1.1 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.