DataLearner logo

Muse Glimmer-30B Benchmark Details

Muse Glimmer-30B currently shows benchmark results led by IF Bench (19 / 282, score 77), Creative Writing (21 / 106, score 1789.70), AA-LCR (47 / 174, score 80). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Muse Glimmer-30B

Benchmark Results

Thinking
Tool usage
Internet

Knowledge Exams

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
High
22
274 / 568

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
83.50
185 / 461

Repository Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
76
37 / 116
51.20
44 / 62

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1789.70
21 / 106

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
77
19 / 282

Long Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
High
80
47 / 174

Mathematics

1 evaluations
Benchmark / mode
Score
Rank/total
94.70
15 / 30

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
65.90
22 / 28

Agentic Development

1 evaluations
Benchmark / mode
Score
Rank/total
51.70
134 / 199

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
75.50
25 / 44

Cross-industry Work

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
HighTools
953
96 / 110

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
DeepSearchQA
HighToolsInternet
74.60
3 / 3

Visual Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
74
99 / 229

Documents & Charts

2 evaluations
Benchmark / mode
Score
Rank/total
78.80
19 / 19
75.80
3 / 3

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
43.60
98 / 134

Service Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
τ³-Banking
HighTools
23.50
82 / 167

Competitor Comparison

Benchmark scores for Muse Glimmer-30B compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkMuse Glimmer-30BCurrentQwen3.8-27BGemma 4 31BQwen3.6-35B-A3B
HLE
Accuracy
Knowledge Exams
22.00Thinking Level · High
33.90Thinking Level · Extra High
26.50Thinking Enabled | Tools
22.20Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
83.50Thinking Level · High
90.50Thinking Level · Extra High
84.30Thinking Enabled
84.85Standard Mode
SWE-Bench Pro - Public
Accuracy
Repository Engineering
51.20Thinking Level · High | Tools
61.70Thinking Enabled | Tools
--
49.50Thinking Enabled
SWE-bench Verified
Accuracy
Repository Engineering
76.00Thinking Level · High | Tools
--
--
73.40Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
Writing
1789.70Standard Mode
1668.40Standard Mode
1366.00Standard Mode
--
IF Bench
Accuracy
Instruction Following
77.00Thinking Level · High
--
75.60Thinking Enabled
64.40Thinking Enabled
AA-LCR
Accuracy
Long Reasoning
80.00Thinking Level · High
82.00Thinking Level · Extra High
46.70Standard Mode
64.30Standard Mode
AIME 2026
Accuracy
Mathematics
94.70Thinking Level · High
--
89.20Thinking Enabled
92.70Thinking Enabled
OSWorld-Verified
Accuracy
Desktop Workflows
65.90Thinking Level · High | Tools
84.30Thinking Enabled | Tools
--
--
Terminal-Bench 2.1
Accuracy
Agentic Development
51.70Thinking Level · High | Tools
79.80Thinking Level · Extra High | Tools
43.40Thinking Enabled | Tools
44.90Thinking Enabled | Tools
GDPval-AA v2
Elo score
Cross-industry Work
953.00Thinking Level · High | Tools
1463.00Thinking Level · Extra High | Tools
691.00Standard Mode | Tools
953.00Standard Mode | Tools
MMMU-Pro
Accuracy
Visual Understanding
74.00Thinking Level · High
76.30Thinking Level · Extra High
76.90Thinking Level · High
75.00Thinking Enabled
4 additional benchmarks remain in the chart above.

Standard API Pricing: Muse Glimmer-30B vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.

Version History

How each version of the Muse Glimmer-30B series stacks up on benchmark tests

Muse Glimmer-30BMuse Spark 1.1Muse Spark
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkMuse Glimmer-30BCurrentMuse Spark 1.1Muse Spark
HLE
Accuracy
Knowledge Exams
22.00Thinking Level · High
62.10Thinking Enabled | Tools
58.00Deep Thinking Mode
GPQA Diamond
Accuracy
Scientific Reasoning
83.50Thinking Level · High
89.80Thinking Level · Extra High
89.50Thinking Enabled
SWE-Bench Pro - Public
Accuracy
Repository Engineering
51.20Thinking Level · High | Tools
61.50Thinking Enabled | Tools
--
SWE-bench Verified
Accuracy
Repository Engineering
76.00Thinking Level · High | Tools
--
77.40Thinking Enabled | Tools
Creative Writing
Elo、大模型评判两两对战
Writing
1789.70Standard Mode
1916.10Standard Mode
--
IF Bench
Accuracy
Instruction Following
77.00Thinking Level · High
--
75.90Thinking Enabled
AA-LCR
Accuracy
Long Reasoning
80.00Thinking Level · High
77.70Thinking Level · Extra High
--
OSWorld-Verified
Accuracy
Desktop Workflows
65.90Thinking Level · High | Tools
80.80Thinking Enabled | Tools
--
Terminal-Bench 2.1
Accuracy
Agentic Development
51.70Thinking Level · High | Tools
80.00Thinking Enabled | Tools
62.20Thinking Enabled | Tools
MCP-Atlas
Pass rate / claim coverage
Tool Orchestration
75.50Thinking Level · High | Tools
88.10Thinking Enabled | Tools
82.20Standard Mode | Tools
GDPval-AA v2
Elo score
Cross-industry Work
953.00Thinking Level · High | Tools
1294.00Thinking Level · Extra High | Tools
--
MMMU-Pro
Accuracy
Visual Understanding
74.00Thinking Level · High
--
80.40Thinking Enabled
2 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Muse Glimmer-30B Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.