DataLearner logo

Mistral Medium 3.5 Benchmark Details

Mistral Medium 3.5 currently shows benchmark results led by Terminal Bench Hard (86 / 244, score 33.30), ECI (90 / 167, score 141.42), τ³-Banking (112 / 167, score 15.10). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Mistral Medium 3.5

Benchmark Results

Thinking
Tool usage
Internet

Agentic Development

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Thinking EnabledTools
33.30
86 / 244

Memory & Persistence

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
17.64
125 / 126
Context Arena
Thinking Level · High
32.05
114 / 126

Capability Indices

3 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
141.42
90 / 167
Vals Index
Thinking Level · HighTools
17.95
40 / 42

Office & Business

1 evaluations
Benchmark / mode
Score
Rank/total
AA-Briefcase
Standard ModeTools
519.11
3 / 3

Visual Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Thinking Enabled
64.90
154 / 229

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Enabled
39.58
108 / 134

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
SAGE
Thinking Level · HighTools
37.61
46 / 64
τ³-Banking
Thinking EnabledTools
15.10
112 / 167

Legal

3 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Thinking EnabledTools
69.10
34 / 44
Legal Research Bench
Thinking Level · HighTools
9.13
40 / 42
Harvey's Legal Agent Benchmark
Thinking Level · HighTools
0.42
35 / 43

Finance

2 evaluations
Benchmark / mode
Score
Rank/total
Finance Agent v2
Thinking Level · HighTools
32.10
40 / 43
EMB (Excel Modeling)
Thinking Level · HighTools
11.19
41 / 42

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Enabled
2.80
106 / 122

Data Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
Standard ModeToolsInternet
12.50
23 / 29

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
Code Migration
Thinking Level · HighTools
5.13
41 / 43

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Thinking Level · HighTools
2.89
57 / 61

Mathematics

1 evaluations
Benchmark / mode
Score
Rank/total
ProofBench v1.1 (Lean 4)
Thinking Level · HighTools
9
28 / 28

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Thinking Level · HighTools
67.73
63 / 66
MedCode
Thinking Level · HighTools
33.75
56 / 64

Competitor Comparison

Benchmark scores for Mistral Medium 3.5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkMistral Medium 3.5CurrentQwen3.6-27BGLM-5.2Gemma 4 31B
Terminal Bench Hard
Accuracy
Agentic Development
33.30Thinking Enabled | Tools
34.80Thinking Enabled | Tools
50.80Thinking Level · High | Tools
36.40Thinking Enabled | Tools
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
32.05Thinking Level · High
82.17Thinking Enabled
72.34Thinking Level · High
--
ECI
ECI score (capability index, higher is better)
Capability Indices
141.42Thinking Level · High
146.47Thinking Level · High
151.76Thinking Level · High
--
Vals Index
跨行业任务准确率综合指数
Capability Indices
17.95Thinking Level · High | Tools
--
53.12Thinking Level · High | Tools
--
MMMU-Pro
Accuracy
Visual Understanding
64.90Thinking Enabled
74.60Thinking Enabled
--
76.90Thinking Level · High
SciCode
Score
Scientific Computing
39.58Thinking Enabled
42.80Thinking Enabled
51.20Thinking Level · High
45.50Thinking Enabled
τ³-Banking
Score
Service Workflows
15.10Thinking Enabled | Tools
16.70Thinking Enabled | Tools
37.11Thinking Level · Extra High | Tools
14.80Thinking Enabled | Tools
Harvey Lab-AA
Criterion pass rate (some entries report all-pass rate)
Legal
69.10Thinking Enabled | Tools
82.30Thinking Enabled | Tools
90.97Thinking Level · High | Tools
47.23Thinking Enabled | Tools
0.42Thinking Level · High | Tools
--
7.08Thinking Level · High | Tools
--
9.13Thinking Level · High | Tools
--
31.25Thinking Level · High | Tools
--
EMB (Excel Modeling)
Accuracy (%)
Finance
11.19Thinking Level · High | Tools
--
61.53Thinking Level · High | Tools
--
Finance Agent v2
Score
Finance
32.10Thinking Level · High | Tools
--
49.70Thinking Level · High | Tools
--
5 additional benchmarks remain in the chart above.

Standard API Pricing: Mistral Medium 3.5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Mistral Medium 3.5
MistralAI$1.5 / 1M tokens$7.5 / 1M tokens—
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—

Version History

How each version of the Mistral Medium 3.5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

5 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkMistral Medium 3.5CurrentMistral Large 3Devstral Medium
Terminal Bench Hard
Accuracy
Agentic Development
33.30Thinking Enabled | Tools
15.90Standard Mode | Tools
9.10Standard Mode | Tools
MMMU-Pro
Accuracy
Visual Understanding
64.90Thinking Enabled
55.70Standard Mode
--
SciCode
Score
Scientific Computing
39.58Thinking Enabled
36.60Standard Mode
--
τ³-Banking
Score
Service Workflows
15.10Thinking Enabled | Tools
5.80Standard Mode | Tools
--
GDP.pdf
Task pass rate (%)
Documents & Charts
2.80Thinking Enabled
1.20Standard Mode
--

Single-Benchmark Version Trend

Viewing: Terminal Bench Hard · Agentic Development

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Mistral Medium 3.5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Mistral Medium 3.5
MistralAI$1.5 / 1M tokens$7.5 / 1M tokens—
Mistral Large 3
MistralAI$0.5 / 1M tokens$1.5 / 1M tokens—
Devstral Medium
MistralAI$0.4 / 1M tokens$2 / 1M tokens—