DataLearner logo

M2.1 Benchmark Details

M2.1 currently shows benchmark results led by MMLU Pro (7 / 133, score 88), SWE-bench Verified (40 / 114, score 74.80), IF Bench (15 / 33, score 70).

Benchmark Results

M2.1

Benchmark Results

Thinking
Tool usage

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Thinking Mode
88
7 / 133
HLE
Thinking Mode
22
117 / 181

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
81
111 / 224

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Mode
74.80
40 / 114
SWE-Bench Pro - Public
Thinking ModeTools
32.60
56 / 57

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
81
56 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
34.70
49 / 67

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking ModeTools
87
22 / 35

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking ModeTools
70
15 / 33

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeTools
47.40
44 / 54

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Thinking ModeTools
47.90
37 / 48

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
84.30
19 / 38