DataLearner logo

GPT-3.5 Benchmark Details

GPT-3.5 currently shows benchmark results led by GSM8K (37 / 70, score 57.10), C-Eval (27 / 48, score 54.40), MMLU (73 / 124, score 70).

Benchmark Results

GPT-3.5

Benchmark Results

Thinking

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
70
73 / 124
C-Eval
Standard Mode
54.40
27 / 48

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
GSM8K
Standard Mode
57.10
37 / 70

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
MBPP
Standard Mode
52.20
63 / 96
HumanEval
Standard Mode
48.10
87 / 140