Terminal-Bench-Science 0.1 evaluates scientific research agents on 70 expert-authored terminal tasks spanning life, physical, Earth, mathematical, and engineering sciences. Each model-agent system is run three times per task. Scores therefore measure the named model and agent harness together, not an agent-independent model capability.
Browse the latest scores, model modes, release dates, and parameter sizes for Terminal-Bench-Science 0.1.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
![]() Claude Opus 5 Thinking Level · MaxTools | 30.00 | 2026-07-24 | Unknown | Closed | |
![]() GPT-5.6 Sol Thinking Level · MaxTools | 22.40 | 2026-06-26 | Unknown | Closed | |
![]() Claude Fable 5 Thinking Level · MaxTools | 21.40 | 2026-06-09 | Unknown | Closed | |
4 | ![]() Claude Opus 4.8 Thinking Level · MaxTools | 10.50 | 2026-05-28 | Unknown | Closed |
5 | ![]() GPT-5.6 Terra Thinking Level · MaxTools | 8.60 | 2026-06-26 | Unknown | Closed |
6 | ![]() GLM-5.3 Thinking Level · MaxTools | 8.10 | 2026-08-14 | 744B | — |
7 | ![]() Kimi K3 Thinking Level · MaxTools | 7.10 | 2026-07-16 | 2800B | Free Commercial |
8 | Grok 4.6 Thinking Level · HighTools | 7.10 | 2026-08-12 | Unknown | Closed |
9 | ![]() GPT-5.6 Luna Thinking Level · MaxTools | 3.30 | 2026-06-26 | Unknown | Closed |
In the official leaderboard reviewed on August 30, 2026, Claude Opus 5 with Claude Code leads at 30.0%, followed by GPT-5.6 Sol with Codex at 22.4% and Claude Fable 5 with Claude Code at 21.4%.