See key specs and per-benchmark scores for each model/mode. Scroll horizontally for all columns. 当前对比 2 个模型的评测数据与核心参数。

Claude Opus 5
Anthropic
Best overall
Claude Opus 5 · 1216.60
Best single
Claude Opus 5 · GDPval-AA v2 1861.00
Modality coverage
Claude Opus 5 · 2 modalities
Head to head
3
Benchmarks
3
Wins
0
Losses
+84.63
Average diff
Compare benchmark results across thinking modes and tool usage.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
Complete scores for each model/mode across selected benchmarks.
3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Claude Opus 5 | Grok 4.6 |
|---|---|---|
DeepSWE 编程与软件工程 | 68.80Thinking Level · High | Tools | 65.90Thinking Level · High | Tools |
AA-Briefcase 生产力知识 | 1720.00Thinking Level · High | Tools | 1577.00Thinking Level · High | Tools |
GDPval-AA v2 生产力知识 | 1861.00Thinking Level · High | Tools | 1753.00Thinking Level · High | Tools |
Side-by-side input/output token pricing
Licensing, MoE architecture, and multi-modality support.
| Features & specs | Claude Opus 5Anthropic | Grok 4.6xAI |
|---|---|---|
Core specsRelease | 2026-07-24 | 2026-08-12 |
Context length | 1M | 500K |
Max output | 128000 | Not provided |
MoE | No | No |
LicenseCode Open Source | Closed Source | Closed Source |
Weights Open Source | Closed Source | Closed Source |
Commercial use | 不开源 | 不开源 |
Modality supportText Input/Output | / | / |
Image Input/Output | / | / |
ResourcesPaper / report | Introducing Claude Opus 5 | Introducing Grok 4.6 |