Claude 3 Opus was the most capable model in Anthropic's Claude 3 family, released on March 4, 2024. It outperformed GPT-4 on many benchmarks at launch, added image input and offered a 200K-token context window, and was known for strong writing and nuanced reasoning. It has since been retired.
数据优先来自官方发布(GitHub、Hugging Face、论文),其次为评测基准官方结果,最后为第三方评测机构数据。 了解数据收集方法
模型基本信息
开源和体验地址
官方介绍与博客
API接口信息
| 类型 | 适用条件 | 输入 | 输出 |
|---|---|---|---|
| 文本 | - | $15.00/ 1M tokens | $75.00/ 1M tokens |
| 图像 | - | $15.00/ 1M tokens | — |
| 类型 | 适用条件 | 输入 | 输出 |
|---|---|---|---|
| 文本 | - | $7.50/ 1M tokens | $37.50/ 1M tokens |
| 图像 | - | $7.50/ 1M tokens | — |
| 类型 | 有效期 | 写入 | 读取 |
|---|---|---|---|
| 文本 | - | $18.75/ 1M tokens | $1.50/ 1M tokens |
| 图像 | - | $18.75/ 1M tokens 缓存状态 = write | — |
| 图像 | - | — | $1.50/ 1M tokens 缓存状态 = hit |
表中「—」表示该模态在此方向不计费,或供应商未公开对应价格。
评测结果
Claude3-Opus 当前已收录的代表性评测结果包括 GSM8K(9 / 68,得分 95)、MMLU(24 / 117,得分 86.80)、HumanEval(21 / 101,得分 84.90)。 本页还汇总了参数规模、上下文长度与 API 价格,便于结合评测结果与部署约束一起判断模型适配度。
发布机构
模型解读
Claude3-Opus是Anthropic公司发布的第三代多模态大语言模型。第三代的Claude-3模型包含3个版本,这里说的Claude3-Opus是其中能力最强的模型。各项评测人任务结果都非常好,甚至超过了GPT-4。
在多模态方面,Claude3-Opus也有强大的能力。

Claude2最受诟病的就是无效的拒绝回答。由于Anthropic在对齐方面做了严格的工作,导致Claude2.1经常出现拒绝回答的情况。在Claude3-Opus上。Anthropic做了改进,在内部测试中,Claude2.1错误地拒绝比例大概在26%左右,而Claude3-Opus上这个比例下降到了11%,进步明显!

DataLearner 官方微信
欢迎关注 DataLearner 官方微信,获得最新 AI 技术推送
