Model comparisons
Head-to-head write-ups: what the benchmark gap actually means, and which model fits which job.
GPT-6 Astra vs GPT-5.6 Sol:按评测版本、测试配置与任务成本比较
更新至 2026-09-05 核验快照:AA Coding Agent Index v1.4 的 Codex max 对照为 67 对 65,Sol 历史 80 分不与新版本混比。Astra 在多项官方 Agent 和长上下文评测领先,但 AA 编程测试中也有 Sol 更强的子项及更短的耗时。标准 token 单价相差 2.5 倍,不等于任务账单相差 2.5 倍。
GPT-6 Astra对比Claude Fable 5.1,谁更强,哪个价格更有优势
GPT-6 Astra 与 Claude Fable 5.1 是 OpenAI 和 Anthropic 在 2026 年推出的两款旗舰级模型,两者都面向复杂推理、编程、Agent 和长时程知识工作。从目前可直接对比的 8 项评测来看,GPT-6 Astra 以 5 项领先、3 项落后的成绩取得小幅整体优势,在 ARC-AGI、AutomationBench、Terminal-Bench 及科学工具任务上表现更突出;Claude Fable 5.1 则在 HLE、AA Intelligence Index 和 OSWorld 2.0 等知识推理与计算机操作评测中占优。两款模型均提供约 100 万 token 上下文和 128K 最大输出,API 基础价格也同为每百万 token 输入 10 美元、输出 50 美元,因此实际选型更取决于任务类型、Agent 工作流以及缓存使用方式,而非单纯的价格或综合分数。
Claude Fable 5.1 / Fable 5 / Opus 5 与 GPT-5.6 Sol 四方对比:评测、价格与选型
四款模型共有的 4 项 0–100 量表评测里,Fable 5.1 平均 58.8 分居首,Opus 5 50.3、Fable 5 46.7、GPT-5.6 Sol 42.5;计算机操作(OSWorld 2.0 partial)上 Fable 5.1 以 77.9 对 75.4 小幅领先 Opus 5,但它的价格是 Opus 5 的两倍、Sol 的 2.5 倍。
AI Model List FAQ
How often is this AI model list updated?
New models, version bumps, pricing changes, and benchmark results are added as soon as they are published — typically within hours of an official announcement. Once the data is in, the page refreshes within 5 minutes so visitors always see the latest information.
What do the "Open Source" filter options mean?
A model can be open source (weights publicly available) yet still restrict commercial use through its license — for example, some Llama variants prohibit large-scale commercial deployment. The three options reflect this: "Free for commercial use" means no restrictions for production; "Paid commercial" means a license fee applies; "Not for commercial use" means the model cannot legally be used in commercial products.
How do I pick the right model for my use case?
Start by filtering on capability (chat, coding, reasoning, multimodal) to narrow the field, then compare the top candidates on the benchmark closest to your workload. For production, also weigh API pricing, context length, and license — a model that wins on benchmarks may still lose on total cost of ownership.
Which organizations' models are included?
The list covers mainstream model publishers worldwide — including OpenAI, Anthropic, Google, Meta, Mistral, DeepSeek, Alibaba (Qwen), Zhipu (GLM), Moonshot (Kimi), and many others. Use the publisher filter to narrow to a specific lab.
Can I see the full benchmark results for a specific model?
Yes — click any model card to open its detail page, where you will find the complete benchmark table, parameter sizes, context window, license, API pricing, and links to the official paper, Hugging Face, and GitHub repository.



















