AI Model Leaderboards
Live rankings across ARC-AGI-2, HLE, AIME 2025, SWE-bench Verified, and more — browse composite scores or drill into math, coding, and agent categories.
Composite Rankings
There is no single, universally agreed-upon comprehensive AI model ranking, so we selected two representative leaderboards that approach the question from different angles. Artificial Analysis Intelligence Index aggregates scores from 10 standardized benchmarks (coding, math, reasoning, etc.) to measure objective capability. LMArena (formerly Chatbot Arena) ranks models by Elo ratings derived from anonymous crowd-sourced A/B voting, reflecting real-world user preference. Together they offer both an objective and a subjective perspective.
AA Intelligence Index
Full rankingComposite of 10 standardized benchmarks across coding, math, science, reasoning, and agentic tasks.
Updated 2026-09-13
LMArena Text Generation
Full rankingElo ratings from anonymous crowdsourced A/B voting, reflecting real user preference for response quality.
Updated 2026-09-11

Recent Rank Changes
Risers, decliners, and new entrants across the coding, math, and agent leaderboards over the last 30 days.
Coding
Full ranking- DeepSeek V4.1 FlashNew#1·DeepSeek-AI·3471
- Gemini 3 Deep Think February 2026 UpgradeUpdated#2·Google Deep Mind·3455
- DeepSeek V3.2 SpecialeUpdated#3·DeepSeek-AI·1395.3
- Google Gemma 4 31BUpdated#4·DeepMind·1115
- GPT Opensources 120BUpdated#5·OpenAI·1060.72
- GPT Opensources 20BUpdated#6·OpenAI·984.58
- OpenAI o4 - miniUpdated#7·OpenAI·957.67
- Google Gemma 4 26B A4BUpdated#8·DeepMind·897.55
Math
Full ranking- Grok 4 HeavyUpdated#1·xAI·100
- Gemini 2.5 Deep ThinkUpdated#2·Google Deep Mind·99.2
- GPT-5 CodexUpdated#3·OpenAI·98.7
- Anthropic Claude Opus 4.6Updated#4·Anthropic·98.69
- Step 3.5 FlashUpdated#5·StepFunAI·98.55
- GPT-5-ProUpdated#6·OpenAI·98.35
- Kimi K2 ThinkingUpdated#7·Moonshot AI·97.87
- Kimi k1.5 (Long-CoT)Updated#8·Moonshot AI·96.2
Per-Benchmark Rankings
Filter by math, coding, agent, and more. Switch benchmarks below or jump into a category leaderboard for the full ranking. View all benchmarks.
Recommended models
Ranked by Terminal Bench 2.0LLM Performance Results
Data source: DataLearnerAIClick any row to open the model page. Tick the checkboxes to compare up to 4 models side by side. Scores shown are the best result across all evaluation modes.
Leaderboard FAQ
Where does the leaderboard data come from?
Scores are aggregated from primary sources: official model cards, technical reports, papers, vendor blog posts, and reproducible third-party evaluations. Each row links back to the underlying model detail page where the source is cited.
Why do scores for the same model differ across benchmarks?
Each benchmark measures a different capability — reasoning (HLE, ARC-AGI-2), math (AIME, FrontierMath), coding (SWE-bench Verified), agent use (τ²-Bench), and so on. A model tuned for one capability may perform very differently on another, which is exactly why we surface per-benchmark scores rather than a single number.
How often is the leaderboard updated?
Data is revalidated every 5 minutes, and new models or evaluation results are added as soon as they are published. The "Updated on" indicator at the top of the page reflects the most recent data refresh.
How should I read the composite ranking?
The composite view aggregates a model's standing across multiple core benchmarks. It is a useful first filter, but for production decisions you should drill into the specific benchmark closest to your workload — for example, SWE-bench Verified for coding agents, or τ²-Bench for tool-use scenarios.
How do I compare an open-source model with a closed API model?
Use the license filter at the top to mix open and closed models in the same view, then look at the same benchmark column for both. Beyond raw scores, consider total cost of ownership: API pricing for closed models vs. self-hosting cost for open weights.
As of 2026-09, AA Intelligence Index leaders include Claude Fable 5.1, Claude Fable 5.1, GPT-6 Astra (max), based on 10 standardized capability benchmarks.
On the user-preference side, LMArena Text Generation currently ranks Claude Fable 5, Claude Opus 4.6 (high), Opus 4.7 (high) near the top via anonymous A/B voting.
Scroll down for per-benchmark breakdowns in math, coding, and agent categories. See Data Methodology for scoring details, or browse LLM Blogs for in-depth commentary.
Leading model developers
View all 102 organizationsJump to a developer to explore its full model lineup, series, and product lines.
阿里巴巴
OpenAI
Google Deep Mind
Facebook AI研究实验室
智谱AI
DeepSeek-AI
MistralAI
Google Research
Anthropic
Microsoft Azure
百度
Stability AI
字节跳动Seed团队Model comparisons
Head-to-head write-ups: what the benchmark gap actually means, and which model fits which job.
DeepSeek V4.1 Flash 与 DeepSeek V4 Flash 怎么选?限时预览与成熟 API 版本对比
DeepSeek V4.1 Flash 是限时 API 预览版,官方确认其原生多模态,但暂未公开独立评测与完整规格;DeepSeek V4 Flash 则已有 1M 上下文、284B MoE、开源权重和 0731 API 版本的 Agent 评测。若生产工作负载需要稳定接口、可复现实测或本地部署,应优先 V4 Flash;V4.1 Flash 更适合在有效期内做隔离试用,不应仅因版本号更高就默认替换现有方案。
GPT-6 Astra vs GPT-5.6 Sol:按评测版本、测试配置与任务成本比较
更新至 2026-09-05 核验快照:AA Coding Agent Index v1.4 的 Codex max 对照为 67 对 65,Sol 历史 80 分不与新版本混比。Astra 在多项官方 Agent 和长上下文评测领先,但 AA 编程测试中也有 Sol 更强的子项及更短的耗时。标准 token 单价相差 2.5 倍,不等于任务账单相差 2.5 倍。
GPT-6 Astra对比Claude Fable 5.1,谁更强,哪个价格更有优势
GPT-6 Astra 与 Claude Fable 5.1 是 OpenAI 和 Anthropic 在 2026 年推出的两款旗舰级模型,两者都面向复杂推理、编程、Agent 和长时程知识工作。从目前可直接对比的 8 项评测来看,GPT-6 Astra 以 5 项领先、3 项落后的成绩取得小幅整体优势,在 ARC-AGI、AutomationBench、Terminal-Bench 及科学工具任务上表现更突出;Claude Fable 5.1 则在 HLE、AA Intelligence Index 和 OSWorld 2.0 等知识推理与计算机操作评测中占优。两款模型均提供约 100 万 token 上下文和 128K 最大输出,API 基础价格也同为每百万 token 输入 10 美元、输出 50 美元,因此实际选型更取决于任务类型、Agent 工作流以及缓存使用方式,而非单纯的价格或综合分数。
Claude Fable 5.1 / Fable 5 / Opus 5 与 GPT-5.6 Sol 四方对比:评测、价格与选型
四款模型共有的 4 项 0–100 量表评测里,Fable 5.1 平均 58.8 分居首,Opus 5 50.3、Fable 5 46.7、GPT-5.6 Sol 42.5;计算机操作(OSWorld 2.0 partial)上 Fable 5.1 以 77.9 对 75.4 小幅领先 Opus 5,但它的价格是 Opus 5 的两倍、Sol 的 2.5 倍。
Explore more
The leaderboard covers benchmarked models. Browse the full catalog by model, organization, or benchmark.
Browse every tracked model — filter by organization, type, and release date, not just benchmark scores.
BrowseExplore the labs and companies behind these models and their full model lineups.
BrowseDive into each benchmark — what it measures, how it scores, and the full ranking.
Browse







