Artificial Analysis Intelligence Index
Artificial Analysis Intelligence Index aggregates multiple rigorous benchmarks to compare AI model intelligence across coding, reasoning, science, tool use, and agentic tasks.
Top Model
Claude Opus 5 (max)
Top Score
63
Model Count
256
Data version
2026年08月18日
Data source: Artificial Analysis
Ranking Table
| Rank | Model | Intelligence Index | Organization |
|---|---|---|---|
Claude Opus 5 (max)Anthropic | 63 | Anthropic | |
Claude Opus 5 (xhigh)Anthropic | 63 | Anthropic | |
Claude Fable 5Anthropic | 62 | Anthropic | |
| 4 | Claude Opus 5 (high)Anthropic | 61 | Anthropic |
| 5 | GPT-5.6 Sol (max)OpenAI | 61 | OpenAI |
| 6 | 61 | xAI | |
| 7 | Kimi K3 (max)Moonshot AI | 60 | Moonshot AI |
| 8 | GPT-5.6 Sol (xhigh)OpenAI | 59 | OpenAI |
| 9 | Claude Opus 5 (medium)Anthropic | 59 | Anthropic |
| 10 | Qwen3.8-Max阿里巴巴 | 58 | 阿里巴巴 |
| 11 | Qwen3.8 2.4T A95BAlibaba | 58 | Alibaba |
| 12 | GPT-5.6 Sol (high)OpenAI | 57 | OpenAI |
| 13 | Muse Spark 1.2 (xhigh)Meta | 57 | Meta |
| 14 | GPT-5.6 Terra (max)OpenAI | 57 | OpenAI |
| 15 | Gemini 3.7 Flash (high)Google Deep Mind | 56 | Google Deep Mind |
| 16 | 56 | xAI | |
| 17 | GPT-5.6 Sol (medium)OpenAI | 56 | OpenAI |
| 18 | Claude Sonnet 5 (max)Anthropic | 55 | Anthropic |
| 19 | Gemini 3.7 Flash (medium)Google Deep Mind | 53 | Google Deep Mind |
| 20 | DeepSeek-V4-Pro (max)DeepSeek-AI | 53 | DeepSeek-AI |
| 21 | GPT-5.6 Terra (xhigh)OpenAI | 53 | OpenAI |
| 22 | GLM-5.2 (max)智谱AI | 53 | 智谱AI |
| 23 | Claude Opus 5 (low)Anthropic | 52 | Anthropic |
| 24 | GPT-5.6 Luna (max)OpenAI | 52 | OpenAI |
| 25 | Qwen3.8-27B阿里巴巴 | 52 | 阿里巴巴 |
| 26 | DeepSeek-V4-Flash (max)DeepSeek-AI | 52 | DeepSeek-AI |
| 27 | Gemini 3.6 FlashGoogle Deep Mind | 52 | Google Deep Mind |
| 28 | Gemini 3.7 Flash (low)Google Deep Mind | 51 | Google Deep Mind |
| 29 | GPT-5.6 Sol (low)OpenAI | 51 | OpenAI |
| 30 | GPT-5.6 Terra (high)OpenAI | 50 | OpenAI |
| 31 | GPT-5.6 Luna (xhigh)OpenAI | 50 | OpenAI |
| 32 | Kimi K3 (low)Moonshot AI | 48 | Moonshot AI |
| 33 | Gemini 3.1 Pro PreviewGoogle Deep Mind | 48 | Google Deep Mind |
| 34 | Motif 3Motif Technologies | 47 | Motif Technologies |
| 35 | GPT-5.6 Luna (high)OpenAI | 47 | OpenAI |
| 36 | GPT-5.6 Terra (medium)OpenAI | 47 | OpenAI |
| 37 | Gemini 3.5 Flash (medium)Google Deep Mind | 47 | Google Deep Mind |
| 38 | GPT-5.3 Codex (xhigh)OpenAI | 46 | OpenAI |
| 39 | MiniMax M3MiniMaxAI | 45 | MiniMaxAI |
| 40 | Motif 3 (Beta)Motif Technologies | 45 | Motif Technologies |
| 41 | DeepSeek-V4-Pro (max)DeepSeek-AI | 45 | DeepSeek-AI |
| 42 | DeepSeek-V4-Pro (high)DeepSeek-AI | 44 | DeepSeek-AI |
| 43 | Kimi K2.7 CodeMoonshot AI | 43 | Moonshot AI |
| 44 | MiMo-V2.5-ProXiaomi | 43 | Xiaomi |
| 45 | Claude Sonnet 5 (non-reasoning)Anthropic | 43 | Anthropic |
| 46 | InklingThinking Machines Lab | 42 | Thinking Machines Lab |
| 47 | Hy3腾讯AI实验室 | 42 | 腾讯AI实验室 |
| 48 | GPT-5.6 Sol (non-reasoning)OpenAI | 42 | OpenAI |
| 49 | Nex-N2-ProNex AGI | 42 | Nex AGI |
| 50 | Solar Pro 4Upstage | 42 | Upstage |
| 51 | GPT-5.6 Terra (low)OpenAI | 41 | OpenAI |
| 52 | Inkling SmallThinking Machines | 41 | Thinking Machines |
| 53 | JT-4.1 Flash 236B A21BChina Mobile | 40 | China Mobile |
| 54 | Agnes 2.5 Pro AlphaSapiens AI | 40 | Sapiens AI |
| 55 | Qwen3.7-Plus阿里巴巴 | 39 | 阿里巴巴 |
| 56 | GPT-5.6 Luna (medium)OpenAI | 39 | OpenAI |
| 57 | Nemotron 3 UltraNVIDIA | 38 | NVIDIA |
| 58 | MiMo-V2.5Xiaomi | 38 | Xiaomi |
| 59 | Ling 3.0 FlashInclusionAI | 38 | InclusionAI |
| 60 | Qwen3.6-27B阿里巴巴 | 38 | 阿里巴巴 |
| 61 | Gemini 3.5 Flash-LiteGoogle Deep Mind | 37 | Google Deep Mind |
| 62 | Solar Open2 250BUpstage | 37 | Upstage |
| 63 | Nova 2 Omni(Preview)亚马逊 | 37 | 亚马逊 |
| 64 | 37 | xAI | |
| 65 | 36 | xAI | |
| 66 | MiMo-V2-OmniXiaomi | 36 | Xiaomi |
| 67 | Gemini 3.5 Flash (minimal)Google Deep Mind | 36 | Google Deep Mind |
| 68 | Claude Sonnet 4.6 (Non-reasoning, Low Effort)Anthropic | 35 | Anthropic |
| 69 | Muse Glimmer-30B (high)Facebook AI研究实验室 | 35 | Facebook AI研究实验室 |
| 70 | Qwen2-57B-A14B阿里巴巴 | 35 | 阿里巴巴 |
| 71 | GLM-5.2智谱AI | 35 | 智谱AI |
| 72 | GPT-5.6 Terra (non-reasoning)OpenAI | 35 | OpenAI |
| 73 | Qwen3.5-397B-A17B阿里巴巴 | 34 | 阿里巴巴 |
| 74 | Gemini 2.0 Flash ExperimentalDeepMind | 34 | DeepMind |
| 75 | LongCat 2.0LongCat | 34 | LongCat |
| 76 | GPT-5.6 Luna (low)OpenAI | 34 | OpenAI |
| 77 | KAT-Coder-Pro V2KwaiKAT | 34 | KwaiKAT |
| 78 | Qwen3.5-122B-A10B阿里巴巴 | 33 | 阿里巴巴 |
| 79 | Qwen3.5-397B-A17B阿里巴巴 | 33 | 阿里巴巴 |
| 80 | Qwen3.6-35B-A3B阿里巴巴 | 32 | 阿里巴巴 |
| 81 | DeepSeek-V4-ProDeepSeek-AI | 32 | DeepSeek-AI |
| 82 | Ring-2.6-1TInclusionAI | 32 | InclusionAI |
| 83 | G9v3-39A5BAI9Stars | 32 | AI9Stars |
| 84 | Qwen3.5-Omni-Plus阿里巴巴 | 31 | 阿里巴巴 |
| 85 | Qwen3.6-27B阿里巴巴 | 31 | 阿里巴巴 |
| 86 | OpenAI o3OpenAI | 31 | OpenAI |
| 87 | K-EXAONE 2.0LG AI Research | 31 | LG AI Research |
| 88 | Step 3.7 FlashStepFunAI | 31 | StepFunAI |
| 89 | Mistral Medium 3.5MistralAI | 30 | MistralAI |
| 90 | Haiku 4.5Anthropic | 30 | Anthropic |
| 91 | Gemma 4 31BDeepMind | 30 | DeepMind |
| 92 | DeepSeek-V4-FlashDeepSeek-AI | 29 | DeepSeek-AI |
| 93 | GPT-5.5 InstantOpenAI | 29 | OpenAI |
| 94 | JT-35B-FlashChina Mobile | 29 | China Mobile |
| 95 | MiMo-V2.5-ProXiaomi | 28 | Xiaomi |
| 96 | Qwen3.5-122B-A10B阿里巴巴 | 28 | 阿里巴巴 |
| 97 | GPT-5.6 Luna (non-reasoning)OpenAI | 27 | OpenAI |
| 98 | Doubao Seed CodeByteDance Seed | 26 | ByteDance Seed |
| 99 | Gemma 4 26B A4BDeepMind | 26 | DeepMind |
| 100 | Nemotron 3 SuperNVIDIA | 26 | NVIDIA |
| 101 | MiMo-V2-FlashXiaomi | 25 | Xiaomi |
| 102 | 25 | xAI | |
| 103 | Qwen3.6-35B-A3B阿里巴巴 | 25 | 阿里巴巴 |
| 104 | Ling 3.0 TinyInclusionAI | 25 | InclusionAI |
| 105 | Qwen3.5-35B-A3B阿里巴巴 | 24 | 阿里巴巴 |
| 106 | Haiku 4.5Anthropic | 24 | Anthropic |
| 107 | GPT OSS 120B (high)OpenAI | 24 | OpenAI |
| 108 | Nemotron 3.5 LightningNVIDIA | 24 | NVIDIA |
| 109 | C4AI Command A (202503)CohereAI | 23 | CohereAI |
| 110 | K-EXAONELG AI Research | 22 | LG AI Research |
| 111 | ERNIE 5.0 Thinking Preview百度 | 22 | 百度 |
| 112 | Gemma 4 31BDeepMind | 22 | DeepMind |
| 113 | Gemma 4 12BGoogle | 22 | |
| 114 | Nova 2 Pro(Preview) (medium)亚马逊 | 22 | 亚马逊 |
| 115 | Mercury 2Inception | 22 | Inception |
| 116 | Qwen3.5-9B阿里巴巴 | 22 | 阿里巴巴 |
| 117 | Qwen3-Coder-Next阿里巴巴 | 21 | 阿里巴巴 |
| 118 | Nova 2 Omni(Preview) (medium)亚马逊 | 21 | 亚马逊 |
| 119 | Apriel-v1.6-15B-ThinkerServiceNow | 21 | ServiceNow |
| 120 | Nova 2 Lite (high)亚马逊 | 21 | 亚马逊 |
| 121 | Qwen3.5-9B阿里巴巴 | 21 | 阿里巴巴 |
| 122 | EXAONE 4.5 33BLG AI Research | 21 | LG AI Research |
| 123 | Gemma 4 26B A4BDeepMind | 20 | DeepMind |
| 124 | Qwen3.5 4BAlibaba | 20 | Alibaba |
| 125 | North Mini CodeCohere | 20 | Cohere |
| 126 | Nova 2 Pro(Preview) (low)亚马逊 | 20 | 亚马逊 |
| 127 | Mistral Small 4Mistral | 20 | Mistral |
| 128 | Devstral 2Mistral | 19 | Mistral |
| 129 | Nova 2 Lite (medium)亚马逊 | 19 | 亚马逊 |
| 130 | Qwen3.5-Omni-Flash阿里巴巴 | 19 | 阿里巴巴 |
| 131 | JT-MINIChina Mobile | 19 | China Mobile |
| 132 | Trinity Large ThinkingArcee AI | 19 | Arcee AI |
| 133 | HyperNova 60B 2605Multiverse Computing | 18 | Multiverse Computing |
| 134 | Magistral Medium 1.2Mistral | 18 | Mistral |
| 135 | Nova 2 Lite (low)亚马逊 | 18 | 亚马逊 |
| 136 | Nemotron Cascade 2 30B A3BNVIDIA | 18 | NVIDIA |
| 137 | Devstral Small 2Mistral | 18 | Mistral |
| 138 | K2 Think V2MBZUAI | 17 | MBZUAI |
| 139 | LongCat Flash LiteLongCat | 17 | LongCat |
| 140 | HyperCLOVA X SEED Think (32B)Naver | 17 | Naver |
| 141 | K-EXAONELG AI Research | 17 | LG AI Research |
| 142 | Qwen3-Next阿里巴巴 | 17 | 阿里巴巴 |
| 143 | Nova 2 Omni(Preview) (low)亚马逊 | 17 | 亚马逊 |
| 144 | Mi:dm K 2.5 ProKorea Telecom | 17 | Korea Telecom |
| 145 | G9v3-3BAI9Stars | 16 | AI9Stars |
| 146 | Qwen3.5 4BAlibaba | 16 | Alibaba |
| 147 | Mistral Large 3MistralAI | 16 | MistralAI |
| 148 | INTELLECT-3Prime Intellect | 16 | Prime Intellect |
| 149 | Solar Open 100BUpstage | 15 | Upstage |
| 150 | GPT OSS 20B (high)OpenAI | 15 | OpenAI |
| 151 | Qwen3-Omni-30B-A3B阿里巴巴 | 15 | 阿里巴巴 |
| 152 | GPT OSS 120B (low)OpenAI | 15 | OpenAI |
| 153 | Nemotron 3 NanoNVIDIA | 15 | NVIDIA |
| 154 | Solar Pro 3Upstage | 14 | Upstage |
| 155 | Llama 4 MaverickFacebook AI研究实验室 | 14 | Facebook AI研究实验室 |
| 156 | Nova 2 Pro(Preview)亚马逊 | 14 | 亚马逊 |
| 157 | GPT OSS 20B (low)OpenAI | 14 | OpenAI |
| 158 | K2-V2 (high)MBZUAI | 14 | MBZUAI |
| 159 | Qwen3-Next阿里巴巴 | 14 | 阿里巴巴 |
| 160 | DiffusionGemma 26B A4BGoogle | 13 | |
| 161 | Gemma 4 12B (Non-reasoning)Google | 13 | |
| 162 | Motif-2-12.7BMotif Technologies | 13 | Motif Technologies |
| 163 | Nova PremierAmazon | 13 | Amazon |
| 164 | K2-V2 (medium)MBZUAI | 12 | MBZUAI |
| 165 | Llama Nemotron Super 49B v1.5Meta | 12 | Meta |
| 166 | Celeris-1Celeris | 12 | Celeris |
| 167 | Mistral Small 4Mistral | 12 | Mistral |
| 168 | Tri-21B-ThinkTrillion Labs | 12 | Trillion Labs |
| 169 | Gemma 4 E4BDeepMind | 12 | DeepMind |
| 170 | MiniCPM5-1BOpenBMB | 12 | OpenBMB |
| 171 | Sarvam 105B (high)Sarvam | 12 | Sarvam |
| 172 | Nova 2 Lite亚马逊 | 12 | 亚马逊 |
| 173 | MiniCPM5-1BOpenBMB | 12 | OpenBMB |
| 174 | Magistral Small 1.2Mistral | 11 | Mistral |
| 175 | Ministral 3 14BMistralAI | 11 | MistralAI |
| 176 | Nanbeige4.1-3BNanbeige | 11 | Nanbeige |
| 177 | EXAONE 4.0 32BLG AI Research | 11 | LG AI Research |
| 178 | Nova 2 Omni(Preview)亚马逊 | 10 | 亚马逊 |
| 179 | Llama 4 ScoutFacebook AI研究实验室 | 10 | Facebook AI研究实验室 |
| 180 | Hermes 4 70BNous Research | 10 | Nous Research |
| 181 | Falcon-H1R-7BTII UAE | 10 | TII UAE |
| 182 | Gemma 4 E2BDeepMind | 10 | DeepMind |
| 183 | Qwen3-Omni-30B-A3B阿里巴巴 | 10 | 阿里巴巴 |
| 184 | Step3 VL 10BStepFun | 9 | StepFun |
| 185 | Llama3.3-70B-InstructFacebook AI研究实验室 | 9 | Facebook AI研究实验室 |
| 186 | Ministral 3 8BMistralAI | 9 | MistralAI |
| 187 | Llama Nemotron UltraNVIDIA | 9 | NVIDIA |
| 188 | ERNIE-4.5-300B-A47B百度 | 9 | 百度 |
| 189 | Hermes 4 405BNous Research | 9 | Nous Research |
| 190 | NVIDIA Nemotron Nano 12B v2 VLNVIDIA | 9 | NVIDIA |
| 191 | Gemma 4 E4BDeepMind | 9 | DeepMind |
| 192 | Granite 4.1 30BIBM | 9 | IBM |
| 193 | NVIDIA Nemotron Nano 9B V2NVIDIA | 9 | NVIDIA |
| 194 | Hermes 4 405BNous Research | 9 | Nous Research |
| 195 | Nemotron 3 Nano 4BNVIDIA | 9 | NVIDIA |
| 196 | Llama Nemotron Super 49B v1.5Meta | 9 | Meta |
| 197 | K2-V2 (low)MBZUAI | 8 | MBZUAI |
| 198 | Kimi Linear 48B A3B InstructKimi | 8 | Kimi |
| 199 | Llama3.1-405BFacebook AI研究实验室 | 8 | Facebook AI研究实验室 |
| 200 | LFM2.5-8B-A1BLiquid AI | 8 | Liquid AI |
| 201 | Ring-flash-2.0InclusionAI | 8 | InclusionAI |
| 202 | Olmo 3.1 32B ThinkAI2 | 8 | AI2 |
| 203 | C4AI Command A (202503)CohereAI | 7 | CohereAI |
| 204 | Qwen3.5 2BAlibaba | 7 | Alibaba |
| 205 | Llama 3.1 Nemotron 70BNVIDIA | 7 | NVIDIA |
| 206 | Nemotron 3 NanoNVIDIA | 7 | NVIDIA |
| 207 | NVIDIA Nemotron Nano 9B V2NVIDIA | 7 | NVIDIA |
| 208 | Ministral 3 3BMistral | 7 | Mistral |
| 209 | Hermes 4 70BNous Research | 7 | Nous Research |
| 210 | Granite 4.1 8BIBM | 6 | IBM |
| 211 | Sarvam 30B (high)Sarvam | 6 | Sarvam |
| 212 | Olmo 3.1 32B InstructAI2 | 6 | AI2 |
| 213 | Gemma 4 E2BDeepMind | 6 | DeepMind |
| 214 | R1 1776Perplexity | 6 | Perplexity |
| 215 | Llama 3.2-Vision-90BFacebook AI研究实验室 | 6 | Facebook AI研究实验室 |
| 216 | Phi-4-mini-instruct (3.8B)Microsoft Azure | 6 | Microsoft Azure |
| 217 | EXAONE 4.0 32BLG AI Research | 6 | LG AI Research |
| 218 | Qwen3.5 2BAlibaba | 5 | Alibaba |
| 219 | Qwen3.5 0.8BAlibaba | 5 | Alibaba |
| 220 | DeepHermes 3 - Mistral 24BNous Research | 5 | Nous Research |
| 221 | Jamba 1.7 LargeAI21 Labs | 5 | AI21 Labs |
| 222 | Granite 4.0 H SmallIBM | 5 | IBM |
| 223 | Qwen3-Omni-30B-A3B阿里巴巴 | 5 | 阿里巴巴 |
| 224 | LFM2 24B A2BLiquid AI | 5 | Liquid AI |
| 225 | Phi-4-reasoningMicrosoft Azure | 5 | Microsoft Azure |
| 226 | Amazon Nova Micro亚马逊 | 4 | 亚马逊 |
| 227 | Granite 4.1 3BIBM | 4 | IBM |
| 228 | NVIDIA Nemotron Nano 12B v2 VLNVIDIA | 4 | NVIDIA |
| 229 | Phi-4-multimodal-instruct Microsoft Azure | 4 | Microsoft Azure |
| 230 | MiniCPM-V 4.6 1.3BOpenBMB | 4 | OpenBMB |
| 231 | Jamba Reasoning 3BAI21 Labs | 4 | AI21 Labs |
| 232 | Gemini 3.0 FlashGoogle Deep Mind | 4 | Google Deep Mind |
| 233 | Olmo 3 7B ThinkAI2 | 4 | AI2 |
| 234 | Molmo 7B-DAllen Institute for AI | 3 | Allen Institute for AI |
| 235 | Llama 3.2-Vision-11BFacebook AI研究实验室 | 3 | Facebook AI研究实验室 |
| 236 | Qwen3.5 0.8BAlibaba | 3 | Alibaba |
| 237 | Exaone 4.0 1.2BLG AI Research | 3 | LG AI Research |
| 238 | Olmo 3 7BAI2 | 2 | AI2 |
| 239 | Exaone 4.0 1.2BLG AI Research | 2 | LG AI Research |
| 240 | LFM2.5-1.2B-ThinkingLiquid AI | 2 | Liquid AI |
| 241 | Jamba 1.7 MiniAI21 Labs | 2 | AI21 Labs |
| 242 | LFM2 2.6BLiquid AI | 2 | Liquid AI |
| 243 | LFM2.5-1.2B-InstructLiquid AI | 2 | Liquid AI |
| 244 | Granite 4.0 H 1BIBM | 2 | IBM |
| 245 | Gemma 3-270MGoogle Deep Mind | 2 | Google Deep Mind |
| 246 | Apertus 70B InstructSwiss AI | 2 | Swiss AI |
| 247 | Granite 4.0 MicroIBM | 2 | IBM |
| 248 | DeepHermes 3 - Llama-3.1 8BNous Research | 2 | Nous Research |
| 249 | Granite 4.0 1BIBM | 2 | IBM |
| 250 | Molmo2-8BAI2 | 2 | AI2 |
| 251 | LFM2 8B A1BLiquid AI | 1 | Liquid AI |
| 252 | LFM2.5-VL-1.6BLiquid AI | 1 | Liquid AI |
| 253 | Granite 4.0 350MIBM | 1 | IBM |
| 254 | Tiny Aya GlobalCohere | 1 | Cohere |
| 255 | Apertus 8B InstructSwiss AI | 1 | Swiss AI |
| 256 | Granite 4.0 H 350MIBM | 1 | IBM |
Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.
Benchmark Components (Intelligence Index v4.0)
The Intelligence Index aggregates 10 rigorous benchmarks to provide a holistic measure of AI capabilities, preventing narrow specialization.
GDPval-AA
Agentic real-world tasks
τ²-Bench
Agentic tool use
Terminal-Bench
Agentic coding
SciCode
Coding proficiency
AA-LCR
Long context reasoning
AA-Omniscience
Knowledge & hallucination
IFBench
Instruction following
Humanity's Last Exam
Reasoning & knowledge
GPQA Diamond
Scientific reasoning
CritPt
Physics reasoning
FAQ
What is the Artificial Analysis Intelligence Index?▼
The Artificial Analysis Intelligence Index v4.0 is a composite benchmark that aggregates performance across 10 challenging evaluations — spanning mathematics, science, coding, agentic tasks, and reasoning — to measure AI capabilities holistically. It is designed to prevent narrow specialization and provide a single score for tracking progress.
How is the Intelligence Index calculated?▼
The index aggregates scores from 10 benchmarks: GDPval-AA (agentic real-world tasks), τ²-Bench (tool use), Terminal-Bench Hard (agentic coding), SciCode (coding), AA-LCR (long context reasoning), AA-Omniscience (knowledge & hallucination), IFBench (instruction following), Humanity's Last Exam (reasoning), GPQA Diamond (scientific reasoning), and CritPt (physics). All tests are independently run by Artificial Analysis on standardized hardware.
How does this differ from LMArena?▼
LMArena rankings are based on crowdsourced user votes (Elo ratings from blind A/B tests), reflecting subjective human preferences. The Artificial Analysis Intelligence Index uses standardized automated benchmarks with objective scoring, measuring technical capabilities across specific domains. Both perspectives are valuable — LMArena captures real-world user experience, while AA Intelligence Index provides reproducible technical measurements.
Where can I find the original data?▼
The original leaderboard and detailed methodology are available at artificialanalysis.ai. The Intelligence Index methodology is documented at Intelligence Index page.




















