DataLearner logo

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index aggregates multiple rigorous benchmarks to compare AI model intelligence across coding, reasoning, science, tool use, and agentic tasks.

Top Model

Kimi K3 (max)

Top Score

60

Model Count

256

Data version

2026年08月18日

Data source: Artificial Analysis

Origin:AllChina
Leaderboard snapshot month:

Ranking Table

RankModelIntelligence IndexOrganization
7Moonshot AIKimi K3 (max)Moonshot AI60Moonshot AI
11AlibabaQwen3.8 2.4T A95BAlibaba58Alibaba
20DeepSeek-AIDeepSeek-V4-Pro (max)DeepSeek-AI53DeepSeek-AI
26DeepSeek-AIDeepSeek-V4-Flash (max)DeepSeek-AI52DeepSeek-AI
32Moonshot AIKimi K3 (low)Moonshot AI48Moonshot AI
39MiniMaxAIMiniMax M3MiniMaxAI45MiniMaxAI
41DeepSeek-AIDeepSeek-V4-Pro (max)DeepSeek-AI45DeepSeek-AI
42DeepSeek-AIDeepSeek-V4-Pro (high)DeepSeek-AI44DeepSeek-AI
43Moonshot AIKimi K2.7 CodeMoonshot AI43Moonshot AI
81DeepSeek-AIDeepSeek-V4-ProDeepSeek-AI32DeepSeek-AI
88StepFunAIStep 3.7 FlashStepFunAI31StepFunAI
92DeepSeek-AIDeepSeek-V4-FlashDeepSeek-AI29DeepSeek-AI
98ByteDance SeedDoubao Seed CodeByteDance Seed26ByteDance Seed
124AlibabaQwen3.5 4BAlibaba20Alibaba
146AlibabaQwen3.5 4BAlibaba16Alibaba
184StepFunStep3 VL 10BStepFun9StepFun
198KimiKimi Linear 48B A3B InstructKimi8Kimi
204AlibabaQwen3.5 2BAlibaba7Alibaba
218AlibabaQwen3.5 2BAlibaba5Alibaba
219AlibabaQwen3.5 0.8BAlibaba5Alibaba
236AlibabaQwen3.5 0.8BAlibaba3Alibaba

Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.

Benchmark Components (Intelligence Index v4.0)

The Intelligence Index aggregates 10 rigorous benchmarks to provide a holistic measure of AI capabilities, preventing narrow specialization.

GDPval-AA
Agentic real-world tasks
τ²-Bench
Agentic tool use
Terminal-Bench
Agentic coding
SciCode
Coding proficiency
AA-LCR
Long context reasoning
AA-Omniscience
Knowledge & hallucination
IFBench
Instruction following
Humanity's Last Exam
Reasoning & knowledge
GPQA Diamond
Scientific reasoning
CritPt
Physics reasoning

FAQ

What is the Artificial Analysis Intelligence Index?
The Artificial Analysis Intelligence Index v4.0 is a composite benchmark that aggregates performance across 10 challenging evaluations — spanning mathematics, science, coding, agentic tasks, and reasoning — to measure AI capabilities holistically. It is designed to prevent narrow specialization and provide a single score for tracking progress.
How is the Intelligence Index calculated?
The index aggregates scores from 10 benchmarks: GDPval-AA (agentic real-world tasks), τ²-Bench (tool use), Terminal-Bench Hard (agentic coding), SciCode (coding), AA-LCR (long context reasoning), AA-Omniscience (knowledge & hallucination), IFBench (instruction following), Humanity's Last Exam (reasoning), GPQA Diamond (scientific reasoning), and CritPt (physics). All tests are independently run by Artificial Analysis on standardized hardware.
How does this differ from LMArena?
LMArena rankings are based on crowdsourced user votes (Elo ratings from blind A/B tests), reflecting subjective human preferences. The Artificial Analysis Intelligence Index uses standardized automated benchmarks with objective scoring, measuring technical capabilities across specific domains. Both perspectives are valuable — LMArena captures real-world user experience, while AA Intelligence Index provides reproducible technical measurements.
Where can I find the original data?
The original leaderboard and detailed methodology are available at artificialanalysis.ai. The Intelligence Index methodology is documented at Intelligence Index page.