Image-to-Video Arena Leaderboard
The latest AI image-to-video leaderboard based on anonymous Arena voting. Covers Elo scores, confidence intervals, and vote counts for leading video animation models.
Top Model
Seedance 2.0
Top Score
1,478
Model Count
43
Data version
2026年08月02日
Data source: LM Arena
About This Leaderboard
This leaderboard ranks AI image-to-video models by animation quality. Data comes from LMArena's Image-to-Video Arena track, evaluated through anonymous blind testing by real users.
Methodology Overview
Blind testing: Users upload an image, two anonymous models generate animated videos, and users vote for the more natural result.
Elo scoring: Based on the Bradley-Terry model, scientifically measuring each model's relative strength in image-to-video tasks.
Diverse animation scenarios: Covers portrait animation, landscape motion, object transformation, artistic creation, and more.
DataLearner provides in-depth analysis on top of the raw data, linking leaderboard models to the DataLearner model database so you can quickly access model details, API pricing, benchmark scores, and more.
Ranking Table
| Rank | Model | Score | 95% CI | Votes | Organization | License |
|---|---|---|---|---|---|---|
| Seedance 2.0字节跳动Seed团队 | 1,478 | +/-10 | 100,246 | 字节跳动Seed团队 | Proprietary | |
minimax-h3MiniMax | 1,476 | +/-19 | 1,190 | MiniMax | minimax-h3-community-license-agreement | |
| 1,462 | +/-6 | 87,626 | SpaceXAI | Proprietary | ||
| 4 | gemini-omni-flashGoogle | 1,462 | +/-7 | 44,400 | Proprietary | |
| 5 | happyhorse-1.0Alibaba-ATH | 1,442 | +/-10 | 70,454 | Alibaba-ATH | Proprietary |
| 6 | wan2.7-i2vAlibaba | 1,427 | +/-7 | 47,521 | Alibaba | Proprietary |
| 7 | 1,417 | +/-6 | 504,475 | xAI | Proprietary | |
| 8 | Veo 3.1 Generate (Preview)Google Deep Mind | 1,397 | +/-11 | 25,114 | Google Deep Mind | Proprietary |
| 9 | Veo 3.1 Generate (Preview)Google Deep Mind | 1,390 | +/-9 | 52,988 | Google Deep Mind | Proprietary |
| 10 | Veo 3.1 Fast (Preview)Google Deep Mind | 1,384 | +/-9 | 99,749 | Google Deep Mind | Proprietary |
| 11 | 1,383 | +/-8 | 19,407 | xAI | Proprietary | |
| 12 | Veo 3.1 Fast (Preview)Google Deep Mind | 1,371 | +/-10 | 54,306 | Google Deep Mind | Proprietary |
| 13 | vidu-q3-proShengshu | 1,361 | +/-8 | 36,667 | Shengshu | Proprietary |
| 14 | kling-v3-proKlingAI | 1,359 | +/-7 | 156,829 | KlingAI | Proprietary |
| 15 | Veo 3.1 Generate (Preview)Google Deep Mind | 1,330 | +/-12 | 32,382 | Google Deep Mind | Proprietary |
| 16 | Veo 3.1 Fast (Preview)Google Deep Mind | 1,324 | +/-9 | 41,213 | Google Deep Mind | Proprietary |
| 17 | Wan2.1-T2V-14B阿里巴巴 | 1,322 | +/-10 | 16,926 | 阿里巴巴 | Proprietary |
| 18 | Wan2.6 I2V阿里巴巴 | 1,311 | +/-8 | 94,087 | 阿里巴巴 | Proprietary |
| 19 | Seedance 2.0字节跳动Seed团队 | 1,307 | +/-7 | 244,917 | 字节跳动Seed团队 | Proprietary |
| 20 | pixverse-v5.6 Proprietary | 1,299 | +/-7 | 126,647 | — | — |
| 21 | Kling 2.5 Turbo昆仑万维 | 1,293 | +/-8 | 201,921 | 昆仑万维 | Proprietary |
| 22 | Kling 2.5 Turbo昆仑万维 | 1,274 | +/-12 | 3,791 | 昆仑万维 | Proprietary |
| 23 | Seedance 2.0字节跳动Seed团队 | 1,272 | +/-8 | 34,027 | 字节跳动Seed团队 | Proprietary |
| 24 | 1,260 | +/-6 | 311,267 | MiniMaxAI | Proprietary | |
| 25 | Veo 3.1 Fast (Preview)Google Deep Mind | 1,256 | +/-10 | 26,297 | Google Deep Mind | Proprietary |
| 26 | Veo 3.1 Generate (Preview)Google Deep Mind | 1,256 | +/-10 | 26,105 | Google Deep Mind | Proprietary |
| 27 | p-video Proprietary | 1,243 | +/-16 | 23,356 | — | — |
| 28 | vidu-q2-turboShengshu | 1,242 | +/-17 | 2,506 | Shengshu | Proprietary |
| 29 | Kling 2.5 Turbo昆仑万维 | 1,234 | +/-8 | 29,849 | 昆仑万维 | Proprietary |
| 30 | 1,227 | +/-11 | 21,751 | MiniMaxAI | Proprietary | |
| 31 | Kling 2.5 Turbo昆仑万维 | 1,227 | +/-8 | 29,952 | 昆仑万维 | Proprietary |
| 32 | ray-3 LumaAI | 1,225 | +/-19 | 1,588 | AI | Proprietary |
| 33 | 1,222 | +/-9 | 21,782 | MiniMaxAI | Proprietary | |
| 34 | vidu-q2-proShengshu | 1,222 | +/-17 | 2,608 | Shengshu | Proprietary |
| 35 | Hunyuan-A13B-Instruct腾讯AI实验室 | 1,196 | +/-15 | 5,475 | 腾讯AI实验室 | tencent-hunyuan-community |
| 36 | 1,193 | +/-11 | 22,549 | MiniMaxAI | Proprietary | |
| 37 | Seedance 2.0字节跳动Seed团队 | 1,184 | +/-8 | 33,754 | 字节跳动Seed团队 | Proprietary |
| 38 | Wan2.1-T2V-14B阿里巴巴 | 1,169 | +/-10 | 27,067 | 阿里巴巴 | Apache 2.0 |
| 39 | Veo 3.1 Generate (Preview)Google Deep Mind | 1,164 | +/-16 | 10,319 | Google Deep Mind | Proprietary |
| 40 | ltx-2-19b ltx-2-community-license-agreement | 1,151 | +/-6 | 210,709 | — | — |
| 41 | ray2 LumaAI | 1,106 | +/-16 | 9,527 | AI | Proprietary |
| 42 | runway-gen4-turboRunway | 1,051 | +/-13 | 6,811 | Runway | Proprietary |
| 43 | pika-v2.2Pika | 996 | +/-14 | 8,655 | Pika | Proprietary |
Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.
2026-08 Market Signals
Current Best (SOTA)
Grok Imagine Video 720p
Veo 3.1 Audio 1080p
Veo 3.1 Audio
Best China Model
Vidu-Q3-Pro
Wan2.5-I2V-Preview
Kling-2.6-Pro
Best Open Model
- •Wan-V2.2-A14B
- •LTX-2-19B
- •Pika-V2.2
FAQ
What is the difference between image-to-video and text-to-video?
Text-to-video generates a clip from a prompt alone. Image-to-video starts from a reference image, which gives stronger control over subject identity, composition, and visual style.
Which model should I use to animate old photos?
For portrait animation, compare models on facial expression stability, motion naturalness, and identity preservation. Specialized lip-sync tools may be better when speech alignment is the main requirement.
How can I keep characters consistent?
Use a strong reference image as the first frame, keep the prompt specific, and avoid large changes in clothing, camera angle, or style unless the model supports identity conditioning.
What is first-frame fidelity?
First-frame fidelity measures how closely the generated video preserves the uploaded reference image at the beginning of the clip. Higher fidelity means the video feels like motion extending from the source image rather than a loose reinterpretation.



