DataLearner logo

Text-to-Video Arena Leaderboard

The latest AI video generation leaderboard based on Text-to-Video Arena anonymous user voting. Covers Elo scores, confidence intervals, and vote counts for leading video models.

Top Model

minimax-h3

Top Score

1,455

Model Count

43

Data version

2026年08月02日

Data source: LM Arena

About This Leaderboard

This leaderboard ranks AI text-to-video models by generation quality. Data comes from LMArena's Text-to-Video Arena track, evaluated through anonymous blind testing by real users.

Methodology Overview

Blind testing: Users submit text descriptions, two anonymous models generate videos, and users vote for the better result.

Elo scoring: Based on the Bradley-Terry model. Higher scores indicate stronger user preference for that model's video output.

Diverse generation scenarios: Covers natural landscapes, human motion, creative animation, product showcases, and more.

DataLearner provides in-depth analysis on top of the raw data, linking leaderboard models to the DataLearner model database so you can quickly access model details, API pricing, benchmark scores, and more.

Origin:AllChina
Leaderboard snapshot month:

Ranking Table

RankModelScore95% CIVotesOrganizationLicense
4MiniMaxminimax-h3MiniMax1,455+/-191,060MiniMaxminimax-h3-community-license-agreement
5Alibaba-ATHhappyhorse-1.0Alibaba-ATH1,428+/-1321,979Alibaba-ATHProprietary
13Alibabawan2.7-t2vAlibaba1,347+/-915,246AlibabaProprietary
28MiniMaxAIHailuo 2.3MiniMaxAI1,202+/-772,127MiniMaxAIProprietary
29MiniMaxAIHailuo 2.3MiniMaxAI1,198+/-139,365MiniMaxAIProprietary
31MiniMaxAIHailuo 2.3MiniMaxAI1,180+/-129,333MiniMaxAIProprietary

Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.

2026-08 Market Signals

Current Best (SOTA)

01

Veo 3.1 Audio 1080p

02

Veo 3.1 Fast-Audio 1080p

03

Sora-2-Pro

Best China Model

Wan2.6-T2V

Seedance-V1.5-Pro

Kling-2.6-Pro

Best Open Model

  • Wan-V2.2-A14B
  • Kandinsky-5.0-T2V-Pro
  • Mochi-V1

FAQ

01

How does Text-to-Video Arena rank models?

Rankings are based on side-by-side anonymous votes. Users enter the same prompt, compare outputs from two hidden models, and choose the better video. Elo-style scoring then aggregates those comparisons into a leaderboard.

02

What is audio-video sync, and why does it matter?

Audio-video sync means generated sound effects or speech match the motion and timing in the video. It matters because synchronized audio can make generated clips usable with less post-production work.

03

What use cases are text-to-video models good for?

Common uses include short-form video creation, marketing assets, e-commerce product clips, storyboarding, game cinematics, and educational demos.

04

Which models support the longest generation length?

Long generation limits change quickly by product tier and release. In practice, check the current model documentation and compare both maximum duration and quality consistency across longer clips.