DataLearner logo

Arcada Labs Code Categories Arena Leaderboard

The latest AI design-code model leaderboard based on Arcada Labs Code Categories Arena anonymous user voting. Focused on Website, UI components, game development, and data visualization code generation.

Top Model

Kimi K3

Top Score

1394.00

Model Count

163

Data version

2026年09月08日

Data source: Arcada Labs

Origin:AllChina
Leaderboard snapshot month:

Ranking Table

RankModelScore95% CIVotesOrganizationLicense
Moonshot AIKimi K3Moonshot AI1394.00+/-9.65,844Moonshot AIOpen Source
14AlibabaQwen3.8 MaxAlibaba1311.00+/-9.35,698AlibabaOpen Source
16GLM-5.3-FlashZhipu AI1308.00+/-5.815,817Zhipu AIOpen Source
25Moonshot AIKimi K2.6Moonshot AI1290.00+/-4.133,999Moonshot AIOpen Source
35Moonshot AIKimi K2.7 CodeMoonshot AI1273.00+/-5.517,077Moonshot AIOpen Source
38MiniMaxAIMiniMax M3MiniMaxAI1268.00+/-4.922,330MiniMaxAIOpen Source
42DeepSeek-AIDeepSeek-V4-ProDeepSeek-AI1259.00+/-4.526,466DeepSeek-AIOpen Source
44Moonshot AIKimi K2.5 (thinking)Moonshot AI1254.00+/-3.741,601Moonshot AIOpen Source
46MiniMaxAIMiniMax-M2.7MiniMaxAI1252.00+/-3.839,539MiniMaxAIOpen Source
49DeepSeek-AIDeepSeek-V4-FlashDeepSeek-AI1249.00+/-7.29,702DeepSeek-AIOpen Source
58MiniMaxAIMiniMax M2.5MiniMaxAI1225.00+/-6.711,506MiniMaxAIOpen Source
60DeepSeek-AIDeepSeek-V4-FlashDeepSeek-AI1223.00+/-4.428,074DeepSeek-AIOpen Source
62MiniMaxAIM2.1MiniMaxAI1208.00+/-520,805MiniMaxAIOpen Source
67StepFunAIStep 3.7 FlashStepFunAI1198.00+/-4.922,607StepFunAIOpen Source
74DeepSeek-AIDeepSeek-V3.1 (thinking)DeepSeek-AI1193.00+/-5.616,257DeepSeek-AIOpen Source
76DeepSeek-AIDeepSeek V3.2-ExpDeepSeek-AI1188.00+/-5.219,487DeepSeek-AIOpen Source
87DeepSeek-AIDeepSeek V3.2DeepSeek-AI1182.00+/-4.329,120DeepSeek-AIOpen Source
102DeepSeek-AIDeepSeek-R1-0528DeepSeek-AI1156.00+/-5.417,951DeepSeek-AIOpen Source
107MiniMaxAIMiniMax M2MiniMaxAI1153.00+/-6.810,828MiniMaxAIOpen Source
114DeepSeek-AIDeepSeek-V3.1DeepSeek-AI1129.00+/-5.120,334DeepSeek-AIOpen Source
116DeepSeek-AIDeepSeek-V3-0324DeepSeek-AI1126.00+/-5.219,271DeepSeek-AIOpen Source
120Moonshot AIKimi K2 0905Moonshot AI1115.00+/-17.91,504Moonshot AIOpen Source
125Moonshot AIKimi K2 Turbo PreviewMoonshot AI1101.00+/-15.22,094Moonshot AIOpen Source
136Moonshot AIKimi K2Moonshot AI1051.00+/-19.51,352Moonshot AIOpen Source
138AlibabaQwen3-235B-A22B-Thinking-2507Alibaba1050.00+/-9.16,175AlibabaOpen Source

Data is for reference only. Official sources are authoritative. Click model names to view DataLearner model profiles.

About This Leaderboard

This leaderboard uses data from Design Arena developed by Arcada Labs, a Y Combinator-backed platform for anonymous head-to-head evaluation of AI design-code generation.

Unlike LMArena's general text and coding evaluations, Design Arena's code leaderboard focuses on the ability to generate front-end code with visual output. Tasks include Website, UI components, game development, data visualization, SVG, web apps, mobile, and related subcategories.

This page shows the Code Categories aggregate ranking. Votes across subcategories are pooled and scored with a Bradley-Terry model. Votes are counted equally rather than category-weighted, so categories with more votes can influence the aggregate more.

FAQ

01

What is Arcada Labs Code Categories Arena?

Arcada Labs Code Categories Arena is an anonymous evaluation platform focused on AI design-code generation. It covers categories such as websites, UI components, game development, and data visualization, then aggregates votes into an overall ranking.

02

How is Arcada Code Arena different from LMArena Coding Arena?

LMArena Coding Arena focuses on general programming tasks such as code generation, debugging, and algorithms. Arcada Code Arena focuses on visual front-end outputs such as HTML pages, interactive UI components, charts, SVG, and prototypes.

03

What is the ranking methodology?

Arcada Labs pools raw votes from code subcategories and fits a Bradley-Terry model. Votes are equal rather than category-weighted, so higher-volume categories can influence the aggregate more.

04

Which model types perform best for design-code tasks?

Large models with strong visual reasoning and front-end coding ability tend to do well. Specialized UI and code-generation models can also perform strongly when tasks emphasize layout, interaction, and visual polish.