Artificial Analysis Coding Agent Index v1.5 is an equal-weight composite of DeepSWE v1.1 (113 tasks), Terminal-Bench 4.0 (66 tasks), and SWE-Atlas-QnA (124 tasks), covering 303 tasks in total. Each component averages pass@1 over three attempts per task before the three component scores are equally weighted.
Browse the latest scores, model modes, release dates, and parameter sizes for AA Coding Agent Index v1.5.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
No benchmark data available yet