SWE-Atlas-QnA is Scale AI's repository-comprehension benchmark for coding agents. It contains 124 tasks from 11 production open-source repositories across Go, Python, C, and TypeScript. Agents must inspect and run code, trace behavior across files, and answer deep technical questions; Task Resolve Rate is binary and requires all rubric criteria to pass.
Browse the latest scores, model modes, release dates, and parameter sizes for SWE-Atlas-QnA.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
No benchmark data available yet