Featured · Frontier
Scale AI cut 89 tasks from its coding benchmark after an audit found gaming, then graded its own fix
Scale AI cut 89 tasks from the public SWE-Bench Pro leaderboard and locked its evaluation runtime after an independent preprint found the benchmark gameable. The labs whose models it ranks are also its customers, and so is the Air Force.
· 3 min read



