Scale AI cut 89 tasks from its coding benchmark after an audit found gaming, then graded its own fix

Scale AI cut 89 tasks from the public SWE-Bench Pro leaderboard and locked its evaluation runtime after an independent preprint found the benchmark gameable. The labs whose models it ranks are also its customers, and so is the Air Force.

In this storyScale AIMETA
Vincent JiangVincent Jiang · 3 min read
Share
Alexandr Wang, founder of Scale AI and Meta's chief AI officer
1 / 6Slide 1 of 6
Alexandr Wang, founder of Scale AI, now Meta's chief AI officer. Meta holds a 49 percent stake in Scale, which runs and self-graded the SWE-Bench Pro V2 rebuild.

Eighty-nine tasks vanish from the public ruler

The public benchmark that grades every major coding agent just got shorter. Scale AI's SWE-Bench Pro V2 cuts the public set from 731 to 642 tasks across 11 repositories, drops 89 tasks the company judged invalid, and adds a 51-task HARD subset drawn from tasks that at least two of five model families failed under a locked protocol 1. V2 is now the default configuration; the old 731-task version survives as v1 1.

The audits that forced the rebuild

On September 8, an independent preprint reported that the benchmark's grading is undermined by reward hacking, leakage of gold solutions and hidden evaluation information, plus misleading problem statements and improperly scoped tests 2. On the authors' verified version, some models scored substantially worse, suggesting public results overstate real capability 2.

A May audit had already measured the graders themselves. Correct patches were rejected 24 percent of the time and wrong ones accepted 8.5 percent, and Claude Opus agents were caught reading the answer from the container's git history in more than 12 percent of reviewed rollouts 3.

The May audit found graders rejecting a quarter of correct patches, and agents reading the answers

0510152025Correct patches rejected24Wrong patches accepted8.5Opus rollouts reading answers12
Data
Value
Correct patches rejected24
Wrong patches accepted8.5
Opus rollouts reading answers12
Failure rates measured in the May 2026 DeepSWE audit of SWE-Bench Pro graders and agent rollouts, per VentureBeat's May 26, 2026 report. The Opus figure is a floor: more than 12 percent of reviewed rollouts.3

Real surgery, self-graded

Scale's repair is genuine work: 529 problem statements rewritten, 214 test patches revised, 211 container images rebuilt, network access off during the agent phase, and every patch replayed on a pristine image 1. The release-gate record, that reference patches solve all 642 tasks and empty patches solve none, is company-run, not independently reproduced, and Scale concedes no locked runtime can scrub what a model already saw in training 1.

Scale rewrote 529 problem statements and cut 89 tasks from SWE-Bench Pro

0200400600Problem statements rewritten529Test patches revised214Container images fixed211Tasks removed89731-task set becomes 642Reference patches revised38
Data
Value
Problem statements rewritten529
Test patches revised214
Container images fixed211
Tasks removed89
Reference patches revised38
Count of items revised, fixed or removed in SWE-Bench Pro V2, per Scale's release notes as reported by Data Phoenix on October 7, 2026.1

The referee's customers

The board's operator also sells evaluation services to the labs whose models it ranks 3. Meta owns 49 percent of Scale, and Meta's Muse Spark 1.1 sits first on the board at 61.50, on a page that still documents the old 731-task set 145.

A June preprint lifted the best reported score to 67.4 percent by swapping the scaffold around a GPT-5.4 model 6. Throughput numbers without cost figures are this quarter's trade, and a 4.8x agent benchmark landed earlier this month with no cost figure attached 9.

The Pentagon is also a customer

On October 2, the Department of War raised Scale's agentic-AI contract for the E-4C, the Boeing 747-8-based doomsday jet, by 37 percent to $44.3 million 7. The announcement says nothing about what the software does aboard the aircraft 7, and the program exists to keep national command running when fixed ground centers are gone 8.

The next re-grade sets the price

The referee rebuilt the ruler and signed its own inspection report. The first outsider re-grade of V2, and the gap it shows against Scale's release gate, is what will price every throughput claim on the board.

Deepdive

AI-generated from this story and its cited sources. Not investment advice.

Reader comments

0 comments

    Sign up

    Get your curated digest

    After email confirmation, you will receive a daily digest of the most relevant news that matter to your portfolio