Mercor's benchmark flunked 12 real CPAs; the fix it sells is more humans
A study from the $10 billion AI-labor broker shows frontier models acing accounting tasks on which licensed CPAs averaged 37%. The same document shows no model fully solving almost 60% of the full benchmark, and that gap is what Mercor sells.
Vincent Jiang · 3 min read
Twelve CPAs sat the test and averaged 37 percent
Twelve licensed CPAs, averaging five and a half years in practice, worked four month-end close scenarios simplified from Mercor's APEX-Accounting benchmark 12. They met about 37% of the grading criteria 12. Frontier models now ace the same tasks, which the best models missed eighteen months ago 12.
The paper also prices the defeat. Models finished the suites in under ten minutes, at a cost per grading criterion more than an order of magnitude below the humans 211. The study went out on 1 October 2026, and founder-CEO Brendan Foody spent the same week at CBS, ahead of a Sunday 60 Minutes appearance, arguing AI "may change human work rather than simply replace it" 3.
A $20 billion mark rests on the humans being scored
Mercor, founded in 2023, supplies expert contractors to AI labs including OpenAI and Anthropic, roughly 30,000 of them as of October 2025, by the company's count 5. Tracker tallies put its total funding near $485.6 million 7.
The company says it crossed $2 billion in annualized revenue in June, four months after crossing $1 billion 4. In July it was in talks to raise $500 million at a $20 billion valuation, double last fall's $10 billion Series C; Forbes estimated the three founders at $4.3 billion apiece if it closed 46.
The page that flunks accountants also sells the supervision
The paper's quieter finding is the business model. On the full 160-task benchmark, built by more than 40 professionals averaging 11 years in the field, no model fully solved almost 60% of the tasks 1.
Mercor's research hub turns that gap into inventory. APEX data is offered under training license, 50,000-plus tasks across 30-plus domains with samples the same day, alongside 30,000-plus experts and human evaluation "at the scale frontier labs rely on" 89. The tasks are written by professionals from Goldman Sachs, McKinsey and Skadden, the same rung being scored, with the accounting set built with the spend-management firm Ramp 8.
Candid caveats, and a ceiling that sells
Mercor criticizes itself plainly: the tasks "tested what AI is best at," the setting stripped out colleagues and accumulated job context, and task authors expected humans near 30% for juniors and 55% for mid-levels 2. The team says it considered not publishing at all 2. The full-benchmark ceiling still tells labs what to buy next.
Best model clears 61.8% of the full 160-task accounting benchmark
Data
| Value | |
|---|---|
| Claude Opus 5.5 Max | 61.8% |
| Fable 5.1 Max | 61% |
| GPT-6 Astra Max | 57.9% |
| Fable 5 Max | 56.4% |
| Claude Opus 5 Max | 54.5% |
The company's defense is its own catalog
Mercor concludes accountants are not replaceable and expects productivity gains rather than cuts 1. The bench carries the cost.
Former workers describe hostile conditions and week-one resignations without bonus or equity, and an analysis found one in five labellers homeless while many draw public assistance 10. The company calls performance-tied bonuses industry standard 10. Its homepage advertises an average contracted rate of $109 an hour 9.
Two tells, one dated Sunday
Watch whether the $20 billion round Forbes expected to close in July closed, and whether labs keep paying expert rates for a gap the paper itself says persists 46. The LiteLLM supply-chain hack disclosed in April, which reportedly exposed up to four terabytes, and the lawsuits from contract workers ride along into any raise 456.
Either way, the repricing lands first on junior accounting work and finance outsourcing. Foody takes the bigger stage Sunday 3. The machines aced the exam, and the exam's owner still sells the graders.
Deepdive
AI-generated from this story and its cited sources. Not investment advice.


