Datalab built the benchmark that ranks Datalab first, 0.38 points ahead of Reducto
Datalab's new OmniExtractBench crowns Datalab, grades rival Reducto on a suite Reducto itself paid for, and its only coverage is Datalab-sponsored. The buried finding: OpenAI, Microsoft and Mistral are losing the extraction layer to both specialists.
Vincent Jiang · 2 min read
The layer every AI agent's paperwork flows through, document extraction, just got its first shared report card, and the teacher is also the top student. On October 2, Datalab released OmniExtractBench, pooling 620 documents from four vendor-built suites and grading them all with one deterministic, Apache 2.0 scorer that explains every value it judges 1. Datalab's own accurate mode leads at 93.85, its balanced mode sits at 93.48, and rival Reducto's deep_extract v2 is effectively tied at 93.47; 0.38 points separate the tollbooth's two camps 1.
Datalab tops its own extraction benchmark by 0.38 points over Reducto
Data
| Value | |
|---|---|
| Datalab accurate | 93.85% |
| Datalab balanced | 93.48% |
| Reducto deep_extract v2 | 93.47% |
Biased leaderboards were the argument, and Datalab wins the fix
Datalab's launch case is that vendor leaderboards favor the vendor that built them; it then published the first board built to fix that, and it tops the board 1. The announcing write-up on MarkTechPost discloses Datalab sponsorship 1. Reducto, for its part, is being graded partly on LongExtractBench, a 47-document suite it commissioned from micro1 1, after that suite previously scored Datalab at 34% recall before Datalab rebuilt and hit 99.1% 4.
The specialists' rival is losing ground to the frontier's failures
Reducto, closed-source and carrying more than $100M from a16z, Benchmark and First Round, answers with revenue up 8x year over year and customers including Harvey$15.5B — Harvey, private, latest valuation $15.5B, Vanta and Scale 3. Its own site says it has processed over 5 billion pages 5.
Frontier models fail in opposite directions: misses or inventions
- Precision
- Recall
Data
| Precision | Recall | |
|---|---|---|
| GPT 5.6-sol | 95.11% | 84.99% |
| LlamaExtract | 86.57% | 93.13% |
The sponsored post also soft-pedals the finding that should worry buyers most: the specialists are beating the frontier. OpenAI's$1.18T — OpenAI, private, latest valuation $1.18T GPT 5.6-sol posts 95.11 precision but only 84.99 recall, missing 11.88% of values as unfound; LlamaExtract invents, losing 9.03% of its score to fabricated values; Microsoft's+2.38% — Microsoft, up 2.38 percent today Azure Content Understanding and Mistral OCR 4.1 trail on both metrics 1. Reducto co-founder and CEO Adit Abraham frames the stakes plainly: once agents act on extracted data rather than showing it to a person, "the original source document is the only source of truth," and a bad field can propagate through an entire workflow 2.
The ruler is the prize
For enterprises the check is unusually cheap: the scorer installs from PyPI, and rerunning vendors needs only your own API keys 1. Until a customer, not a vendor, runs the yardstick, neither vendor's number moves pricing 1. Whichever benchmark sticks sets the next round's marks, and both camps are racing to make theirs the one that does. Reducto has already used a rival's board to claim the top slot 7; Datalab's index says the same of itself, twice 4.
Deepdive
AI-generated from this story and its cited sources. Not investment advice.



