Datalab built the benchmark that ranks Datalab first, 0.38 points ahead of Reducto

Datalab's new OmniExtractBench crowns Datalab, grades rival Reducto on a suite Reducto itself paid for, and its only coverage is Datalab-sponsored. The buried finding: OpenAI, Microsoft and Mistral are losing the extraction layer to both specialists.

Vincent JiangVincent Jiang · 2 min read
Share
Adit Abraham, Reducto's co-founder and chief executive, speaking into a microphone in a recorded conversation
1 / 7Slide 1 of 7
Adit Abraham, Reducto co-founder and CEO, whose deep_extract v2 sits 0.38 points behind Datalab's accurate mode on OmniExtractBench.

The layer every AI agent's paperwork flows through, document extraction, just got its first shared report card, and the teacher is also the top student. On October 2, Datalab released OmniExtractBench, pooling 620 documents from four vendor-built suites and grading them all with one deterministic, Apache 2.0 scorer that explains every value it judges 1. Datalab's own accurate mode leads at 93.85, its balanced mode sits at 93.48, and rival Reducto's deep_extract v2 is effectively tied at 93.47; 0.38 points separate the tollbooth's two camps 1.

Datalab tops its own extraction benchmark by 0.38 points over Reducto

0%20%40%60%80%100%Datalab accurate93.85%0.38 points cover all threeDatalab balanced93.48%Reducto deep_extract v293.47%Graded partly on a suite Reducto commissioned
Data
Value
Datalab accurate93.85%
Datalab balanced93.48%
Reducto deep_extract v293.47%
Accuracy, percent of values matched, OmniExtractBench, 620-document pool, as published by Datalab, October 2, 2026.1

Biased leaderboards were the argument, and Datalab wins the fix

Datalab's launch case is that vendor leaderboards favor the vendor that built them; it then published the first board built to fix that, and it tops the board 1. The announcing write-up on MarkTechPost discloses Datalab sponsorship 1. Reducto, for its part, is being graded partly on LongExtractBench, a 47-document suite it commissioned from micro1 1, after that suite previously scored Datalab at 34% recall before Datalab rebuilt and hit 99.1% 4.

The specialists' rival is losing ground to the frontier's failures

Reducto, closed-source and carrying more than $100M from a16z, Benchmark and First Round, answers with revenue up 8x year over year and customers including Harvey$15.5B — Harvey, private, latest valuation $15.5B, Vanta and Scale 3. Its own site says it has processed over 5 billion pages 5.

Frontier models fail in opposite directions: misses or inventions

  • Precision
  • Recall
0%50%100%GPT 5.6-solLlamaExtract95.11%93.13%
Data
PrecisionRecall
GPT 5.6-sol95.11%84.99%
LlamaExtract86.57%93.13%
Precision and recall, percent, OmniExtractBench 620-document pool, as published by Datalab, October 2, 2026. GPT 5.6-sol misses fields; LlamaExtract fabricates them.1

The sponsored post also soft-pedals the finding that should worry buyers most: the specialists are beating the frontier. OpenAI's$1.18T — OpenAI, private, latest valuation $1.18T GPT 5.6-sol posts 95.11 precision but only 84.99 recall, missing 11.88% of values as unfound; LlamaExtract invents, losing 9.03% of its score to fabricated values; Microsoft's+2.38% — Microsoft, up 2.38 percent today Azure Content Understanding and Mistral OCR 4.1 trail on both metrics 1. Reducto co-founder and CEO Adit Abraham frames the stakes plainly: once agents act on extracted data rather than showing it to a person, "the original source document is the only source of truth," and a bad field can propagate through an entire workflow 2.

The ruler is the prize

For enterprises the check is unusually cheap: the scorer installs from PyPI, and rerunning vendors needs only your own API keys 1. Until a customer, not a vendor, runs the yardstick, neither vendor's number moves pricing 1. Whichever benchmark sticks sets the next round's marks, and both camps are racing to make theirs the one that does. Reducto has already used a rival's board to claim the top slot 7; Datalab's index says the same of itself, twice 4.

Deepdive

AI-generated from this story and its cited sources. Not investment advice.

Reader comments

0 comments

    Sign up

    Get your curated digest

    After email confirmation, you will receive a daily digest of the most relevant news that matter to your portfolio