---
title: "Datalab built the benchmark that ranks Datalab first, 0.38 points ahead of Reducto"
description: "Datalab's new OmniExtractBench crowns Datalab, grades rival Reducto on a suite Reducto itself paid for, and its only coverage is Datalab-sponsored. The buried finding: OpenAI, Microsoft and Mistral are losing the extraction layer to both specialists."
publisher: "The Inference"
section: "Platforms"
published: 2026-10-09T15:07:18.707Z
modified: 2026-10-09T15:07:18.707Z
canonical: https://theinference.org/article/datalab-built-the-benchmark-that-ranks-datalab-first-0-38-points-ahead-of-reducto
language: en
keywords: "Benchmarks, Document AI, Enterprise AI, Startups"
---

# Datalab built the benchmark that ranks Datalab first, 0.38 points ahead of Reducto

> Datalab's new OmniExtractBench crowns Datalab, grades rival Reducto on a suite Reducto itself paid for, and its only coverage is Datalab-sponsored. The buried finding: OpenAI, Microsoft and Mistral are losing the extraction layer to both specialists.

The layer every AI agent's paperwork flows through, document extraction, just got its first shared report card, and the teacher is also the top student. On October 2, Datalab [released OmniExtractBench](https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/), pooling 620 documents from four vendor-built suites and grading them all with one deterministic, Apache 2.0 scorer that explains every value it judges [1]. Datalab's own accurate mode leads at 93.85, its balanced mode sits at 93.48, and rival [Reducto's deep_extract v2](https://theinference.org/markets/companies/reducto) is effectively tied at 93.47; 0.38 points separate the tollbooth's two camps [1].

*Chart: **Datalab tops its own extraction benchmark by 0.38 points over Reducto** Bars showing Datalab accurate at 93.85, Datalab balanced at 93.48 and Reducto deep_extract v2 at 93.47 accuracy on OmniExtractBench.*

*Accuracy, percent of values matched, OmniExtractBench, 620-document pool, as published by Datalab, October 2, 2026. [1]*

## Biased leaderboards were the argument, and Datalab wins the fix

Datalab's launch case is that vendor leaderboards favor the vendor that built them; it then published the first board built to fix that, and it tops the board [1]. The announcing write-up on MarkTechPost discloses Datalab sponsorship [1]. [Reducto](https://reducto.ai/about), for its part, is being graded partly on LongExtractBench, a 47-document suite it commissioned from micro1 [1], after that suite previously [scored Datalab at 34% recall before Datalab rebuilt and hit 99.1%](https://www.datalab.to/blog) [4].

## The specialists' rival is losing ground to the frontier's failures

Reducto, closed-source and carrying more than $100M from [a16z](https://theinference.org/markets/companies/a16z), Benchmark and First Round, [answers with revenue up 8x year over year](https://jobs.a16z.com/jobs/reducto/c670e390-451c-4ec0-8379-ae50d268f9be--founding-data-engineer) and customers including [Harvey](https://theinference.org/markets/companies/harvey), Vanta and Scale [3]. Its own site [says it has processed over 5 billion pages](https://reducto.ai/about) [5].

*Chart: **Frontier models fail in opposite directions: misses or inventions** Column chart comparing precision and recall for GPT 5.6-sol and LlamaExtract, showing frontier models trading misses for inventions.*

*Precision and recall, percent, OmniExtractBench 620-document pool, as published by Datalab, October 2, 2026. GPT 5.6-sol misses fields; LlamaExtract fabricates them. [1]*

The sponsored post also soft-pedals the finding that should worry buyers most: the specialists are beating the frontier. [OpenAI](https://theinference.org/markets/companies/openai)'s GPT 5.6-sol posts 95.11 precision but only 84.99 recall, missing 11.88% of values as unfound; LlamaExtract invents, losing 9.03% of its score to fabricated values; [Microsoft](https://theinference.org/markets/companies/microsoft)'s Azure Content Understanding and Mistral OCR 4.1 trail on both metrics [1]. Reducto co-founder and CEO Adit Abraham [frames the stakes plainly](https://www.computerweekly.com/blog/CW-Developer-Network/Reducto-targets-the-document-bottleneck-in-enterprise-AI): once agents act on extracted data rather than showing it to a person, "the original source document is the only source of truth," and a bad field can propagate through an entire workflow [2].

## The ruler is the prize

For enterprises the check is unusually cheap: the scorer installs from PyPI, and rerunning vendors needs only your own API keys [1]. Until a customer, not a vendor, runs the yardstick, neither vendor's number moves pricing [1]. Whichever benchmark sticks sets the next round's marks, and both camps are racing to make theirs the one that does. Reducto has already [used a rival's board to claim the top slot](https://www.morningstar.com/news/pr-newswire/20260701sf96113/reducto-deep-extract-ranks-first-overall-in-longextractbench-an-independent-benchmark-for-complex-document-extraction) [7]; Datalab's index says the same of itself, twice [4].

## Takeaway

OmniExtractBench crowns its own builder, and the sponsored write-up buries the real finding: specialists are beating frontier models at extraction, and whichever vendor's yardstick sticks sets the next round's marks.

## Sources

1. [MarkTechPost, Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks, 2 October 2026](https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/)
2. [Computer Weekly Developer Network, Reducto targets the document bottleneck in enterprise AI, October 2026](https://www.computerweekly.com/blog/CW-Developer-Network/Reducto-targets-the-document-bottleneck-in-enterprise-AI)
3. [Andreessen Horowitz jobs board, Founding Data Engineer at Reducto, 5 October 2026](https://jobs.a16z.com/jobs/reducto/c670e390-451c-4ec0-8379-ae50d268f9be--founding-data-engineer)
4. [Datalab blog index, posts 049 and 063, accessed 9 October 2026](https://www.datalab.to/blog)
5. [Reducto, About Reducto, accessed 9 October 2026](https://reducto.ai/about)
7. [PR Newswire via Morningstar, Reducto Deep Extract Ranks First Overall in LongExtractBench, 1 July 2026](https://www.morningstar.com/news/pr-newswire/20260701sf96113/reducto-deep-extract-ranks-first-overall-in-longextractbench-an-independent-benchmark-for-complex-document-extraction)


---

Datalab built the benchmark that ranks Datalab first, 0.38 points ahead of Reducto — The Inference. Canonical: https://theinference.org/article/datalab-built-the-benchmark-that-ranks-datalab-first-0-38-points-ahead-of-reducto
