Coding & Development · 2 Oct 2026 · 17:23 CEST
Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
Publisher preview · OZZZER analysis pending editorial review.
Datalab has released OmniExtractBench, an open benchmark for structured document extraction. It tests how accurately a system fills a JSON schema from a PDF. The benchmark pools 620 documents from 4 existing benchmarks. One deterministic scorer grades all of them and explains each decision. The release lands while extraction vendors publish their own leaderboards. Datalab argues those leaderboards are hard to compare or audit.
OmniExtractBench is its attempt at a shared yardstick. Is it deployable? Yes, the scorer installs from PyPI as omni-extract-bench (v0.1.7, Python 3.11+, SciPy only) under Apache 2.0. Rerunning vendors requires your own API keys and paid credits. What is OmniExtractBench? OmniExtractBench is a structured extraction benchmark built by Datalab. Each task gives a system a PDF and a JSON schema.
The system returns JSON, which is scored value by value against a gold file. The code is on GitHub, and the data is on Hugging Face under CC BY 4.0. The 4 flaws it targets Datalab’s launch post names 4 recurring problems with existing extraction benchmarks: Bias: documents and scoring can favor the vendor that built…
Excerpt supplied by the publisher.
Source
MarkTechPost · 2 Oct 2026 · 17:23 CEST
Open the original at MarkTechPost ↗