Document Parsing: Accuracy vs. Cost Per Correct Answer
By Sygnet Research, checked before publication

Key takeaways
- Benchmark accuracy is measured per page or per field on clean, curated corpora; your cost is paid per correct, usable record, which is a different denominator entirely.
- A parser that is 90% accurate but overconfident leaks errors into production, while a 78% parser with calibrated confidence routes its own failures to review, so the cheap-looking option is often the less accurate one.
- The accounts payable industry average touchless rate is 32.6% and best-in-class is 49.2%, per Ardent Partners figures cited by Parseur, nowhere near the 90%+ numbers model cards imply.
- Rerun the arithmetic with your own review labor rate, your own error cost, and your own document mix before you pick on leaderboard rank.
Why does a higher benchmark score not mean a lower cost per correct answer?
Because the benchmark number and the cost driver measure different things. Leaderboard accuracy is an average over a fixed corpus; cost per correct answer is the total spend (inference plus human review plus the cost of errors that escape) divided by the number of records you can actually use downstream. The link between the two runs through two variables benchmarks ignore: how the errors are distributed across fields, and whether the system knows when it is wrong.
OmniDocBench, one of the most widely used parsing benchmarks, evaluates parsing accuracy on 1,355 pages using edit-distance and table-structure metrics, and its overall score is a weighted blend: Overall = ((1 − TextEdit) + FormulaCDM + TableTEDS) / 3 × 100. That is a useful research metric and a terrible procurement metric. A model can gain two overall points by transcribing display formulas better while getting worse at the four fields your AP workflow actually needs. The authors of OmniDocBench themselves note that it is restricted to two languages and largely to clean or semi-clean scans.
Benchmark accuracy is an average over someone else's documents, weighted by someone else's priorities.
What is the real cost formula?
Cost per correct answer = (inference cost + review labor + leaked-error cost) ÷ correct records. Write it out and the ranking flips surprisingly often.
Review labor usually dominates. In most financial institutions, manual verification of extracted fields still dominates the per-document cost, which makes straight-through processing the key driver of productivity. Inference, by comparison, is cheap and getting cheaper: Google Document AI charges $0.10 per document of up to 10 pages for the invoice parser, and $30 per 1,000 pages for the form parser and custom extractor, while Amazon Textract prices forms at $0.05 per page and lending at $0.07. A reviewer at €35 fully loaded costs roughly €0.29 per 30 seconds of attention. One touched document wipes out several pages of model inference.
Leaked errors are the third term and the hardest to price. In high-stakes pipelines, including financial reconciliation, compliance verification, and procurement automation, an extraction that is silently wrong is more dangerous than one that is visibly absent. If a wrong IBAN costs you €400 in recovery work, a 0.5% leak rate on 50,000 documents a year is €100,000 that never appears on any vendor invoice.
How does calibration beat raw accuracy?
Calibration converts accuracy into automation. A parser that reliably flags its own 22% of uncertain fields lets you auto-approve the other 78% with a bounded error rate; a parser that is right 90% of the time but confident on all of it forces you to review everything or accept 10% leakage.
The formal version of this is a selective gate: accept a field only if its probability exceeds a threshold chosen on a held-out calibration set so that the error rate among accepted fields stays at or below an operator-chosen budget, using Learn-Then-Test for a distribution-free guarantee, with everything below the threshold routed to human review. The relevant metric is not accuracy, it is the percentage of fields that can be auto-approved while holding accepted-tier error at or below a target.
Which is why raw VLM confidence is a trap. Modern VLMs extract key-values out of the box, but their verbalized confidence signals are unreliable and weakly track field correctness. A survey of calibration on document extraction found quality ranging from near-perfect to severely overconfident, so measure against your own documents before you trust a score. The test is cheap: sample 100 extractions the system marked 0.90+ and verify them manually; if more than 10% contain errors, the confidence scores are miscalibrated.
A worked comparison
The table below is a model, not measured data: 10,000 documents a month, 12 fields each, €0.29 of reviewer time per touched document, €400 per leaked error. Parser A scores 90% field accuracy with overconfident scores; Parser B scores 78% with calibrated per-field confidence. Plug your own numbers in, the structure is what matters.
| Input | Parser A (90%, overconfident) | Parser B (78%, calibrated) |
|---|---|---|
| Field accuracy | 90% | 78% |
| Usable confidence signal | No | Yes, per field |
| Documents auto-approved | 35% (hand-set rules) | 72% (threshold at 0.5% accepted-tier error) |
| Documents touched / month | 6,500 | 2,800 |
| Review cost / month | €1,885 | €812 |
| Leaked errors / month | ~45 | ~8 |
| Leaked-error cost | €18,000 | €3,200 |
| Inference cost @ €0.05/doc | €500 | €500 |
| Total / month | €20,385 | €4,512 |
| Cost per correct record | €2.08 | €0.45 |
The twelve-point accuracy gap is worth less than the calibration gap, by a factor of four in this model. Note also what the comparison does not include: integration effort, schema drift, and the fact that Amazon's A2I review service is closed to new customers and Google deprecated Document AI's human-in-the-loop, so on those APIs your engineers build the review screens themselves. That is a real line item, and it is the core of any honest build versus buy analysis.
What should engineering teams measure instead?
Measure straight-through rate at a fixed accepted-tier error budget, then cost per correct record. Everything else is diagnostic.
Three lanes, not two, is the operational pattern that makes this measurable: straight-through, field-level review, and full manual exception; collapse them into one "needs review" pile and the pile wins. Set per-field thresholds by consequence, not uniformly: price thresholds by what a wrong value costs, with AWS putting the sane range at roughly 50% for archival text and 90% or higher for financial decisions. For reference targets, 60% to 75% straight-through after three months of tuning is a realistic ambition, with leaked error rate below 0.5% and field-level accuracy above 95% on critical fields.
Then audit. A random sample of auto-approved documents is the only check that finds a threshold set too low before somebody downstream does. Sygnet builds document extraction with per-field confidence and review routing, and publishes its own notes on setting confidence thresholds for claims workflows.
The question is never how accurate the model is, but how many fields it will let you auto-approve at an error rate you can defend to an auditor.
When does the 90% parser actually win?
When your documents resemble the benchmark and your error cost is low. If you are building a RAG corpus over clean digital PDFs, raw parsing quality is close to the whole story: there is no review queue, and a 1% text edit distance improvement propagates straight into retrieval quality.
The picture changes with messy inputs. HunyuanOCR's team built a "Wild" version of OmniDocBench by printing documents and re-capturing them under folding, bending and varying illumination precisely because clean-scan scores do not predict phone-photo performance, and most KYC or claims intake traffic is phone photos. Table-heavy work is another split: on OmniDocBench v1.5, reported TEDS figures range from 57.88 for Marker-1.8.2 to 86.78 for dots.ocr, a 29-point spread that an overall score flattens into something much less alarming. If your value lives in tables, read the sub-metric, not the headline, and see our notes on table extraction metrics.
FAQ
How do I calculate cost per correct answer for a document parser?
Run 300 to 500 of your own documents through each candidate. Count fully correct records (all required fields right), measure average reviewer seconds per touched document with a stopwatch, and assign a euro value to a leaked error based on what the downstream fix actually costs. Then: (inference + review labor + leak cost) ÷ correct records. Do this per document type; the ranking often differs between invoices and ID documents.
Is field-level accuracy or document-level accuracy the right metric?
Document-level, if you care about automation. Document accuracy asks what percentage of documents had zero extraction errors across all fields, and it is the number that determines your touchless processing rate. Field accuracy of 98% across 12 fields implies roughly 78% of documents containing at least one error if failures are independent. They are not independent in practice, which is why you measure rather than compute.
Can I fix an overconfident parser without changing models?
Often yes, with a calibration layer. One study reports that a Lasso calibration layer fitted on about 165 in-domain samples appears sufficient for a new deployment domain, and that negative sample quality dominates dataset size, with a 26% natural failure rate providing richer supervision than a 4% one. Practically: collect your reviewer corrections, fit a small model mapping signals to correctness, and gate on that instead of the vendor score.
What straight-through rate should I expect in year one?
Lower than the vendor deck says. Industry benchmarks for accounts payable put the average touchless rate at 32.6% and best-in-class at 49.2%. A well-tuned modern pipeline on a narrow document type can beat that, but plan budgets around 50% to 70% on mixed traffic, and model the review queue as a permanent cost line rather than a transitional one.
One email when we publish something worth your time.