Lire en français →

CHECKLISTS

How to evaluate extraction accuracy: a checklist

By Sygnet Research. Written by Sygnet, sourced, checked before publication.

This checklist is for anyone about to sign off on a document AI vendor or model, whether you're running a pilot, renegotiating a contract, or auditing a system already in production. Use it before you trust an accuracy number anyone gives you, including your own team's.

Build the test set first

  • Pull documents from real production sources, not vendor demo samples (vendors curate for success)
  • Include at least 100-200 documents per document type before trusting any aggregate number
  • Cover edge cases deliberately: scanned at an angle, handwritten fields, low-resolution faxes, multi-page variants
  • Hold out a set the vendor or model never saw during tuning
  • Label the ground truth yourself, or have a second person check the labels (vendor-supplied "correct" answers are not independent)
  • Refresh the test set every few months as document formats or suppliers change

Define field-level metrics, not just document-level ones

  • Measure accuracy per field, not just per document (a 95% document accuracy can hide a field that's wrong 40% of the time)
  • Separate "field missing" from "field wrong" from "field hallucinated" (these need different fixes)
  • Track precision and recall separately for fields that are sometimes absent on the source document (totals, tax lines, optional dates)
  • Define what "correct" means for ambiguous fields in advance (date formats, currency, partial names) so your team doesn't improvise mid-review
  • Weight fields by business impact: a wrong IBAN matters more than a wrong invoice description

Measure straight-through processing (STP) rate honestly

  • Define STP as "needed zero human correction," not "model produced an answer"
  • Report STP at the confidence threshold you'll actually use in production, not the one that looks best in a demo
  • Check what happens to the STP rate if you tighten the threshold by 5-10 points (steep drops suggest a fragile model)
  • Separate STP rate by document type: an average across types can mask a type that's failing
  • Track STP rate over time after deployment, not just at the pilot stage

Test under real operating conditions

  • Run the same test on documents from your worst-quality source (old scanner, mobile photo, bad fax)
  • Check performance on your actual document mix, weighted by real volume, not an even split across types
  • Test with the file formats you'll actually receive: PDF, TIFF, JPEG, email attachments
  • Verify the system handles multi-page and multi-document files correctly, not just single-page samples
  • Compare OCR-based and VLM-based approaches on the same set if you're unsure which fits your documents (see OCR vs VLM)

Validate confidence scores

  • Check that low-confidence flags actually correlate with errors (plot confidence against actual accuracy)
  • Decide your review threshold based on cost of error, not a default setting (see AI confidence thresholds for claims processing)
  • Confirm confidence scores are calibrated per field, not just per document
  • Test whether confidence scores stay stable when document quality drops

Put a number on the business impact

  • Calculate cost per correctly extracted field, not just cost per document (see document parsing: accuracy vs. cost per correct answer)
  • Estimate time saved by staff reviewing exceptions versus processing everything manually
  • Model the ROI at your actual volume and error tolerance, not a generic benchmark (see the ROI calculator)
  • Factor in the cost of errors that reach downstream systems undetected

Common mistakes

  • Trusting a vendor's headline accuracy number without checking what document set it was measured on.
  • Measuring document-level accuracy and ignoring which specific fields drive the errors.
  • Testing only on clean samples, then being surprised by real-world scans.
  • Treating STP rate as fixed, when it moves with every threshold and document mix change.
  • Skipping a second pass on your own ground truth labels, so your "baseline" is already wrong.
  • Comparing two systems on different test sets and drawing conclusions anyway.

FAQ

How many documents do I need in a test set to trust the results?

It depends on how many document types and variants you handle, but fewer than 100 per type usually produces noisy, unstable numbers. If you have five document types with different layouts, you need a few hundred per type to catch edge cases. For a detailed walkthrough of testing invoice extraction specifically, see invoice AI accuracy: how to test and measure real performance.

Should I compare vendors on their own test sets or mine?

Always insist on your own. Vendors naturally report numbers from sets that favor their system, and even honest vendors tune on data that doesn't match your document mix. If you're comparing multiple vendors, see our IDP vendor evaluation checklist for the full process beyond accuracy alone.

Does a high STP rate always mean low error rate?

No. A high STP rate can hide a system that's confidently wrong on fields nobody checks, like a secondary tax code or a rarely-used date field. Always pair STP rate with field-level accuracy and check what fraction of "straight-through" documents actually had an error that simply wasn't flagged for review.

NEXT STEP

See it on your own documents

One email when we publish something worth your time.