Invoice AI Accuracy: How to Test & Measure Real Performance
By Sygnet Research, checked before publication

Key takeaways
- An AI tool that advertises "98% accuracy" tells you nothing useful unless you know which fields, which type of invoices, and which measurement method produced that number.
- The right verification method fits in one sentence: assemble a batch of 150 to 300 real invoices, enter them manually once, and compare field by field against the tool's output.
- IOFM (Institute of Finance & Management) benchmarks put the error rate of automated accounts payable departments at around 0.8%, versus roughly 2% for manual processing: aiming for zero errors costs more than the errors themselves.
- From 1 September 2026, all VAT-registered businesses in France must be able to receive electronic invoices, which changes the nature of the verification work but does not eliminate it.
Why should you test an AI tool on your own invoices before adopting it?
Because the accuracy figures published by vendors are measured on their own datasets, not yours. A tool that excels on Orange or EDF invoices can fall apart on a delivery note scanned at an angle by a tradesman, or on a credit note with three different VAT rates.
Take an accounting firm of 40 people that receives 300 supplier invoices a month for its clients. Converging studies put the cost of manually processing a single invoice at between €12.80 and €20, averaging around €15 for a French SME. Across 3,600 invoices a year, that represents roughly €54,000. But an undetected error also has a cost: the same research estimates the cost of correcting a misallocated invoice at around €150, factoring in reprocessing, re-validation and chasing the supplier.
In other words, a 2% rate of silent errors across 3,600 invoices wipes out a good share of the gain. That is exactly why the testing phase is not a box-ticking exercise with the vendor.
An accuracy rate without a measurement protocol is not data, it's a marketing claim.
How do you build a representative test set?
Take 150 to 300 real invoices from your last three months, and make a point of including your most troublesome cases. A sample that is too clean will give you a false sense of security.
The actual composition for our 40-person accounting firm:
- 60% "normal" invoices (native PDF, single VAT rate, recurring supplier);
- 20% poor-quality scans, phone photos, skewed documents;
- 10% multi-page invoices with line-item tables, the most common breaking point (see Sygnet's guide on evaluating table extraction);
- 10% awkward cases: credit notes, foreign-currency invoices, VAT reverse charge, deposits, invoices with no purchase order.
Next, build your ground truth: have one person manually enter the 8 to 12 fields that actually matter to you (invoice number, supplier SIRET business registration number, date, due date, net total, VAT by rate, gross total, IBAN, order reference). Budget 3 to 4 hours for 200 invoices. This is the only serious investment in the whole protocol, and it can be reused to test a second vendor later.
Which metrics should you look at (and which should you ignore)?
The only metric that matters is the percentage of invoices that are usable without human correction, not average character-level accuracy. An invoice with 11 out of 12 correct fields still has to be reopened by a human, so it counts as a failure.
Measure three things:
- Field-level accuracy: for each field, the percentage of values matching your manual entry. You'll quickly find that the due date and the VAT breakdown by rate are the weak points.
- Straight-through processing rate: invoices where every critical field is correct. This is your real time saving.
- Silent errors: cases where the tool is wrong while reporting high confidence. These are the only dangerous errors, since they slip past the checks.
Everything else is noise. A tool claiming 99% character-level accuracy can still produce 20% of invoices that need reopening. To dig into the trade-off between accuracy and real cost, Sygnet has documented the reasoning in terms of cost per correct answer rather than raw percentages.
How should you use confidence scores without being misled?
A confidence score is meant for sorting, not reassurance. Its only legitimate use is to automatically decide which invoices go to human review and which can move straight into the accounting system.
The method is simple. On your test set, plot the actual observed error rate for each confidence tier. If a confidence of 0.95 corresponds to a 0.5% error rate and 0.85 corresponds to 6%, your automatic-approval threshold sits somewhere in between, depending on your risk tolerance. This calibration is specific to each business: a threshold that works for €40 expense claims makes no sense for €80,000 subcontracting invoices.
Common-sense rule: index the threshold to the amount. Any invoice above a given threshold (say, €5,000) goes to human review regardless of the confidence score shown. And any change to a supplier's IBAN goes to human review, full stop. That is the main entry point for payment fraud, and no extraction tool can protect you from a perfectly legible fake bank details document (RIB).
| What you measure | What it tells you | Common trap |
|---|---|---|
| Character-level accuracy | Raw reading quality | Flattering, with little bearing on work actually saved |
| Field-level accuracy | Where the tool breaks down | Hides invoices with multiple errors |
| Rate of perfect invoices | Real time saved | Drops quickly on line-item detail |
| Silent errors | Financial risk | Ignored in most vendor demos |
| Cost per correctly processed invoice | Profitability | Forgets the cost of human review |
What automated checks should you put in place in production?
Arithmetic and business-rule checks catch most extraction errors, with no AI and no human review needed. They are easy to write and should run on 100% of invoices.
The minimum checklist, for our 40-person accounting firm:
- net total + VAT = gross total (1-cent rounding tolerance);
- sum of line items = net total;
- VAT rate belongs to the list of rates currently in force in France (20%, 10%, 5.5%, 2.1%);
- valid SIRET number (Luhn checksum) consistent with the supplier database;
- duplicate detection on the triplet supplier + invoice number + amount, the best safeguard against double payment;
- due date later than the invoice date.
Any invoice that fails one of these checks goes into the human review queue, regardless of its confidence score. And keep a full audit trail: who validated what, when, and what the original value was. You will need this in the event of a tax audit, and invoices must in any case be kept for 10 years under French commercial law (Code de commerce), a rule confirmed in e-invoicing compliance guides.
Does e-invoicing make these checks unnecessary?
No, it simply shifts part of the burden. From 1 September 2026, all VAT-registered businesses in France will need to be able to receive structured electronic invoices, with no revenue threshold; the obligation to issue such invoices will apply first to large companies and mid-sized firms (ETI), then from 1 September 2027 to SMEs, small businesses and micro-enterprises.
An invoice in a structured format (Factur-X, a French hybrid PDF/XML e-invoicing standard, or UBL or CII) no longer needs to be "read": the data arrives already structured. But three realities remain. First, for months you will still receive PDFs and scans from foreign suppliers or those lagging in compliance. Second, attachments in the purchasing cycle, such as delivery notes, quotes and contracts, are not covered by the reform. Finally, structured data does not mean correct data: a supplier can perfectly well issue a wrong IBAN within an entirely compliant data flow.
On the penalty side, failure to issue invoices electronically is punishable by a fine of €15 per invoice, capped at €15,000 per calendar year, while failure to e-report is punishable by €250 per missing transmission. Sygnet has published a detailed overview of French e-invoicing and of the applicable penalties.
The reform guarantees the format of the invoices you receive, never their accuracy.
How long does it take to properly validate a tool?
Four to six weeks, including two weeks of parallel running. The trap is switching to production after a 20-invoice demo hand-picked by the vendor.
The approach that works:
- Week 1: build the test set and manually enter the ground truth.
- Week 2: run the batch through the tool, calculate the three metrics, calibrate the thresholds.
- Weeks 3 to 5: dual entry in production. The tool processes everything, humans process everything, and you compare the two every day. It's costly, but it's the only phase where you discover the cases your sample missed.
- Week 6: make the decision, backed by one figure: cost per correctly processed invoice, including human review.
A useful benchmark: top-performing accounts payable teams spend $2.78 per invoice processed, against $12.88 for average organisations. If your fully-loaded cost after automation does not drop clearly below your current manual cost, the problem isn't the AI, it's how human review has been configured.
FAQ
What accuracy rate should you demand from an invoice extraction tool?
Demand a straight-through processing rate of at least 85% on your own documents, and a residual error rate below 1% on financial fields. Industry benchmarks put automation at around 0.8% errors versus roughly 2% for manual processing. Don't accept any figure that hasn't been measured on your own test set.
Do invoices still need human validation?
Yes, but in a targeted way. Three categories should always go to review: invoices above an amount threshold you define, invoices that fail an arithmetic check, and any change to a supplier's bank details. Everything else can be processed automatically once your confidence thresholds have been calibrated on real data.
Should you build your own tool or buy a solution?
Below roughly 1,000 invoices a month, buying is almost always more cost-effective: the hidden cost of an in-house build lies in maintaining edge cases and keeping up with regulatory format changes. A detailed comparison of both approaches is available in Sygnet's report on the build versus buy trade-off.
One email when we publish something worth your time.