INDUSTRY USE CASES

Benchmarking LLM Contract Review: Accuracy & Efficiency

By Sygnet Research, checked before publication

Key takeaways

  • You never need 5,000 labeled contracts: a stratified sample of roughly 400 documents gives you a ±5 point accuracy estimate at 95% confidence, and per-clause strata matter more than total volume.
  • Public expert-labeled corpora do the heavy lifting for free: CUAD ships over 500 contracts labeled by legal experts across 41 clause types, totaling more than 13,000 annotations, and MAUD adds 47,457 annotations over 152 public merger agreements.
  • An LLM judge is usable only after you measure its agreement with human reviewers; published agreement rates cluster near 0.83 and collapse when the reference standard shifts from "one expert" to "majority of experts".
  • Report recall per clause type, not a single global F1. On CUAD-derived benchmarks the best models sit around F1 0.64, and the variance across clause categories is where diligence risk actually lives.

Why can't you just hand-label the data room?

Because expert clause annotation costs more than the deal review itself. MAUD's authors put real numbers on this: annotating it took over 10,000 hours of work by law students and experienced lawyers, with each law student completing 70-100 hours of training before labeling, and the team estimated the dataset's value at over $5 million using a prevailing rate of $500 per hour in M&A legal fees. That was 152 agreements. Scale that labeling model to 5,000 contracts with 20 to 40 target clauses each and you are funding a second diligence team.

The second reason is structural. Clause extraction is a needle-in-haystack problem: in CUAD, only about 0.25% of each contract is highlighted for any given label on average, and more than 99% of sliding-window features contain none of the 41 relevant labels, so models trained naively learn to always output the empty span. Random sampling of pages, or of documents, buys you almost no signal on the clauses you care about. You need sampling designed around positives.

Expert clause annotation at data-room scale costs more than the diligence review it is supposed to accelerate.

What does a benchmark architecture look like at 5,000 contracts?

Four layers, each answering a different question, with only the smallest layer requiring human labels.

Layer 1: public gold sets to validate the method before touching client data. CUAD's test split alone gives you 4,128 data points across 41 clause categories from 102 unique contracts, with context length from 0.6k to 301k characters and a 30% positive / 70% negative label distribution. That negative mass is the point: it tests whether your pipeline hallucinates clauses that aren't there.

Layer 2: a stratified human-labeled slice of your own corpus, 300 to 500 documents, allocated by contract type, language, scan quality and page count rather than drawn uniformly.

Layer 3: multi-model consensus pseudo-labels across the remaining documents. Where two heterogeneous extractors agree on a span and a normalized value, treat it as a provisional label; where they disagree, you have a free, high-yield queue for human review.

Layer 4: deterministic invariants that need no labels at all: date parse validity, governing-law values inside a closed vocabulary, notice periods within plausible ranges, party names reconciled against the cap table. Invariant violations are errors by construction.

LayerVolumeHuman costWhat it measures
Public gold (CUAD, MAUD)100-500 contractszeroMethod validity, negative-class discipline
Stratified human slice300-500 docs2-4 expert weeksUnbiased precision/recall with CIs
Consensus pseudo-labelsall 5,000review of disagreements onlyDrift, coverage, document-type blind spots
Deterministic invariantsall 5,000zeroFormat, range and cross-document contradictions

How big does the human-labeled sample actually need to be?

For a global accuracy figure, around 385 documents. That is plain binomial arithmetic: 1.96² × 0.25 / 0.05² ≈ 384 samples for a ±5 point margin at 95% confidence, independent of whether the corpus holds 5,000 or 500,000 contracts. Tighten to ±2 points and the number jumps to about 2,400, which is why ±5 is the sane operating choice for a one-off diligence benchmark.

The trap is that a global figure is nearly useless. You need per-clause recall, and rare clauses are rare: a change-of-control provision might appear in 8% of the corpus, so a 400-document random sample yields roughly 32 positives and a confidence interval wide enough to drive a truck through. Fix it by allocating per-stratum: 30 to 50 positives per clause type you actually care about, found by over-sampling documents the pipeline flagged plus keyword-retrieved candidates it did not. Then reweight to corpus proportions when you report. Our checklist for evaluating extraction accuracy covers the stratum design in more detail.

Can an LLM judge replace the human annotator?

It can multiply your annotators, not replace them, and only after you measure its agreement on the stratified slice. The published numbers are encouraging but not blank-cheque. A scoping review of LLM-as-a-judge in healthcare found that across 13 studies reporting agreement rate, both the median and mean sat at 0.83, with the caveat that well-scoped tasks with anchored rubrics outperform open-ended generation. In statute-centric legal QA, judge-human agreement reached 88.30%, but false negatives accounted for 10.80% of cases, meaning the judge was stricter than human raters.

The warning sign every data team should internalize: in one study the same judge that reached 0.83 agreement with at least one expert dropped to 0.51 against majority-of-experts judgments. Your judge's score is a function of your reference standard. Define the standard first (majority of three reviewers, adjudicated disagreements), then calibrate. Run the judge on 300 human-labeled items, compute Cohen's kappa and a confusion matrix, and record its directional bias so you can correct the headline number. Also note that consistency is not correctness: a recent study shows high self-consistency does not necessarily indicate high agreement with human judgments.

Your LLM judge's agreement score is a property of your reference standard, not of the judge.

Which metrics matter for due diligence specifically?

Recall per clause type, abstention rate, and cost per correct answer. Precision-weighted F1 hides the only failure that kills a deal: a missed clause nobody re-reads. Use F2 alongside F1 to weight recall, as ContractEval does, where proprietary models led on correctness and GPT-4.1 and GPT-4.1 mini posted the top F1 scores of 0.641 and 0.644 with F2 at 0.672 and 0.678.

Track false abstention explicitly. ContractEval measured a "laziness" rate and found that open-source models produce "no related clause" responses more often even when relevant clauses are present, a failure mode that looks like clean output and reads like a clean contract. GPT-4.1's false rate on that axis was 0.071.

Two more findings worth designing around: performance varies across legal categories, with significantly lower correctness on less common or longer clauses, and reasoning mode improved output effectiveness but reduced correctness. Benchmark reasoning on and off rather than assuming it helps. For the economics of the trade-off, see our note on accuracy versus cost per correct answer.

How do you catch errors in the 4,600 documents you never labeled?

Through disagreement signals and cross-document consistency rather than more labels. Three mechanisms carry most of the weight.

Consensus disagreement: run two extractors with different failure profiles (a layout-aware parser and a vision-language model, say) and route every mismatch to review. Because the extraction task is sparse, mismatches are a small fraction of total fields but concentrate a large share of errors.

Self-consistency sampling: re-run the same model at temperature > 0 on a slice and flag fields whose answers vary. Treat variance as a proxy for difficulty, keeping in mind the caveat above that stability alone does not prove correctness.

Corpus-level contradiction checks: the same counterparty appearing with two different governing laws, an assignment clause marked absent in a contract that the master agreement says must contain one, termination notice periods outside the distribution for that contract family. These checks find real errors with zero annotation. Sygnet builds extraction pipelines where field-level confidence and span provenance are exposed per answer, which is what makes this kind of post-hoc auditing possible at all; the same logic applies to M&A and due-diligence document processing more broadly.

Set your review threshold from the calibration slice, not from intuition. If a clause type shows 0.92 recall above a confidence cut-off and 0.55 below it, that cut-off is your human-in-the-loop boundary, and it should be set per clause type.

FAQ

How many contracts do I need to label to trust a clause extraction benchmark?

Roughly 400 documents for a global ±5 point accuracy estimate at 95% confidence, plus 30 to 50 positive examples per clause type you care about. The per-clause allocation matters more than the total. Rare provisions like change-of-control or exclusivity need deliberate over-sampling, since a uniform draw will produce too few positives to estimate recall with a usable interval.

Is CUAD still a valid benchmark for LLM pipelines?

Yes, as a method check rather than a performance ceiling. Its contracts come from real transactions across a broad range of clause types, and annotations were made by law students with 70-100 hours of specialized training under lawyer supervision. It remains the backbone of newer work: ContractEval evaluated 4 proprietary and 15 open-source LLMs on it. It will not reflect your document types, languages or scan quality, so pair it with your own stratified slice.

What accuracy should I expect from an off-the-shelf LLM on clause extraction?

Lower than vendor demos suggest. On CUAD-based clause-level risk identification, the top models scored F1 around 0.64, with Jaccard similarity used to measure output usability separately from correctness. The ContractEval authors conclude that most LLMs perform at a level comparable to junior legal assistants and open-source models need targeted fine-tuning for high-stakes legal settings. Plan for human review on low-confidence and rare clauses.

Should I benchmark quantized or full-precision models?

Benchmark both, then decide on cost. ContractEval found that <cite name="" index="33-8">quantization speeds up inference at the cost of a performance drop</cite> and that it slightly lowers performance, especially in thinking mode, so full-precision models may be preferable for legal tasks unless resource constraints are strict. For diligence work, where a missed clause costs far more than GPU hours, the full-precision default is usually right.

One email when we publish something worth your time.