Lire en français →

SOLUTIONS

Document processing for M&A and due diligence

By Sygnet Research. Written by Sygnet, sourced, checked before publication.

A data room for a mid-sized acquisition can hold anywhere from a few hundred to tens of thousands of documents: contracts, cap tables, financial statements, corporate registries, leases, employment agreements, IP filings. A deal team's real job is judgment, not retyping figures out of PDFs, yet that's where most associate hours actually go. Once extraction and cross-checking are automated, the team spends its time on what the numbers mean rather than on finding them. That shift doesn't just save hours, it changes how much diligence a firm can actually do inside a tight exclusivity window.

Documents and use cases

Use caseDocuments involvedWhat gets extracted or checked
Corporate structure reviewCompany statutes, Kbis extracts, cap tablesShareholders, share classes, registered address, legal representatives
Contract risk scanContracts, leases, supplier agreementsChange-of-control clauses, termination rights, renewal dates, indemnities
Financial statement pullBank statements, tax notices, audited accountsRevenue lines, balances, tax liabilities, account numbers
Employment liability checkPayslips, employment contractsSalary bands, notice periods, benefits obligations
Real estate and lease auditLease agreementsRent, term, break clauses, guarantor details
IP and registry verificationKbis extracts, statutes, filingsEntity identifiers, SIREN/SIRET, filing dates
Litigation and claims exposureInsurance claim forms, correspondenceClaim status, amounts, open liabilities
Cross-document consistencyAll of the aboveMatching names, addresses, dates and figures across sources

Where manual processing breaks

The problem isn't volume alone, it's inconsistency. A data room assembled by the seller's advisors mixes scanned contracts, native PDFs, spreadsheets and email threads, with no shared naming convention and often no index worth trusting. Associates end up building tracking spreadsheets by hand, copying figures from one document into another, and that copying is where errors creep in: a shareholder's name spelled differently in the statutes versus the Kbis extract, a rent figure that doesn't match between the lease and the financial statement.

Deadlines make this worse. Exclusivity periods run four to eight weeks, and the data room can be updated daily as the seller's side adds documents. A checklist done manually on day one is stale by day ten. Nobody has time to re-verify everything, so risk gets triaged by gut feeling about which folders "look important," which means real exposure can sit undiscovered in a contract nobody reopened.

Junior staff carry most of this load, which is expensive and doesn't scale well when three deals are live at once. And because the work is manual, there's rarely a clean audit trail showing who checked what, when, and against which version of a document. If a warranty claim surfaces after closing, reconstructing the diligence trail becomes its own project.

How Sygnet fits

Sygnet reads the mixed document types found in a data room (scanned PDFs, native files, spreadsheets) and infers a schema per document type rather than requiring a rigid template for each one. Every extracted field carries a confidence score, so low-confidence fields (a faint signature, a poorly scanned clause) get flagged for human review instead of silently passing through.

Validation rules catch obvious problems: a lease end date before its start date, a tax ID that doesn't match a known format. Cross-document checks go further, comparing a shareholder's name across statutes, Kbis extract and cap table, or a rent figure across lease and financial statement, and surfacing mismatches automatically. Every extraction keeps field-level provenance: which page, which document version, which model run produced it. That audit trail matters if a dispute arises later.

Data stays on EU infrastructure, which matters for cross-border deals involving sensitive commercial information. Results are available through an API with webhooks, so a data room platform or deal management tool can pull structured output as documents arrive rather than in a batch at the end. See our security and compliance page for infrastructure detail, and contract analysis for how clause-level extraction works in more depth.

Compliance and data protection

Diligence documents routinely contain personal data (names, salaries, addresses in leases and payslips), which brings them under GDPR if any party or the target has EU operations. That means data minimisation, defined retention, and a lawful basis for processing, even when the processing is temporary and deal-specific. Firms handling this data should confirm their processor's hosting location and logging practices; our note on GDPR and SOC 2 compliance for LLM logging covers what to check.

Beyond GDPR, confidentiality obligations in the NDA governing the deal usually impose stricter terms than statute alone, particularly around who can access documents and how long copies persist after the deal closes or falls through. Any automation layer should support prompt, verifiable deletion once the engagement ends.

FAQ

Can this handle a data room with no consistent folder structure or naming?

Yes. Schema inference works per document, not per folder, so it doesn't depend on the seller's organisation scheme. Each file is classified and processed on its own, and results are indexed by document type and extracted fields, which makes it easier to search across a messy data room than to fix the folder structure itself.

How does cross-document checking actually work?

The system links fields it believes refer to the same real-world entity (a company name, an address, a monetary figure) across different documents, then flags disagreements. It's a matching problem, not a rules engine, so it copes with minor formatting differences, but any flagged mismatch still needs a human to decide if it's a real discrepancy or a formatting quirk.

No. It surfaces clauses like change-of-control or termination rights and extracts their terms, but deciding whether a clause is a dealbreaker is still a legal judgment call. The value is in not missing a clause buried on page 40 of a scanned lease, not in automating the risk assessment itself.

NEXT STEP

See it on your own documents

One email when we publish something worth your time.