Lire en français →

CHECKLISTS

GDPR checklist for document AI projects

By Sygnet Research. Written by Sygnet, sourced, checked before publication.

This checklist is for teams building or buying a document AI system that touches personal data: KYC files, payslips, medical reports, lease agreements, identity documents. Use it before you sign a vendor contract or ship a pipeline to production, not after. It won't make you compliant by itself, but it will catch the mistakes that usually surface in an audit or a breach.

Scoping the data

  • List every document type the system will process and flag which ones contain personal data (an invoice is different from a medical report or an ID card)
  • Identify special category data: health info, biometric data on ID cards and passports, union membership mentioned in HR files
  • Confirm a lawful basis for processing (contract, consent, legitimate interest) for each document type
  • Check whether any documents involve minors or vulnerable individuals (stricter rules apply)
  • Map where extracted data lands downstream (CRM, data warehouse, ticketing system)

Vendor and architecture due diligence

  • Confirm where the model runs: your infrastructure, vendor cloud, or a third-party LLM API (this determines who is a processor and who might be a sub-processor)
  • Get a signed Data Processing Agreement from every vendor touching personal data
  • Ask whether documents or extracted text are used to train the vendor's models (many general-purpose LLM APIs do this by default unless you opt out)
  • Check data residency: does processing stay in the EU, or does it transit through US servers
  • Review the vendor's security and compliance documentation, including SOC 2 or ISO 27001 status
  • If comparing providers, run them through a structured process like the IDP vendor evaluation checklist

Logging and retention

  • Decide what gets logged: raw documents, extracted fields, model prompts, confidence scores
  • Set a retention period for each log type and enforce it with automated deletion, not a policy document nobody checks
  • Mask or redact personal data in logs used for debugging or model evaluation
  • Confirm LLM provider logs (if any) also respect your retention rules, not just your own database
  • Review the specifics in our guide on GDPR and SOC 2 compliance for LLM logging

Data subject rights

  • Build a way to find all records tied to one person across documents, extracted fields, and logs (needed for access and erasure requests)
  • Test an erasure request end to end, including backups and any vendor-side copies
  • Document how you'd respond to a request within the one-month statutory window
  • Check that automated decisions (auto-approve, auto-reject) based on extracted data allow for human review if requested

Technical safeguards

  • Encrypt documents at rest and in transit, including temporary storage used during OCR or parsing
  • Apply role-based access so only people who need to see a payslip or a medical report can see it
  • Run a comparison of extraction approaches, since OCR vs VLM choices affect how much raw text gets exposed to third-party models
  • Pseudonymize data where possible before sending it to a general-purpose model
  • Monitor accuracy continuously; a wrong extraction on a bank statement or tax notice can itself become a data protection problem (wrong data linked to the wrong person)

Governance and documentation

  • Complete a Data Protection Impact Assessment if the project involves large-scale or systematic processing (most KYC and insurance claims pipelines qualify)
  • Record the project in your Article 30 processing register
  • Assign an owner for ongoing compliance, not just the launch
  • Decide whether to build in-house or use a vendor, weighing compliance overhead against control; see build vs buy IDP
  • Budget for compliance work when comparing costs, not just licensing fees, using something like the ROI calculator

Common mistakes

  • Treating the DPIA as a one-time document instead of updating it when the document set or model changes
  • Assuming a vendor's "GDPR compliant" marketing claim removes your own obligations as data controller
  • Sending full documents to a general-purpose LLM API without checking its data retention and training policies
  • Logging everything "just in case" without a deletion plan
  • Forgetting that extraction errors on sensitive fields (health status, income, ID numbers) are a data accuracy problem under GDPR, not just a quality metric
  • Skipping a real erasure test until a regulator or a customer actually asks

FAQ

Do we need a DPIA for a document AI project?

Usually yes, if the processing is systematic, large-scale, or involves special category data like health records or biometric ID data. Insurance claims, KYC onboarding, and healthcare document processing almost always qualify. If unsure, run the assessment anyway: it is cheaper than defending the decision not to afterward.

Is using a third-party LLM API automatically non-compliant?

No, but it adds a sub-processor and a data flow you need to document and control. Check contractual training-data opt-outs, residency, and retention. Many vendors offer EU-only processing or zero-retention modes. The risk is not the technology itself; it's using it without a signed DPA or without knowing where the data actually goes.

How does this differ from general vendor evaluation?

A GDPR checklist focuses narrowly on lawful basis, data flows, and rights. The IDP vendor evaluation checklist covers accuracy, cost, and integration too. Use both: compliance can disqualify an otherwise strong vendor, and a non-compliant pipeline is a liability regardless of how accurate it is.

NEXT STEP

See it on your own documents

One email when we publish something worth your time.