CHECKLISTS
GDPR checklist for document AI projects
By Sygnet Research. Written by Sygnet, sourced, checked before publication.
This checklist is for teams building or buying a document AI system that touches personal data: KYC files, payslips, medical reports, lease agreements, identity documents. Use it before you sign a vendor contract or ship a pipeline to production, not after. It won't make you compliant by itself, but it will catch the mistakes that usually surface in an audit or a breach.
Scoping the data
- List every document type the system will process and flag which ones contain personal data (an invoice is different from a medical report or an ID card)
- Identify special category data: health info, biometric data on ID cards and passports, union membership mentioned in HR files
- Confirm a lawful basis for processing (contract, consent, legitimate interest) for each document type
- Check whether any documents involve minors or vulnerable individuals (stricter rules apply)
- Map where extracted data lands downstream (CRM, data warehouse, ticketing system)
Vendor and architecture due diligence
- Confirm where the model runs: your infrastructure, vendor cloud, or a third-party LLM API (this determines who is a processor and who might be a sub-processor)
- Get a signed Data Processing Agreement from every vendor touching personal data
- Ask whether documents or extracted text are used to train the vendor's models (many general-purpose LLM APIs do this by default unless you opt out)
- Check data residency: does processing stay in the EU, or does it transit through US servers
- Review the vendor's security and compliance documentation, including SOC 2 or ISO 27001 status
- If comparing providers, run them through a structured process like the IDP vendor evaluation checklist
Logging and retention
- Decide what gets logged: raw documents, extracted fields, model prompts, confidence scores
- Set a retention period for each log type and enforce it with automated deletion, not a policy document nobody checks
- Mask or redact personal data in logs used for debugging or model evaluation
- Confirm LLM provider logs (if any) also respect your retention rules, not just your own database
- Review the specifics in our guide on GDPR and SOC 2 compliance for LLM logging
Data subject rights
- Build a way to find all records tied to one person across documents, extracted fields, and logs (needed for access and erasure requests)
- Test an erasure request end to end, including backups and any vendor-side copies
- Document how you'd respond to a request within the one-month statutory window
- Check that automated decisions (auto-approve, auto-reject) based on extracted data allow for human review if requested
Technical safeguards
- Encrypt documents at rest and in transit, including temporary storage used during OCR or parsing
- Apply role-based access so only people who need to see a payslip or a medical report can see it
- Run a comparison of extraction approaches, since OCR vs VLM choices affect how much raw text gets exposed to third-party models
- Pseudonymize data where possible before sending it to a general-purpose model
- Monitor accuracy continuously; a wrong extraction on a bank statement or tax notice can itself become a data protection problem (wrong data linked to the wrong person)
Governance and documentation
- Complete a Data Protection Impact Assessment if the project involves large-scale or systematic processing (most KYC and insurance claims pipelines qualify)
- Record the project in your Article 30 processing register
- Assign an owner for ongoing compliance, not just the launch
- Decide whether to build in-house or use a vendor, weighing compliance overhead against control; see build vs buy IDP
- Budget for compliance work when comparing costs, not just licensing fees, using something like the ROI calculator
Common mistakes
- Treating the DPIA as a one-time document instead of updating it when the document set or model changes
- Assuming a vendor's "GDPR compliant" marketing claim removes your own obligations as data controller
- Sending full documents to a general-purpose LLM API without checking its data retention and training policies
- Logging everything "just in case" without a deletion plan
- Forgetting that extraction errors on sensitive fields (health status, income, ID numbers) are a data accuracy problem under GDPR, not just a quality metric
- Skipping a real erasure test until a regulator or a customer actually asks
FAQ
Do we need a DPIA for a document AI project?
Usually yes, if the processing is systematic, large-scale, or involves special category data like health records or biometric ID data. Insurance claims, KYC onboarding, and healthcare document processing almost always qualify. If unsure, run the assessment anyway: it is cheaper than defending the decision not to afterward.
Is using a third-party LLM API automatically non-compliant?
No, but it adds a sub-processor and a data flow you need to document and control. Check contractual training-data opt-outs, residency, and retention. Many vendors offer EU-only processing or zero-retention modes. The risk is not the technology itself; it's using it without a signed DPA or without knowing where the data actually goes.
How does this differ from general vendor evaluation?
A GDPR checklist focuses narrowly on lawful basis, data flows, and rights. The IDP vendor evaluation checklist covers accuracy, cost, and integration too. Use both: compliance can disqualify an otherwise strong vendor, and a non-compliant pipeline is a liability regardless of how accurate it is.
NEXT STEP
See it on your own documents
One email when we publish something worth your time.