Lire en français →

CHECKLISTS

RAG document ingestion checklist

By Sygnet Research. Written by Sygnet, sourced, checked before publication.

This checklist is for teams building a retrieval-augmented generation pipeline who are about to feed PDFs, scans, or mixed-format documents into a vector store. Use it before you write a single line of embedding code, and again whenever you add a new document type to the corpus. Skipping these steps is the most common reason RAG answers look confident but are wrong.

Scope and source quality

  • List every document type that will enter the pipeline, including edge cases like handwritten forms or faxed scans (mixed quality breaks naive chunking)
  • Confirm whether documents are born-digital, scanned, or a mix (this decides whether you need OCR at all, see OCR vs VLM)
  • Check for duplicate or near-duplicate documents in the source set (duplicates skew retrieval relevance)
  • Identify documents with tables, multi-column layouts, or embedded images, since these need different parsing logic than plain text
  • Decide what counts as "stale" content and set a re-ingestion cadence

Parsing and extraction

  • Test extraction accuracy on a representative sample before committing to a parser (see extraction accuracy evaluation)
  • Verify table structure is preserved, not flattened into unreadable text
  • Confirm reading order is correct for multi-column or form-style layouts
  • Check how the parser handles low-quality scans and rotated pages
  • Decide between building your own parsing stack or buying one (see build vs buy IDP)

Chunking strategy

  • Choose a chunk size appropriate to the document type, not a single default for everything
  • Preserve section headers and metadata inside or alongside each chunk (retrieval without context produces nonsense answers)
  • Test overlap settings against actual queries users will ask, not synthetic ones
  • Avoid chunking mid-table or mid-clause, which is common with contracts and claim forms
  • Confirm chunks retain a pointer back to the source document and page number

Metadata and provenance

  • Tag each chunk with document type, source system, and ingestion date
  • Record who or what approved the document for ingestion
  • Store a checksum or hash of the original file for traceability
  • Capture document-level confidence scores if extraction was automated

Security and compliance

  • Classify documents by sensitivity before they reach the vector store (see security and compliance)
  • Confirm whether personal data needs redaction or masking before embedding
  • Check retention rules against your GDPR checklist obligations
  • Verify who has query access to retrieved chunks, not just who has upload access

Evaluation before go-live

  • Run a set of real user questions against the pipeline and manually grade the answers
  • Measure retrieval precision separately from generation quality (bad answers can come from either stage)
  • Spot-check answers that cite low-confidence or OCR-heavy chunks
  • Compare cost per correct answer against your current process, not just raw accuracy (worth reading if you haven't: accuracy vs cost per correct answer)

Common mistakes

  • Treating all documents as plain text and skipping layout-aware parsing entirely
  • Chunking by a fixed character count regardless of document structure
  • Ingesting scanned contracts or claim forms without testing OCR quality first
  • Never revisiting ingestion settings after the first successful test batch
  • Storing no provenance, so nobody can explain why the model cited a given chunk
  • Assuming vector similarity alone guarantees relevance

FAQ

How is RAG ingestion different from standard document extraction?

Standard extraction pulls structured fields for a database. RAG ingestion prepares unstructured text for retrieval, so the priority shifts from precise field mapping to preserving context and meaning across chunks. You still need accurate parsing underneath, but the output target is different: readable, retrievable passages rather than key-value pairs.

Do I need OCR if my documents are already digital PDFs?

Often yes, partially. Many "digital" PDFs contain scanned images embedded in them, or text layers that are garbled from poor export tools. Always test a sample before assuming the text layer is reliable. See OCR vs VLM for how different extraction methods handle this.

How often should I re-run this checklist?

Every time you add a new document type, change your parser or chunking logic, or notice a drop in answer quality. Static corpora can go months without review, but anything fed by live uploads, like contract analysis or insurance claims, needs quarterly checks at minimum.

NEXT STEP

See it on your own documents

One email when we publish something worth your time.