SOC 2 Type II: Verifying Document Content Security in AI/LLM Vendors
By Sygnet Research, checked before publication

Key takeaways
- A SOC 2 Type II opinion covers only the controls the vendor chose to put in scope, so ask for Section III (system description) and Section IV (tests and exceptions), not the one-page certificate. Scope decisions often reveal more about security posture than the controls themselves.
- "We don't log document content" is a claim about persistence, not purpose. Demand it as a tested control with a named control ID, a sample size, and a result in Section IV.
- Check the model provider underneath. OpenAI's default is that abuse monitoring logs, which can contain prompts and responses, are retained for up to 30 days, and that default survives unless the vendor holds an approved ZDR or Modified Abuse Monitoring agreement.
- Pair the report with an Article 28 DPA, a named subprocessor list and a bridge letter if the audit period ended more than 90 days ago.
What does a SOC 2 Type II actually prove about document content?
It proves that an independent CPA firm tested a defined set of controls over a defined period and reported the exceptions. Nothing more. A clean SOC 2 audit means the vendor passed the controls they chose to include, not that they are comprehensively secure. If "no document content is written to persistent storage" is not written as a control objective in Section III and tested in Section IV, the report says nothing about it, however confidently the sales deck does.
So your first request is boring and decisive: the full report under NDA, including the system description and the test matrix. Read the system boundary. A vendor can scope its SOC 2 to the control plane (dashboard, auth, billing) and leave the inference path out. Then check the exceptions table. Exceptions are normal and often informative; a report with zero exceptions across 12 months and 90 controls usually means the controls were written loosely.
A SOC 2 report with no exceptions across a full year is not a flex, it is a hint that the controls were written to be passed.
Also check who signed it, the audit period end date, and whether the Trust Services Criteria include Confidentiality and Privacy or only Security and Availability. Most OCR and LLM pipelines handle customer confidential data, so Security alone is a thin scope.
Which specific controls should appear in Section IV for a no-content-logging claim?
Five, and you can name them when you ask. First, a data retention control stating that document bytes and extracted text are deleted within N seconds or minutes of response delivery, with the deletion job tested on a sample of requests. Second, a logging control stating that application and infrastructure logs exclude request and response bodies, tested by inspecting log schemas and sampled log entries. Third, an access control proving no human role can retrieve customer document content in production, including support and on-call engineers. Fourth, a change management control, because a log-level config change is exactly how content starts leaking into Datadog. Fifth, a subprocessor control for the model provider.
If the vendor fine-tunes or evaluates models, ask separately whether any customer content reaches a training or eval corpus, and which control proves it does not. That is a different data path from inference, with different storage and different people touching it.
Why is the carve-out section the most important page in the report?
Because the layer that holds your documents is often the layer the auditor never tested. Under the carve-out method, the subservice organization's controls are excluded from the audit's scope and testing entirely; the report simply states that certain criteria rely on the subservice organization's own controls, and you are expected to separately obtain and review that subservice organization's report. Your auditor does not test the subservice organization's controls directly.
A carve-out is standard and not a red flag by itself. But it creates a real assurance gap: if a vendor runs on AWS and AWS is carved out, the audit does not cover the infrastructure layer (networking, compute, storage, physical security). The part that matters for your question is whether the foundation model provider is named as a subservice organization at all. When a critical subservice is carved out, the report lists Complementary Subservice Organization Controls, the activities the vendor relies on that provider to perform. Read that list. If OCR output is sent to a third-party LLM and that provider appears nowhere in Section III, the system description is misleading and you should say so in writing.
How do you verify the model provider's retention, not just the vendor's?
Ask for the retention configuration in writing, per endpoint, plus evidence of the approval. The default behaviour of the big APIs is not zero retention. Abuse monitoring logs may contain prompts, responses and derived metadata such as classifier outputs, and by default they are generated for all API feature usage and retained for up to 30 days unless longer retention is required by law. Customer content can be excluded from those logs only for eligible customers approved for Zero Data Retention or Modified Abuse Monitoring.
Two practical consequences. ZDR is not a checkbox: it requires provider approval and acceptance of additional requirements, after which retention controls are configured at organization level and can be overridden per project. So "we have ZDR" needs a screenshot of the org and project settings, not a sentence. And coverage is partial: some endpoints and features are not ZDR eligible, so enabling ZDR for an organization does not convert every stateful API object into an ephemeral one. Conversations, files and vector stores keep application state until deleted, whatever the setting.
Training is a separate question from storage. A training opt-out is a control over purpose; ZDR is a control over persistence. Defaults differ by vendor: OpenAI states API data is not used to train its models unless you opt in, while Mistral trains on API data by default and requires an opt-out through an admin setting, as does Cohere via a Data Controls toggle.
What evidence beats a SOC 2 report for this specific claim?
Evidence you can test yourself, or that an auditor tested with a stated sample. Ranked by how hard it is to fake:
| Evidence | What it actually proves | How hard to fake | Ask for it when |
|---|---|---|---|
| Section IV test of a retention/log-redaction control | Operating effectiveness over the period, with sample size and exceptions | Hard | Always |
| Log schema export plus a live trace of your own test document | Content is absent from the observability stack today | Hard | Pre-contract POC |
| Model provider ZDR/MAM approval evidence and per-project settings | The inference path does not persist content | Medium | Any third-party LLM in the chain |
| Named subprocessor list with locations and transfer mechanism | Who else touches the bytes | Medium | Always, under GDPR |
| Independent penetration test report, current year | Exploitability, not policy | Medium | Regulated workloads |
| Bridge letter | Management says nothing material changed since period end | Easy | Report older than 90 days |
| Completed CAIQ or security questionnaire | Self-assertion as of the date answered | Easy | Triage only |
A bridge letter is issued by management, not the auditor, covers the gap from period end to the letter date, and carries management representation only with no auditor involvement. Past six months, many vendor risk teams stop accepting letters and ask when the new report arrives, because management is then asserting more unexamined time than the auditor examined. Our IDP vendor evaluation checklist sequences these requests so you are not negotiating evidence after the contract is signed.
What must the DPA say that the SOC 2 cannot?
The DPA is where no-content-logging becomes enforceable. Article 28(3)(h) requires the processor to make available all information needed to demonstrate compliance and to allow and contribute to audits and inspections. Offering a SOC 2 report is a reasonable first step, but a clause that rules out any inspection conflicts with that obligation. On subprocessors, the processor "shall not engage another processor without prior specific or general written authorisation of the controller."
AI vendors diverge from standard SaaS on two clauses specifically. Retention must cover prompts, outputs, logs, embeddings and caches, not just stored records and backups, and the subprocessor list must disclose the foundation model provider actually running inference. Add a deletion SLA with a number, and a breach notification window; a review often surfaces a 96-hour window where you wanted 72, or a transfer to a jurisdiction without an adequacy decision. Our GDPR checklist for document AI projects maps these clauses to the document pipeline, and Sygnet publishes its own posture and subprocessors on its security and compliance page.
Retention language written for a database does not cover a vector store, an embedding cache or a prompt log, and that is precisely where document content ends up.
FAQ
Can a vendor be SOC 2 Type II certified and still store my documents?
Yes, easily. The report attests to the controls in scope, and storing customer documents securely is a perfectly auditable control. SOC 2 is not a privacy minimum; it is an opinion on controls the vendor selected. If zero retention matters to you, it has to appear as an explicit control objective in the system description and be tested in Section IV with a stated sample, or it is not covered.
Does zero data retention mean nothing is kept anywhere?
No. ZDR normally means covered prompts and outputs are not written to durable provider storage after processing, and it does not automatically cover metadata, tools, customer logs, files, caches, safety signals or legal exceptions. Stateful endpoints such as files and vector stores keep application state until something deletes it. Ask which endpoints are in use, which are ZDR eligible, and what the vendor's own application retains for retries, queues and human review.
How do I test the no-logging claim myself during a POC?
Send a document with a unique canary string in it, then exercise every retrieval surface you can: support search, the audit log API, error traces, any admin console. Ask the vendor to search their observability stack for the canary and return a dated screenshot. Combine that with a log schema export showing request bodies are excluded. The method is the same one used for extraction accuracy evaluation: define the test before the vendor demo, not after.
One email when we publish something worth your time.