COMPARISONS
Sygnet vs Unstructured
By Sygnet Research. Written by Sygnet, sourced, checked before publication.
Sygnet vs Unstructured covers two tools that both turn documents into JSON, but they sit at different points in a pipeline. Unstructured is document ETL built for retrieval: it parses files into typed elements, chunks them, optionally embeds them, and writes the result into a vector store or warehouse. Sygnet is an intelligent document processing platform and API: you define (or let it infer) a schema, and it returns specific field values with per-field confidence scores and bounding-box provenance, with validation and human review on top. This page is written by Sygnet, so read the comparison with that in mind; every claim about Unstructured below links to its own documentation, pricing page or blog, and where we could not verify something we say so.
At a glance
| Criterion | Unstructured | Sygnet |
|---|---|---|
| Document types | Broad file coverage via format-specific partitioners; the open-source API supports the partitionable document types of the unstructured[all-docs] dependency, and the library routes a file automatically using libmagic-based type detection (docs) | Any document type, zero-shot, including long contracts, forms and scans |
| Setup / training | No model training; you configure a pipeline. Source connectors ingest data, destination connectors write results, and a workflow adds chunking, embedding and scheduling (docs) | No training, no templates; schema inference on first upload |
| Extraction approach | Partitioning into elements: content is converted into structured document elements and metadata in a consistent JSON format (docs), with strategies including auto, fast, hi_res and ocr_only plus a VLM strategy (docs) | Field-level extraction against a schema, returning typed JSON values |
| Confidence and provenance | Elements carry metadata such as page number and coordinates, and table elements expose an HTML representation under text_as_html (docs); we found no published per-field confidence score for extracted business fields | Per-field confidence plus bounding-box provenance for each value |
| Validation | Cleaning and staging functions are part of the library (partition, chunk, clean and stage raw source documents, docs); business validation rules are not a documented product feature we could verify | Validation rules and cross-document checks (totals, dates, identity consistency) |
| Review workflow | Not documented as a product feature on the pages we reviewed; review would typically live in your own application | Human review queue, triggered only when confidence falls below your threshold |
| Deployment / hosting | Managed SaaS plus self-hosting: a self-hosted REST API for partitioning documents with the open-source library (GitHub), and a dedicated instance or VPC option with full data isolation (pricing) | EU-hosted on Google Cloud, europe-west1 |
| Data policy | The open-source library sends lightweight analytics pings on import and per top-level partition call by default (GitHub); Unstructured has publicised SOC 2 Type 2 compliance for its serverless API (blog) | No training on customer documents (security) |
| Pricing model | Per page. The first 10,000 pages are not billed (a one-time account credit), then billing is monthly at $0.015 per page processed (billing docs); third-party sources quote other figures, so confirm against the live page | Per document (pricing) |
| Enrichment | VLM-powered enrichments exist, for example image elements sent to a VLM to generate descriptions that can be embedded alongside text (blog) | Not a focus; output is field values, not chunk enrichment |
| Chunking for RAG | Several strategies, including By Title semantic chunking and By Similarity, which uses an embedding model to group topically similar elements (docs) | No chunking; Sygnet is not a retrieval pipeline |
| API | REST API and SDKs; the API detects file type, selects a partitioner, and returns document elements as JSON or CSV (GitHub) | REST API returning structured JSON (structured output) |
Where Unstructured is strong
If your end goal is retrieval, Unstructured is the more natural fit, and we would say so to any team building RAG. The element model is the reason. Partitioning breaks a document into elements such as Title, NarrativeText and ListItem, letting you decide what content to keep for your application, which is exactly the granularity a chunker wants. From there the pipeline continues in one place: embedding is an optional step rather than a requirement, so you can run a partitioner and a chunker, skip vector generation, and still get clean chunks.
Three other things matter in practice. Cost control: Auto routing sends text-only pages to Fast and sends more complex pages to VLM or High Res, so you are not paying heavy compute on simple pages. Open source: you can run the parser yourself, read the code, and pin versions. Connector breadth: Unstructured offers multiple destination connectors, including all major vector databases, which removes a lot of glue code between parsing and your index. For teams whose question is "how do I get ten million pages of mixed files into a searchable index," that combination is hard to beat.
Where Sygnet is strong
Sygnet is built for the other question: "what is the invoice total, the lease end date, the policyholder's IBAN, and how sure are you?" Extraction is zero-shot against a schema, which Sygnet can infer from the document itself, so onboarding a new form does not mean collecting training samples or drawing template zones.
Each returned field carries a confidence score and a bounding box pointing back to the region it came from, which makes an answer auditable instead of merely plausible (field-level provenance). On top of that sit validation rules and cross-document checks: line items against a stated total, an address on a payslip against the one on a utility bill. Human review is routed only when confidence is low, so reviewer time goes to the small share of documents that actually need a second pair of eyes.
The operational details are deliberate. Processing runs in the EU on Google Cloud europe-west1, customer documents are not used for training, and pricing is per document rather than per page, which keeps a thirty-page contract from costing thirty times a receipt. Sygnet is currently in early access.
Which one for which team
- Building RAG or search over a document corpus. Pick Unstructured. You need elements, chunks and embeddings written to a vector store, and that is what its workflow model produces.
- Running a business process on specific fields. Pick Sygnet: claims triage, KYC checks, invoice posting, lease abstraction. Confidence thresholds and provenance are what let you automate a decision and defend it later (confidence thresholds for claims).
- Both, in sequence. Many teams index everything for search and extract fields from the subset that drives money or compliance. Using two tools here is reasonable, not wasteful.
- Strict EU data residency with a named region. Sygnet states europe-west1. Unstructured offers a dedicated instance or VPC with full data isolation, which can also satisfy residency requirements, though the specifics come from sales rather than a public page.
- A small team with engineering appetite and no budget. Start with the open-source library. You will write more code, but you control the whole path, and the first pages cost nothing either way (build vs buy).
FAQ
Can Unstructured do field extraction like Sygnet?
You can get there, but you assemble it. Unstructured gives you clean elements and chunks; mapping those to named business fields, scoring confidence per field, and validating against rules is work you do in your own code or with an LLM call. Sygnet packages that step. If your fields are few and stable, the DIY route is perfectly sensible.
Can Sygnet feed a RAG pipeline?
Partly. Sygnet returns structured field values, which are excellent as filterable metadata on a document record: counterparty, dates, amounts, jurisdiction. It does not chunk or embed text, so for passage-level retrieval you still want a parser and chunker. A common pattern is Unstructured for the text index and Sygnet for the fields that must be exact.
How do the two pricing models compare?
They are not directly comparable. Unstructured bills per page processed after a one-time 10,000-page credit, which suits large, page-heavy corpora. Sygnet bills per document, which suits workflows where documents vary from one page to eighty (pricing). Model both against your real page distribution before deciding; the ROI calculator helps with the extraction side.
NEXT STEP
See it on your own documents
One email when we publish something worth your time.