Lire en français →

COMPARISONS

Sygnet vs Azure AI Document Intelligence

By Sygnet Research. Written by Sygnet, sourced, checked before publication.

Azure AI Document Intelligence is Microsoft's document extraction service inside the Azure stack: prebuilt models for common forms, a layout engine, and custom models you train on your own labeled samples. Sygnet is an intelligent document processing platform and API built around zero-shot extraction, so you define a schema (or let the system infer one) instead of collecting and labeling a training set. Sygnet wrote this page, which is a reason to read it skeptically; we have tried to keep every claim about Azure sourced to Microsoft's own documentation and pricing pages, and to be clear about what Azure does better. If your documents are invoices, receipts, IDs or tax forms and your company already runs on Azure, the honest answer may well be Azure, and the sections below say why.

At a glance

CriterionAzure AI Document IntelligenceSygnet
Document typesPrebuilt models cover a defined list: Document, Layout, Receipt, Invoice, ID, W-2, 1098 tax forms, Health insurance card and Contract. Anything else needs a custom modelAny document type, including long and irregular ones, without a per-type model
Setup / trainingCustom extraction uses supervised learning on labeled data: a minimum of five completed forms of the same type, and the training set is uploaded to an Azure blob storage containerNo training set. You describe fields, or let Sygnet infer the schema from the document
Extraction approachTwo custom model families: a custom neural model, fine-tuned on your labeled dataset, supporting structured, semi-structured and unstructured documents, plus template models for fixed layoutsZero-shot extraction with schema inference, no per-layout model to maintain
Classification / routingSeparate custom classifier: training requires at least two distinct classes and a minimum of five samples per classDocument type is identified as part of extraction
ConfidenceYes, at field and word level. Field objects include a confidence value, for example an InvoiceTotal field returned with "confidence": 0.945Per-field confidence scores on every extracted value
ProvenanceYes. Bounding polygons are currently returned as 4-vertex quadrilaterals, and fields carry boundingRegions with page number and polygonBounding-box provenance per field, exposed as field-level provenance
ValidationNot documented as a built-in business rules engine in the pages we reviewed; validation is typically built in your own pipelineValidation rules and cross-document checks run inside the platform
Review workflowStudio provides labeling and testing tools; a production human-review queue is not part of the service as documentedHuman review routed only on low confidence
DeploymentCloud service plus containers. Disconnected containers run APIs offline, billed against a purchased commitment tier, currently available for custom and invoice modelsHosted only, EU region (Google Cloud europe-west1)
Data policyData and results are temporarily stored in Azure Storage in the same region as the request, then deleted 24 hours after the analyze request; a delete API is available in v4.0No training on customer documents. Details on the security page
Pricing modelPer page, pay as you go, with a free tier of 0 to 500 pages per month and published per-1,000-page rates per model type; custom extraction was reduced to $30 per 1,000 pages, and commitment tiers offer an upfront monthly fee for high volume at a discountPer document. See pricing
APIREST API, SDKs and Studio: you can use the Document Intelligence Studio, REST API, or client librariesREST API returning JSON, currently in early access

Where Azure AI Document Intelligence is strong

Azure's biggest advantage is that it is Azure. Billing, identity, networking and storage already exist in most enterprises, procurement is a line item rather than a new vendor review, and the service plugs into the rest of the platform. The prebuilt catalogue is broad and needs no work from you: Document, Layout, Receipt, Invoice, ID, W-2, 1098 tax forms, Health insurance card and Contract are all available out of the box.

The output format is unusually disciplined. Fields come back with confidence and bounding regions, for instance a currency field with a typed value, its page polygon and a confidence score, which makes building your own review UI straightforward.

Two more real strengths. Price per page at the low end is hard to beat: the free tier covers up to 500 pages per month and custom extraction dropped to $30 per 1,000 pages in June 2024. And for regulated or air-gapped work, disconnected containers run the APIs with no internet connection and no billing traffic, currently for custom and invoice models. No hosted-only vendor, Sygnet included, can match that.

Where Sygnet is strong

Sygnet starts from a different assumption: you should not need a model per document type. Extraction is zero-shot, with schema inference, so a new form, a foreign-language lease or a 90-page agreement can be processed the first time you see it. That removes the labeling step and the retraining cycle when a supplier changes layout.

The second difference is what surrounds the extraction. Sygnet returns per-field confidence scores and bounding-box provenance, then applies validation rules and cross-document checks (does the total match the lines, does the name on the payslip match the ID) and routes to human review only when confidence is low. That logic lives in the platform rather than in code you maintain, which is the whole argument in build vs buy.

Third, hosting and data handling are narrow on purpose: EU-hosted on Google Cloud europe-west1, and no training on customer documents. Pricing is per document, not per page, so a long contract does not cost twenty times a one-pager. Sygnet is in early access, which is a genuine caveat: fewer integrations, fewer years of production history, and a smaller ecosystem than a Microsoft service.

Which one for which team

  • You are an Azure shop with high-volume, stable forms. Invoices, receipts, IDs, tax forms. Use Azure. The prebuilt models cost little per page and your platform team already knows the tooling.
  • You need on-premises or air-gapped processing. Azure, via disconnected containers that work without internet connectivity. Sygnet cannot serve this.
  • Your documents are a long tail that keeps changing. Claims packs, due diligence folders, supplier paperwork from 400 vendors. Training a model per layout is the wrong shape of work. See insurance claims or contract analysis.
  • You need validation and review as product, not as project. If your real cost is the people checking output, pick the tool that decides what to check. Our view on thresholds is in AI confidence thresholds for claims processing.
  • EU data residency is a hard constraint and you cannot run containers. Sygnet is EU-hosted by default; with Azure you choose a region, and data is temporarily stored in the same region as the request.

FAQ

Can Azure handle document types outside its prebuilt list?

Yes, through custom models. You label examples and train: five completed forms of the same type minimum, with supervised learning on your labeled set, and the neural model is fine-tuned on that data. That works well for forms you see repeatedly. It works less well when every document is slightly different, because each new variant means more labeling.

How do the pricing models actually compare?

They are not directly comparable. Azure bills per page by model type, with separate per-1,000-page rates for Read, prebuilt models, custom classification, custom extraction and add-ons and commitment tiers for volume. Sygnet bills per document. Build your estimate on average page counts, and include classification if your pipeline needs it. Our ROI calculator models the review cost too.

Does either service train on my documents?

Sygnet does not train on customer documents. For Azure, Microsoft documents a short retention window: data and results sit in Azure Storage in the request's region and are deleted 24 hours after the analyze request, with an API to delete sooner. We did not find a statement about model training on the Microsoft pages we reviewed, so check Microsoft's data, privacy and security documentation directly rather than taking our word for it.

A closing position, since you are comparing vendors and deserve one. Azure wins on cost per page, platform integration and offline deployment. Sygnet wins when the documents are varied, when you cannot wait to label a training set, and when confidence, provenance and routing matter more than the raw price of OCR. If you want the underlying technical trade-off rather than the vendor framing, read OCR vs VLM.

NEXT STEP

See it on your own documents

One email when we publish something worth your time.