OCR & document parsing pricing compared
Text, tables and layout from documents. Page-priced APIs and token-priced vision models need different cost assumptions.
14 services · sources last checked 2026-10-06 · how we compare
What the most-compared services cost
Estimated yearly bill at 100K pages / month. Enter your own usage below.
- Amazon Textract$18,000/yrAnalyzeDocument Tables + Layout · $15 per 1k pages · Published rates
- Azure AI Document Intelligence (Layout)$10,800/yrCommitment 100K pages · $9 per 1k pages · Published rates
- Mistral OCR$4,800/yr$4 per 1k pages · Published rates
- LlamaParse (LlamaIndex)$4,500/yrStarter · $3.75 per 1k pages · Published rates
Enter your usage to estimate each bill. Services differ in what they include; check the rates, assumptions and extra charges below.
The estimates below use the default usage. Choose your current provider to compare costs.
| Provider | Parent jurisdiction | Features / pricing source | Est. monthly | Est. yearly |
|---|---|---|---|---|
| GPT-6 Luna (via OpenAI API)OCR to markdown with OpenAI's small GPT-6 Luna vision-capable model; cost derived from token prices. | 🇺🇸 United States | EU residencyGDPR DPAEstimated ratesPage cost uses an assumed image-token multiplier plus 300 prompt and 800 output tokens. Confirm actual token usage. | $60$0.599 per 1k pages | $719 |
| Mistral Small 3.2 (via Scaleway Generative APIs)OCR to markdown with the open-weight Mistral Small 3.2 vision model hosted in Paris by Scaleway; cost derived from token prices. | 🇫🇷 France | EU residencyGDPR DPAModeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary. | $70$0.704 per 1k pages | $845 |
| Qwen3-VL-235B (via DeepInfra)OCR to markdown with Alibaba's open-weight Qwen3-VL vision model on DeepInfra's pay-per-token API; cost derived from token prices. | 🇺🇸 United States | Modeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary. | $104$1.037 per 1k pages | $1,244 |
| Gemini 3.1 Flash-Lite (via Gemini API)OCR to markdown with Google's cheapest Gemini vision model via the Gemini Developer API; cost derived from token prices. | 🇺🇸 United States | GDPR DPAModeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary. | $156$1.555 per 1k pages | $1,866 |
| LlamaParse (LlamaIndex)LlamaIndex's hosted document parser with tiered LLM/VLM parsing modes, billed in credits; US and EU regions. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $375Starter · $3.75 per 1k pages | $4,500 |
| Mistral OCRMistral AI's OCR model API (OCR 4.1) returning markdown with tables, images and block-level bounding boxes; EU-hosted. | 🇫🇷 France | EU residencyGDPR DPAPublished rates | $400$4 per 1k pages | $4,800 |
| Azure AI Document Intelligence (Layout)Microsoft's document OCR API; the Layout model returns text, tables, selection marks and structure as markdown. EU regions available. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $900Commitment 100K pages · $9 per 1k pages | $10,800 |
| Google Cloud Document AI (Layout Parser)Google Cloud document processing; Layout Parser returns text, tables, headings and chunks, Enterprise OCR returns text only. EU location available. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $1,000$10 per 1k pages | $12,000 |
| ReductoDocument parsing and extraction API combining layout models and VLMs for complex PDFs, tables and forms. | 🇺🇸 United States | Published rates | $1,000$10 per 1k pages | $12,000 |
| Amazon TextractAWS document OCR API; AnalyzeDocument extracts tables and layout (reading order, headers) from scans and PDFs, with EU regions available. | 🇺🇸 United States | EU residencyGDPR DPAPublished rates | $1,500AnalyzeDocument Tables + Layout · $15 per 1k pages | $18,000 |
| UnstructuredHosted document ETL platform that partitions PDFs, images and Office files into structured elements and tables for RAG pipelines. | 🇺🇸 United States | GDPR DPAopen sourcePublished rates | $1,500$15 per 1k pages | $18,000 |
| ABBYY VantageEnterprise intelligent document processing platform from long-time OCR vendor ABBYY, sold through sales-led volume subscriptions. | 🇺🇸 United States | GDPR DPAcontractNo listed rates | —Contract only · no public price | — |
Text + layout/table extraction. Custom-trained form extractors and handwriting are priced differently, and quality varies by document type.
Need EU residency? Turn on the EU filter to pick a plan that includes it, which may cost more. “EU plan available” alone does not mean the price shown includes residency. Shared links save your inputs and filters, but prices can change.
Self-hosted options
Operating cost not estimated · servers, licenses and engineering time are yours. We list these separately because there’s no calculated bill to compare with managed services.
- Docling (self-hosted) — Open-source (MIT) document converter with layout analysis, reading order and table-structure models; runs locally on CPU or GPU.
- olmOCR (self-hosted) — Allen AI's open (Apache 2.0) 7B vision-language OCR model and pipeline converting PDFs and scans to markdown with tables and equations.
No self-hosted options match these filters.
Head-to-head comparisons
OCR & document parsing pricing explained
What drives the bill
Document APIs charge per page, with different rates for text, tables, layout and custom forms. Vision-language models can extract text or Markdown too, but charge by tokens. Dense pages and long outputs cost more. With open-source tools such as Docling and olmOCR, you pay for hardware, storage and the work of running the pipeline.
Costs the calculator leaves out
- Custom-trained extractors for specific forms (invoices, IDs, receipts), which are priced separately.
- Handwriting, low-quality scans and complex tables, where accuracy varies most between approaches.
- Output tokens. For LLM-based OCR, long dense pages produce long outputs, and that drives cost.
- Data handling. Documents often contain personal data, so check where they are processed and stored.
When paying more makes sense
Confidence scores, bounding boxes and structured fields can help when you need to audit extractions. Check which API and model return them. For search or RAG, a vision model or open-source pipeline may be enough, but test scans, tables and difficult pages from your own documents before switching.