Docling (self-hosted)
Open-source (MIT) document converter with layout analysis, reading order and table-structure models; runs locally on CPU or GPU.
Docling runs locally on CPU or GPU, with no per-page fee. A rough sizing estimate puts 100,000 pages a month on an 8–16-vCPU VM or a small GPU, but scans needing OCR are much slower. Test your documents before buying hardware, and budget for storage and maintenance as well as compute. The project started at IBM Research and is now part of the LF AI & Data Foundation.
- Headquarters
- 🇺🇸 United States
- Parent jurisdiction
- 🇺🇸 United States (parent: LF AI & Data Foundation (started at IBM Research))
- EU self-serve plan
- Available; check eligible plans and regions
- GDPR DPA
- No / not documented
- Billing
- self-hosted
- Free tier
- Free and open source.
- Open source
- Yes · self-hostable
- Website
- github.com
- Source last checked
- 2026-10-06
- Pricing evidence
- Operating cost not estimated
Pricing
Operating cost not estimated · servers, licenses and engineering time are yours. See notes for software editions and license terms.
Self-hosted: budget for hardware, storage and maintenance. Uses layout and TableFormer models plus pluggable OCR engines (EasyOCR, Tesseract, RapidOCR) for scans; the optional Granite-Docling VLM pipeline is slower but handles harder pages. Rough estimate: about 1-3 s per page per CPU worker on born-digital PDFs and several pages/s on one GPU, so 100k pages/month fits on a single 8-16 vCPU VM or a small GPU (e.g. L4); scanned pages needing OCR are several times slower. Data stays wherever you run it.
Managed alternatives
At 100K pages / month. Adjust usage and requirements.
Docling has no bill to estimate: you pay for the servers, storage and time to run it. The cheapest managed option here is GPT-6 Luna (via OpenAI API) at $719/yr, so running Docling yourself saves money only if hosting and upkeep stay below about $60 a month.
- GPT-6 Luna (via OpenAI API): $719/yr estimatedEstimated ratesPage cost uses an assumed image-token multiplier plus 300 prompt and 800 output tokens. Confirm actual token usage.
- Mistral Small 3.2 (via Scaleway Generative APIs): $845/yr estimatedModeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary.
- Qwen3-VL-235B (via DeepInfra): $1,244/yr estimatedModeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary.
- Gemini 3.1 Flash-Lite (via Gemini API): $1,866/yr estimatedModeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary.
- LlamaParse (LlamaIndex): $4,500/yr estimatedPublished rates
No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.
The price difference leaves out migration costs, operating costs and any features you’d need to replace.
Self-hosted options
Operating cost is not estimated. These options are not counted as savings against managed services.
Other OCR & document parsing comparisons
- Amazon Textract vs Azure AI Document Intelligence (Layout)
- Amazon Textract vs Mistral OCR
- LlamaParse (LlamaIndex) vs Mistral OCR
Found a wrong rate or plan restriction? Report a correction with a current source.