olmOCR (self-hosted)
Allen AI's open (Apache 2.0) 7B vision-language OCR model and pipeline converting PDFs and scans to markdown with tables and equations.
olmOCR needs an NVIDIA GPU with at least 12 GB of VRAM. Ai2 estimates rented-GPU compute below $200 per million pages, but that excludes storage and the work of running the pipeline. Unlike Docling's CPU path, this 7B vision-language model requires a GPU. Hosted endpoints such as DeepInfra are available if you don't want to operate one.
- Headquarters
- 🇺🇸 United States
- Parent jurisdiction
- 🇺🇸 United States (parent: Allen Institute for AI (Ai2))
- EU self-serve plan
- Available; check eligible plans and regions
- GDPR DPA
- No / not documented
- Billing
- self-hosted
- Free tier
- Free and open weights.
- Open source
- Yes · self-hostable
- Website
- github.com
- Source last checked
- 2026-10-06
- Pricing evidence
- Operating cost not estimated
Pricing
Operating cost not estimated · servers, licenses and engineering time are yours. See notes for software editions and license terms.
Self-hosted: budget for GPUs, storage and maintenance. Needs a recent NVIDIA GPU with at least 12 GB VRAM (tested on RTX 4090, L40S, A100, H100); model allenai/olmOCR-2-7B-1025-FP8 served with vLLM. Ai2 states under $200 per million pages converted on rented GPUs, i.e. roughly $0.20 per 1,000 pages: about 1-2 GPU-hours per 10,000 pages on an H100-class card. The README also lists hosted endpoints (DeepInfra $0.09/$0.19, Parasail $0.10/$0.20, Cirrascale $0.07/$0.15 per M input/output tokens), roughly $0.30 per 1,000 pages at ~1,500 image + 300 prompt tokens and 800 output tokens per page. Similar open OCR VLMs: DeepSeek-OCR, PaddleOCR-VL, dots.ocr, Granite-Docling.
Managed alternatives
At 100K pages / month. Adjust usage and requirements.
olmOCR has no bill to estimate: you pay for the servers, storage and time to run it. The cheapest managed option here is GPT-6 Luna (via OpenAI API) at $719/yr, so running olmOCR yourself saves money only if hosting and upkeep stay below about $60 a month.
- GPT-6 Luna (via OpenAI API): $719/yr estimatedEstimated ratesPage cost uses an assumed image-token multiplier plus 300 prompt and 800 output tokens. Confirm actual token usage.
- Mistral Small 3.2 (via Scaleway Generative APIs): $845/yr estimatedModeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary.
- Qwen3-VL-235B (via DeepInfra): $1,244/yr estimatedModeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary.
- Gemini 3.1 Flash-Lite (via Gemini API): $1,866/yr estimatedModeled ratesToken cost assumes a 1,000×1,400 px page, a 300-token prompt and 800 output tokens. Extraction quality and retries vary.
- LlamaParse (LlamaIndex): $4,500/yr estimatedPublished rates
No cheaper managed estimate matches your filters. Adjust your usage or compare with your actual contract bill.
The price difference leaves out migration costs, operating costs and any features you’d need to replace.
Self-hosted options
Operating cost is not estimated. These options are not counted as savings against managed services.
Other OCR & document parsing comparisons
- Amazon Textract vs Azure AI Document Intelligence (Layout)
- Amazon Textract vs Mistral OCR
- LlamaParse (LlamaIndex) vs Mistral OCR
Found a wrong rate or plan restriction? Report a correction with a current source.