Mistral OCR 4 Launches With Bounding Boxes, 170 Language Support at $2-4 Per 1,000 Pages
Mistral AI released OCR 4, a compact document extraction model that returns bounding boxes, block classification, and inline confidence scores alongside text. The model supports 170 languages, scores 85.20 on OlmOCRBench, and is priced at $4 per 1,000 pages via API ($2 with batch discount) or $5 per 1,000 pages through Document AI.
Mistral OCR 4 Launches With Bounding Boxes, 170 Language Support at $2-4 Per 1,000 Pages
Mistral AI released OCR 4, a document extraction model that adds bounding boxes, block classification, and inline confidence scores to extracted text. The model runs in a single container for self-hosted deployment and supports 170 languages across 10 language groups.
Pricing and Deployment
OCR 4 is priced at $4 per 1,000 pages via API, with a 50% batch discount reducing the cost to $2 per 1,000 pages. The Document AI interface in Mistral Studio costs $5 per 1,000 pages. The model is compact enough to deploy on a single container for organizations with data sovereignty requirements.
Performance Claims
Mistral claims OCR 4 achieves an 85.20 score on OlmOCRBench, the top result among models tested by the company. The model also scores 93.07 on OmniDocBench and 0.98 on Mistral's internal Crawl Multilingual evaluation.
In human preference evaluations, independent annotators preferred OCR 4 over competing systems with an average win rate of 72%, according to Mistral. The company tested OCR 4 against 600+ documents across 12+ languages in blind comparisons.
Mistral notes significant limitations in automated benchmark scoring, including ground-truth errors in reference data, equivalent LaTeX notation scored as mismatches, and multi-column reading order artifacts. The company recommends evaluating the model on your own documents rather than relying solely on benchmark scores.
Technical Capabilities
OCR 4 returns structured document representations with:
- Bounding boxes for text localization and in-context highlighting
- Block classification identifying titles, tables, equations, signatures, and other document elements
- Inline confidence scores per-page and per-word for verification workflows
- Format support for PDF, DOC, PPT, and OpenDocument files
- 170 languages across English, Western Europe, Eastern Europe, Middle Eastern, Chinese, East Asian, Southeast Asian, and rare language groups
The model integrates with Mistral Search Toolkit, an open-source search framework announced at the AI Now Summit, providing structured inputs for RAG and enterprise search pipelines.
Use Cases and Performance Data
Aidan Donohue, AI Engineer at Rogo, reported OCR 4 matched the accuracy of leading agentic document parsers on financial QA datasets at roughly 8x lower cost and 17x lower latency. Ivan Mihailov, AI engineer at Anaqua, stated the model is roughly 4x faster per page than their previous provider for high-volume docketing workflows.
What This Means
OCR 4 addresses a critical gap in document processing by combining text extraction with spatial and structural metadata. Bounding boxes enable citation-grounded outputs and data pipeline validation, while block classification supports semantic chunking for retrieval systems. The $2-4 per 1,000 page pricing undercuts many enterprise document services, though organizations should verify performance on their specific document types given the benchmark limitations Mistral acknowledges. The single-container deployment option makes this model accessible to organizations that cannot send documents to external APIs for compliance or sovereignty reasons.
Related Articles
Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window
Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.
Meta Releases Muse Spark 1.3, a Free Multimodal Reasoning Model with 1M-Token Context
Meta has released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window, listed as free on OpenRouter. The model targets long-running agentic, multi-agent, and coding workflows, though audio input support remains incomplete.
OpenAI's GPT-6 Astra Cuts Hallucinations, But Indirect Prompt Injection Attacks Still Succeed 8.5% of the Time
OpenAI's new GPT-6 Astra model shows major improvements in hallucination rates and jailbreak resistance over predecessor GPT-5.6 Sol, according to OpenAI's system card. However, indirect prompt injection attacks hidden in documents still succeed 8.5% of the time in external testing by Gray Swan, down from 27% but still above rival Claude Opus 5's 4.8% rate.
OpenAI Ships GPT-6 Astra, But Executives Admit They Can't Fully Monitor What It's Thinking
OpenAI released GPT-6 Astra on Thursday, a model president Greg Brockman says could mark the start of AGI. But the model writes out its reasoning less often than prior versions, and OpenAI's chief scientist says monitoring AI thought processes will keep getting harder.
Comments
Loading...