Mistral OCR 4 Launches With Bounding Boxes, 170 Language Support at $2-4 Per 1,000 Pages
Mistral AI released OCR 4, a compact document extraction model that returns bounding boxes, block classification, and inline confidence scores alongside text. The model supports 170 languages, scores 85.20 on OlmOCRBench, and is priced at $4 per 1,000 pages via API ($2 with batch discount) or $5 per 1,000 pages through Document AI.
Mistral OCR 4 Launches With Bounding Boxes, 170 Language Support at $2-4 Per 1,000 Pages
Mistral AI released OCR 4, a document extraction model that adds bounding boxes, block classification, and inline confidence scores to extracted text. The model runs in a single container for self-hosted deployment and supports 170 languages across 10 language groups.
Pricing and Deployment
OCR 4 is priced at $4 per 1,000 pages via API, with a 50% batch discount reducing the cost to $2 per 1,000 pages. The Document AI interface in Mistral Studio costs $5 per 1,000 pages. The model is compact enough to deploy on a single container for organizations with data sovereignty requirements.
Performance Claims
Mistral claims OCR 4 achieves an 85.20 score on OlmOCRBench, the top result among models tested by the company. The model also scores 93.07 on OmniDocBench and 0.98 on Mistral's internal Crawl Multilingual evaluation.
In human preference evaluations, independent annotators preferred OCR 4 over competing systems with an average win rate of 72%, according to Mistral. The company tested OCR 4 against 600+ documents across 12+ languages in blind comparisons.
Mistral notes significant limitations in automated benchmark scoring, including ground-truth errors in reference data, equivalent LaTeX notation scored as mismatches, and multi-column reading order artifacts. The company recommends evaluating the model on your own documents rather than relying solely on benchmark scores.
Technical Capabilities
OCR 4 returns structured document representations with:
- Bounding boxes for text localization and in-context highlighting
- Block classification identifying titles, tables, equations, signatures, and other document elements
- Inline confidence scores per-page and per-word for verification workflows
- Format support for PDF, DOC, PPT, and OpenDocument files
- 170 languages across English, Western Europe, Eastern Europe, Middle Eastern, Chinese, East Asian, Southeast Asian, and rare language groups
The model integrates with Mistral Search Toolkit, an open-source search framework announced at the AI Now Summit, providing structured inputs for RAG and enterprise search pipelines.
Use Cases and Performance Data
Aidan Donohue, AI Engineer at Rogo, reported OCR 4 matched the accuracy of leading agentic document parsers on financial QA datasets at roughly 8x lower cost and 17x lower latency. Ivan Mihailov, AI engineer at Anaqua, stated the model is roughly 4x faster per page than their previous provider for high-volume docketing workflows.
What This Means
OCR 4 addresses a critical gap in document processing by combining text extraction with spatial and structural metadata. Bounding boxes enable citation-grounded outputs and data pipeline validation, while block classification supports semantic chunking for retrieval systems. The $2-4 per 1,000 page pricing undercuts many enterprise document services, though organizations should verify performance on their specific document types given the benchmark limitations Mistral acknowledges. The single-container deployment option makes this model accessible to organizations that cannot send documents to external APIs for compliance or sovereignty reasons.
Related Articles
Xiaomi Releases MiMo-V2.6-Pro-RL, a 1.02T-Parameter Omnimodal Model with 1M-Token Context
Xiaomi's MiMo team has released MiMo-V2.6-Pro-RL, a 1.02-trillion-parameter sparse mixture-of-experts model with 42B active parameters, 1M-token context, and native text/image/video/audio processing. The model was trained via a single mixed reinforcement learning run spanning coding, agentic, visual, and cybersecurity tasks, with benchmark scores that Xiaomi claims approach or match Claude Opus 5 and GPT-5.6 on several agentic and coding tests.
Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
Xiaomi Releases MiMo-V2.6-Flash: Open-Source MoE Model with 1M-Token Context, $0.14/$0.28 per 1M Tokens
Xiaomi has released MiMo-V2.6-Flash, an open-source Mixture-of-Experts model with 309B total parameters and 15B activated per token, featuring a 1M-token context window and native multimodal capabilities. Priced at $0.14 per 1M input tokens and $0.28 per 1M output tokens, it targets agentic coding and long-horizon task workflows.
Xiaomi Launches MiMo-V2.6-Pro, a 1T+ Parameter Model with 1M-Token Context
Xiaomi has released MiMo-V2.6-Pro, a flagship foundation model exceeding 1 trillion parameters with a 1M-token context window and native multimodal support. The model is priced at $0.435 per 1M input tokens and $0.87 per 1M output tokens, targeting agentic and long-horizon tasks.
Comments
Loading...