Mistral Launches OCR API at $1 Per 1,000 Pages, Claims 94.89% Accuracy on Document Benchmarks
Mistral AI has released Mistral OCR, an API for extracting text and images from documents at $1 per 1,000 pages (approximately $0.50 with batch inference). The company claims 94.89% overall accuracy on its internal test set, comparing favorably to GPT-4o (89.77%), Gemini 2.0 Flash (88.69%), and Azure OCR (89.52%).
Mistral Launches OCR API at $1 Per 1,000 Pages, Claims 94.89% Accuracy on Document Benchmarks
Mistral AI has released Mistral OCR, an API for extracting text and images from documents at $1 per 1,000 pages (approximately $0.50 with batch inference). The company claims 94.89% overall accuracy on its internal test set, comparing favorably to GPT-4o (89.77%), Gemini 2.0 Flash (88.69%), and Azure OCR (89.52%).
The API accepts images and PDFs as input and outputs interleaved text and images in markdown format. Mistral has deployed the model as the default document understanding system on Le Chat, its chatbot platform.
Performance Claims
According to Mistral, the model achieved the following scores on its internal "text-only" test set:
- Overall accuracy: 94.89% (vs GPT-4o's 89.77%)
- Math extraction: 94.29% (vs GPT-4o's 87.55%)
- Scanned documents: 98.96% (vs GPT-4o's 94.58%)
- Tables: 96.12% (vs GPT-4o's 91.70%)
- Multilingual: 89.55% (vs GPT-4o's 86.00%)
The company claims processing speeds of up to 2,000 pages per minute on a single node. Mistral states it extracts embedded images from documents alongside text, a capability not present in the compared models.
Multilingual Support
Mistral claims 99.02% fuzzy match accuracy across multiple languages on its benchmarks, compared to 96.53% for Gemini 2.0 Flash and 97.31% for Azure OCR. The company reports accuracy above 97% for 11 tested languages, including Russian (99.09%), German (99.51%), Spanish (99.54%), Chinese (97.11%), and Hindi (97.55%).
Technical Capabilities
The model handles:
- Mathematical expressions and LaTeX formatting
- Complex tables and interleaved imagery
- Documents as prompts with structured JSON output
- Multiple scripts and fonts across languages
Mistral positions the API for use in RAG (Retrieval-Augmented Generation) systems processing multimodal documents like slides and complex PDFs. Users can chain extracted outputs into downstream function calls for agent-based workflows.
Availability and Deployment
The API is available today on la Plateforme, Mistral's developer platform. The company plans to extend availability to cloud and inference partners, plus on-premises deployment on a selective basis for organizations handling classified information.
Mistral has not disclosed the model's parameter count, architecture details, or training data composition.
What This Means
Mistral OCR enters a competitive market dominated by Google Document AI, Azure OCR, and general-purpose multimodal models like GPT-4o and Gemini. The pricing of $1 per 1,000 pages undercuts typical enterprise OCR pricing, though direct cost comparisons depend on specific use cases and batch processing capabilities. The claimed accuracy advantages—particularly on mathematical content (94.29% vs GPT-4o's 87.55%)—could make it viable for scientific and technical document processing if the benchmarks prove reproducible on external test sets. The key differentiation appears to be simultaneous text and image extraction in a single pass, which existing general-purpose LLMs don't natively support.
Related Articles
ElevenLabs Launches Music v2.5, Adds API Access and Free Tier for AI-Generated Songs
ElevenLabs has released Music v2.5, an updated version of its ElevenMusic generator, now available through both the app and API. The company says blind testing with nearly 48,000 comparison pairs showed listeners preferred v2.5 over the prior version, particularly for R&B, Hip-Hop, and orchestral genres.
Mistral AI Models Now Power Mozilla's Firefox Smart Window Browsing Assistant
Mistral AI will power Mozilla's Firefox Smart Window (beta), a browsing assistant that summarizes searches and tracks tabs. The rollout starts in France and North America, with the UK and Germany to follow later this year.
Meta Launches WhatsApp Business Tools MCP to Let AI Agents Automate Business Setup
Meta released a new MCP server that connects AI coding agents directly to the WhatsApp Business Platform, automating account creation, phone verification, and messaging template setup. The move expands Meta's existing lineup of MCP servers beyond ad management and app monitoring tools.
AWS Details How Amazon Bedrock Prompt Caching Cuts Input Token Costs by Up to 90%
Amazon Bedrock's prompt caching feature can cut input token costs by up to 90% on cache hits by storing repeated context like documents, system prompts, and tool definitions. AWS outlines six implementation patterns and pricing details, including a 25% premium for cache writes and 90% discount on cache reads.
Comments
Loading...