Mistral OCR 3 launches at $2 per 1,000 pages with 74% win rate over previous version
Mistral AI released OCR 3, a document parsing model priced at $2 per 1,000 pages with a 50% batch API discount. The company claims a 74% overall win rate compared to Mistral OCR 2 on forms, scanned documents, complex tables, and handwriting.
Mistral OCR 3 launches at $2 per 1,000 pages with 74% win rate over previous version
Mistral AI released OCR 3 (model ID: mistral-ocr-2512), a document parsing model priced at $2 per 1,000 pages, or $1 per 1,000 pages using the batch API. The company claims a 74% overall win rate compared to its previous OCR 2 model across forms, scanned documents, complex tables, and handwriting.
Pricing and availability
- Standard API: $2 per 1,000 pages
- Batch API: $1 per 1,000 pages (50% discount)
- Available now via API and Document AI Playground in Mistral AI Studio
- Self-hosting option available for organizations with data privacy requirements
- Fully backward compatible with Mistral OCR 2
Technical capabilities
Mistral OCR 3 extracts text and embedded images from documents, outputting markdown with HTML-based table reconstruction. The model handles:
- Handwriting: Cursive text, mixed-content annotations, and handwritten entries on printed forms
- Forms: Invoice processing, receipts, compliance forms, and government documents with improved box and label detection
- Low-quality scans: Handles compression artifacts, skew, distortion, low DPI, and background noise
- Complex tables: Reconstructs structures with headers, merged cells, multi-row blocks, and column hierarchies using HTML tags with colspan/rowspan attributes
Benchmarks
Mistral AI evaluated OCR 3 on internal benchmarks based on customer use cases, comparing outputs to ground truth using fuzzy-match metrics. The company claims the model outperforms "enterprise document processing solutions as well as AI-native OCR solutions," though specific competitor comparisons and benchmark scores were not disclosed.
According to Mistral AI, OCR 3 represents "a significant upgrade across all languages and document form factors" compared to OCR 2. The company describes it as "a much smaller model than most competitive solutions."
Customer applications
Early customers are using Mistral OCR 3 for:
- Invoice processing into structured fields
- Company archive digitization
- Clean text extraction from technical and scientific reports
- Enterprise search enhancement
- Document-to-knowledge transformation pipelines
What this means
Mistral's aggressive pricing at $1-2 per 1,000 pages positions OCR 3 as a cost-competitive alternative to existing document processing services. The 74% win rate claim suggests substantial improvements over the previous generation, though the lack of third-party benchmarks or specific competitor comparisons makes independent verification difficult. The model's smaller size, if accurate, could enable faster processing and lower compute costs for high-volume document workflows. Self-hosting availability addresses enterprise compliance requirements that often block cloud-based document processing adoption.
Related Articles
Google Launches Gemini 3.5 Transcribe, a Speech-to-Text Model That Cleans Up Rambling Speech
Google has released Gemini 3.5 Transcribe, a new speech-to-text model that automatically detects over 85 languages, removes filler words, and structures unstructured speech into clean text. The model powers Android's Rambler feature and is rolling out to Chrome, Docs, Gmail, and other Google products.
Google DeepMind Launches Gemini 3.5 Transcribe, Claims 2.6% Word Error Rate in Testing
Google DeepMind has released Gemini 3.5 Transcribe, a speech-to-text model available via two APIs for real-time streaming and pre-recorded audio. According to Artificial Analysis benchmarks cited by Google, the model achieves a 2.6% word error rate for non-streaming transcription and 4.0% for streaming.
Google Launches Gemini 3.5 Transcribe with 2.6% Word Error Rate, Powers Gboard Rambler
Google has released Gemini 3.5 Transcribe, a speech-to-text model claiming a 4.0% word error rate in streaming mode and 2.6% in non-streaming mode, according to benchmarks from Artificial Analysis. The model already powers Gboard Rambler on Android and the Gemini app for macOS, with Chrome support coming next.
Alibaba Releases Qwen3.8-Flash-Next: 125B-Parameter MoE Model Matches Larger Rivals at $0.16/$0.47 per Million Tokens
Alibaba's Qwen team released Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model that activates just 6 billion parameters per token and previews architecture planned for Qwen4. The model outperforms the much larger Qwen3.7-Plus at roughly one-ninth the training cost and ships at $0.16 per million input tokens and $0.47 per million output tokens.
Comments
Loading...