product updateMistral AI

Mistral Launches OCR API at $1 Per 1,000 Pages, Claims 94.89% Accuracy on Document Benchmarks

TL;DR

Mistral AI has released Mistral OCR, an API for extracting text and images from documents at $1 per 1,000 pages (approximately $0.50 with batch inference). The company claims 94.89% overall accuracy on its internal test set, comparing favorably to GPT-4o (89.77%), Gemini 2.0 Flash (88.69%), and Azure OCR (89.52%).

2 min read
0

Mistral Launches OCR API at $1 Per 1,000 Pages, Claims 94.89% Accuracy on Document Benchmarks

Mistral AI has released Mistral OCR, an API for extracting text and images from documents at $1 per 1,000 pages (approximately $0.50 with batch inference). The company claims 94.89% overall accuracy on its internal test set, comparing favorably to GPT-4o (89.77%), Gemini 2.0 Flash (88.69%), and Azure OCR (89.52%).

The API accepts images and PDFs as input and outputs interleaved text and images in markdown format. Mistral has deployed the model as the default document understanding system on Le Chat, its chatbot platform.

Performance Claims

According to Mistral, the model achieved the following scores on its internal "text-only" test set:

  • Overall accuracy: 94.89% (vs GPT-4o's 89.77%)
  • Math extraction: 94.29% (vs GPT-4o's 87.55%)
  • Scanned documents: 98.96% (vs GPT-4o's 94.58%)
  • Tables: 96.12% (vs GPT-4o's 91.70%)
  • Multilingual: 89.55% (vs GPT-4o's 86.00%)

The company claims processing speeds of up to 2,000 pages per minute on a single node. Mistral states it extracts embedded images from documents alongside text, a capability not present in the compared models.

Multilingual Support

Mistral claims 99.02% fuzzy match accuracy across multiple languages on its benchmarks, compared to 96.53% for Gemini 2.0 Flash and 97.31% for Azure OCR. The company reports accuracy above 97% for 11 tested languages, including Russian (99.09%), German (99.51%), Spanish (99.54%), Chinese (97.11%), and Hindi (97.55%).

Technical Capabilities

The model handles:

  • Mathematical expressions and LaTeX formatting
  • Complex tables and interleaved imagery
  • Documents as prompts with structured JSON output
  • Multiple scripts and fonts across languages

Mistral positions the API for use in RAG (Retrieval-Augmented Generation) systems processing multimodal documents like slides and complex PDFs. Users can chain extracted outputs into downstream function calls for agent-based workflows.

Availability and Deployment

The API is available today on la Plateforme, Mistral's developer platform. The company plans to extend availability to cloud and inference partners, plus on-premises deployment on a selective basis for organizations handling classified information.

Mistral has not disclosed the model's parameter count, architecture details, or training data composition.

What This Means

Mistral OCR enters a competitive market dominated by Google Document AI, Azure OCR, and general-purpose multimodal models like GPT-4o and Gemini. The pricing of $1 per 1,000 pages undercuts typical enterprise OCR pricing, though direct cost comparisons depend on specific use cases and batch processing capabilities. The claimed accuracy advantages—particularly on mathematical content (94.29% vs GPT-4o's 87.55%)—could make it viable for scientific and technical document processing if the benchmarks prove reproducible on external test sets. The key differentiation appears to be simultaneous text and image extraction in a single pass, which existing general-purpose LLMs don't natively support.

Related Articles

product update

OpenRouter Launches Auto Router Beta: Task-Aware Model Routing Based on Community Spend

OpenRouter has released Auto Router Beta, a task-aware routing system that classifies incoming requests and automatically routes them to popular models based on community spending patterns. The router allows users to filter selections by cost-quality tradeoff preferences.

product update

LM Studio launches Bionic, agentic app for local and cloud open-source models

LM Studio released Bionic, a Mac app that runs open-source AI models locally or via cloud for coding, document processing, and research tasks. The app includes offline voice transcription using Mistral's Voxtral model and supports models like GLM 5.2 and Kimi K2.7 Code for codebase editing.

product update

OpenAI restores chat sidebar in Mac app after user backlash over confusing redesign

OpenAI has updated its ChatGPT Mac app to restore direct access to chat conversations through a prominent sidebar toggle. The fix addresses user complaints following a July 10 redesign that replaced the native Mac client with an Electron-based app and buried the standard chat interface behind Work and Codex features.

product update

NVIDIA NeMo Automodel integrates with Hugging Face Diffusers for distributed video and image model fine-tuning

NVIDIA and Hugging Face have integrated NeMo Automodel with the Diffusers library, enabling distributed fine-tuning of video and image diffusion models without checkpoint conversion. The integration supports models including FLUX.1-dev (12B), Wan 2.1 (1.3B/14B), and HunyuanVideo (13B) with full fine-tuning and LoRA options.

Comments

Loading...