Baidu Releases Qianfan-OCR-Fast Model with 66K Context at $0.68 Per 1M Input Tokens
Baidu has released Qianfan-OCR-Fast, a multimodal model specialized for optical character recognition tasks. The model offers a 66,000 token context window and is priced at $0.68 per 1M input tokens and $2.81 per 1M output tokens.
Qianfan-OCR-Fast — Quick Specs
Baidu Releases Qianfan-OCR-Fast Model with 66K Context at $0.68 Per 1M Input Tokens
Baidu has released Qianfan-OCR-Fast, a multimodal model purpose-built for optical character recognition, with a 66,000 token context window and pricing of $0.68 per 1M input tokens.
Specifications
The model is available through OpenRouter with the following specifications:
- Context window: 66,000 tokens
- Input pricing: $0.68 per 1M tokens
- Output pricing: $2.81 per 1M tokens
- Model type: Multimodal (specialized for OCR)
- Release date: Listed as April 20, 2026 (likely an error; actual release date unclear)
Technical Details
According to Baidu, Qianfan-OCR-Fast was trained on specialized OCR data while maintaining broader multimodal capabilities. The company claims it provides improved performance over its predecessor, Qianfan-OCR, though specific benchmark comparisons were not provided.
The model is designed to handle document understanding, text extraction, and related OCR tasks while retaining general multimodal intelligence for image understanding beyond pure text recognition.
Availability
Qianfan-OCR-Fast is currently available through OpenRouter's API routing service, which automatically selects providers based on prompt requirements and maintains fallback options for uptime. Weekly token usage on the platform stands at 273,000 tokens as of the listing date.
No information about direct API access through Baidu's own infrastructure was disclosed in the announcement.
What This Means
Baidu's OCR-specialized model enters a growing market for document understanding AI, competing with models from OpenAI (GPT-4 Vision), Anthropic (Claude 3), and Google (Gemini). The 66K context window is sufficient for processing lengthy documents in a single request, though it falls short of competitors offering 200K+ contexts. At $0.68 per 1M input tokens, pricing is competitive for specialized OCR tasks, particularly for high-volume document processing workflows where domain-specific optimization may justify the cost over general-purpose vision models.
Related Articles
Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier
Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.
Mistral Releases Shieldstral, a 3B Open-Weights Safety Classifier That Matches Models 7x Its Size
Mistral has released Shieldstral, a 3B open-weights safety classifier that reframes content moderation as a policy-adaptive question-answering task. The model claims to match or outperform guard models up to 7x its size on text safety and multimodal benchmarks, and runs on a single 16GB GPU.
Alibaba Releases Qwen3.8 Max, a Multimodal Reasoning Model with 1M Token Context
Alibaba has moved Qwen3.8 Max out of preview into general availability, positioning it as the flagship of the Qwen3.8 series with a 1 million token context window and multimodal input support. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens via OpenRouter.
OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold
OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.
Comments
Loading...