Baidu Releases Qianfan-OCR-Fast Model with 66K Context at $0.68 Per 1M Input Tokens
Baidu has released Qianfan-OCR-Fast, a multimodal model specialized for optical character recognition tasks. The model offers a 66,000 token context window and is priced at $0.68 per 1M input tokens and $2.81 per 1M output tokens.
Qianfan-OCR-Fast — Quick Specs
Baidu Releases Qianfan-OCR-Fast Model with 66K Context at $0.68 Per 1M Input Tokens
Baidu has released Qianfan-OCR-Fast, a multimodal model purpose-built for optical character recognition, with a 66,000 token context window and pricing of $0.68 per 1M input tokens.
Specifications
The model is available through OpenRouter with the following specifications:
- Context window: 66,000 tokens
- Input pricing: $0.68 per 1M tokens
- Output pricing: $2.81 per 1M tokens
- Model type: Multimodal (specialized for OCR)
- Release date: Listed as April 20, 2026 (likely an error; actual release date unclear)
Technical Details
According to Baidu, Qianfan-OCR-Fast was trained on specialized OCR data while maintaining broader multimodal capabilities. The company claims it provides improved performance over its predecessor, Qianfan-OCR, though specific benchmark comparisons were not provided.
The model is designed to handle document understanding, text extraction, and related OCR tasks while retaining general multimodal intelligence for image understanding beyond pure text recognition.
Availability
Qianfan-OCR-Fast is currently available through OpenRouter's API routing service, which automatically selects providers based on prompt requirements and maintains fallback options for uptime. Weekly token usage on the platform stands at 273,000 tokens as of the listing date.
No information about direct API access through Baidu's own infrastructure was disclosed in the announcement.
What This Means
Baidu's OCR-specialized model enters a growing market for document understanding AI, competing with models from OpenAI (GPT-4 Vision), Anthropic (Claude 3), and Google (Gemini). The 66K context window is sufficient for processing lengthy documents in a single request, though it falls short of competitors offering 200K+ contexts. At $0.68 per 1M input tokens, pricing is competitive for specialized OCR tasks, particularly for high-volume document processing workflows where domain-specific optimization may justify the cost over general-purpose vision models.
Related Articles
ByteDance Seed Launches Seed 2.1 Turbo, a 262K-Context Multimodal Model for Coding Agents
ByteDance Seed has released Seed 2.1 Turbo, a multimodal model targeting coding and long-horizon agent workflows with a 262K token context window. The model is priced at $0.50 per 1M input tokens and $2.50 per 1M output tokens, and is now listed on OpenRouter.
NVIDIA Releases Nemotron 3.5 Lightning: 30B MoE Model with 1M Token Context and 3B Active Parameters
NVIDIA released the full-precision BF16 reference weights for Nemotron 3.5 Lightning, a 30B-parameter Mixture-of-Experts model with only 3B active parameters and support for up to 1 million tokens of context. The model uses a hybrid Mamba-2, MoE, and Attention architecture and is licensed under OpenMDW-1.1 for commercial use.
Qwen Releases Qwen3.8 2.4T A95B, a 2.4-Trillion-Parameter Open-Weight MoE Model
Qwen has released Qwen3.8 2.4T A95B, an open-weight sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion active parameters per forward pass. The model is the open-weight variant of Qwen3.8 Max, targeting coding, research, complex reasoning, and agentic workflows with a 262K token context window.
DeepSeek Releases V4 Pro 0813 With 1.05M Token Context Window, Priced at $0.43/M Input
DeepSeek has shipped the general availability release of DeepSeek V4 Pro, codenamed 0813, featuring a 1,049,000-token context window. The mixture-of-experts model is priced at $0.43 per million input tokens and $0.87 per million output tokens, and is live now on OpenRouter.
Comments
Loading...