DeepSeek Releases Experimental V4-Flash-Vision-Exp, Claims Near-Parity With Opus 4.8 on Agent Benchmarks
DeepSeek has released V4-Flash-Vision-Exp, an experimental multimodal extension of V4-Flash that adds image understanding while preserving text reasoning capabilities. The company claims the model approaches or beats Anthropic's Opus 4.8 on its internal multimodal agent benchmarks.
DeepSeek has released Deepseek-V4-Flash-Vision-Exp, an experimental multimodal model that extends its V4-Flash text model with image understanding. According to DeepSeek, the model nearly matches Anthropic's Opus 4.8 on the company's internal multimodal agent benchmarks — a claim that has not been independently verified.
What's new
V4-Flash-Vision-Exp adds image processing to the existing V4-Flash architecture while, DeepSeek says, preserving the base model's text reasoning and world-knowledge performance. The model is built for visual agent workflows: it can describe images, extract text from screenshots, and analyze diagrams, and DeepSeek says it's designed to work with tool-use across different agent frameworks.
The model supports JPEG, PNG, GIF, and WebP formats. Notably, it determines image format by reading actual file content rather than relying on filenames or declared MIME types, according to DeepSeek's API documentation.
DeepSeek also shipped version 0.1.1 of its Harness framework alongside the model, with built-in support for the new vision capabilities. The model is accessible through OpenAI's Chat Completions and Responses APIs as well as Anthropic's Messages endpoint, easing integration for developers already building on those ecosystems.
Image handling and limits
Developers have three ways to send images: direct Base64 encoding, publicly accessible URLs (up to 32 MiB), or DeepSeek's new Files API, which is free to use. The Files API allows a single upload to be referenced by ID across multiple requests, with a 64 MiB size cap per file.
An optional "detail" field can downscale images to 512 x 512 pixels to save tokens when fine visual detail isn't required. By default, the model normalizes images to roughly 800 x 800 pixels based on aspect ratio before processing. Regardless of original resolution, DeepSeek says each image costs no more than 384 tokens — a fixed, predictable cost structure for image-heavy workloads.
A single request can include up to 600 images. Maximum edge length is 8,192 pixels per side, though that drops to 4,096 pixels once a request includes 15 or more images. Images are restricted to user messages only.
Pricing
DeepSeek says pricing for V4-Flash-Vision-Exp follows existing V4-Flash rates, meaning no separate surcharge for the vision capability itself. Exact per-token input and output pricing for V4-Flash was not specified in DeepSeek's release materials reviewed here; pricing not yet disclosed in this report.
What this means
DeepSeek's move to fold vision into its Flash line — rather than shipping vision as a separate, premium model — signals a push toward multimodal agents as a default feature rather than an add-on. The fixed 384-token-per-image cost and free Files API suggest DeepSeek is optimizing for high-volume, image-heavy agent workloads like screenshot parsing and UI automation, where cost predictability matters more than peak resolution.
The comparison to Opus 4.8 comes from DeepSeek's own internal benchmarks, not third-party evaluation, so the "rivals Opus 4.8" claim should be treated as a company assertion until independent testing confirms it. Given the "Exp" label, this is explicitly a preview rather than a production-ready release, and DeepSeek may adjust capabilities, limits, or pricing before a stable version ships.
Related Articles
Reka AI releases Rho-1, a 19B-parameter omni-model for text, image, video and robot control
Reka AI has released a research preview of Rho-1, a 19-billion-parameter omni-model that processes and generates text, images, video, and robot control actions in a single network. Reka says it uses no tool calls or external models. Context window, pricing, and benchmark scores have not been disclosed.
Cloudflare releases Clef decision models, claims 39 ms median latency vs. 524 ms for TypeSafe's Jev
Cloudflare has released Clef and Clef-flash, two open-weight decision models that return probabilities over predefined answer options instead of generating text. The company claims median latencies of 39 ms and 209 ms, against just over 524 ms for TypeSafe AI's Jev. Both support text and images and are API-compatible with Jev.
Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text
Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.
Amazon open-sources Strands Decider 2B, a small decision model built on a Qwen3.5-2B base
Amazon Web Services has released Strands Decider 2B, an open-source model that chooses among pre-decided options and returns a confidence score instead of generating text. It is inspired by TypeSafe's Jev and is small enough to run locally. Amazon says it briefly topped the Jevbench ranking for models of its size.
Comments
Loading...