DeepSeek releases V4 preview, claims parity with GPT-4o and Claude 3.5 Sonnet
DeepSeek released a preview of its V4 model on April 24, 2026, claiming the open-source system matches leading closed-source models from Anthropic, Google, and OpenAI. The company emphasized improved coding capabilities and compatibility with domestic Huawei chips, but did not disclose training costs or hardware specifications.
DeepSeek V4 Preview Released with Claimed Parity to Leading US Models
Chinese AI company DeepSeek released a preview of its V4 model on April 24, 2026, claiming the open-source system can compete with leading closed-source models from Anthropic, Google, and OpenAI.
Key Details
DeepSeek did not disclose:
- Training costs for V4
- Hardware used for training
- Specific benchmark scores
- Context window size
- Pricing information
The company emphasized that V4 represents a "major improvement" over prior models, particularly in coding capabilities. DeepSeek explicitly highlighted compatibility with domestic Huawei technology, marking what the company describes as a milestone for China's chip industry.
Context and Controversy
The V4 preview arrives approximately one year after DeepSeek's R1 model disrupted the US AI industry. DeepSeek claimed R1 was trained at a fraction of the cost of leading American systems, though specific cost figures were never independently verified.
US officials have accused DeepSeek of using banned Nvidia chips for training. Separately, Anthropic has claimed DeepSeek misused Claude to improve its own products, though details of these allegations remain unclear.
Coding Focus
According to DeepSeek, V4's enhanced coding performance targets capabilities that have become central to AI agents and driven adoption of systems like ChatGPT Codex and Claude Code. The company did not provide specific benchmarks comparing V4's coding performance to competitors.
What This Means
DeepSeek's V4 preview continues the pattern established with R1: bold claims about competitive performance without disclosed training costs, hardware specifications, or independent benchmark verification. The emphasis on Huawei chip compatibility suggests China's AI industry is working to reduce dependence on restricted Western semiconductor technology, though the practical performance implications remain unclear. Until DeepSeek releases concrete benchmarks and technical details, the actual capabilities of V4 relative to GPT-4o, Gemini, and Claude 3.5 Sonnet cannot be independently assessed.
Related Articles
Microsoft Releases Fara1.5-27B, a 27B Vision-Only Web Browsing Agent with 262K Context
Microsoft Research AI Frontiers has released Fara1.5-27B, a 27-billion-parameter multimodal agent that completes web tasks by reading screenshots and emitting click/type/scroll commands. The model, fine-tuned from Qwen3.5-27B, ships under MIT license with a 262K-token context window and is designed to run alongside Microsoft's MagenticLite sandbox.
Anthropic's Claude Opus 5 Hits 0% Prompt Injection Success Rate in Browser Agent Tests, With Defenses Enabled
Anthropic's system card for Claude Opus 5 reports a 0% prompt injection success rate across 129 browser agent test scenarios when Auto Mode is enabled. On Gray Swan's broader indirect prompt injection benchmark, Opus 5 posted a 2.0% attacker success rate after 15 attempts, the lowest among tested frontier models.
Anthropic Ships Claude Opus 5, Claims Near-Fable Performance at Half the Price
Anthropic released Claude Opus 5 on July 24, 2026, positioning it as a lower-cost alternative to its more expensive Claude Fable 5 model. Independent evaluators Epoch AI and Artificial Analysis report mixed but largely favorable results, with Opus 5 nearly matching Fable 5 on coding benchmarks while cutting cost-per-task by roughly 20%.
Anthropic Ships Claude Opus 5, Claims It Matches Flagship Fable 5 on Coding at Half the Cost
Anthropic released Claude Opus 5 on July 24, its fourth model launch in under two months, priced at $5 per million input tokens and $25 per million output tokens. The company claims the model matches or beats its flagship Fable 5 on most coding and knowledge-work benchmarks while posting the lowest deception rate of any model it has shipped.
Comments
Loading...