Google delays Gemini 3.5 Pro release after disappointing coding performance in June training update
Google has delayed the release of Gemini 3.5 Pro past its June deadline due to coding performance issues. The company retrained the model in late June with new data but saw disappointing results, according to Bloomberg. An upgraded Flash model is now in testing with partners.
Google delays Gemini 3.5 Pro release after disappointing coding performance in June training update
Google has pushed back the release of Gemini 3.5 Pro past its June deadline after coding performance failed to meet expectations, according to Bloomberg.
What happened
Google announced Gemini 3.5 Flash at I/O 2026 in mid-May and said the Pro version would arrive in June. That deadline has passed with no update.
According to Bloomberg, Google is "taking time to try to improve [Gemini 3.5 Pro's] capabilities, particularly in coding." In late June, Google updated the training data in an attempt to improve coding skills, but the results were disappointing.
The timeline suggests development saw a reset between I/O and the missed launch. Performance in other domains remains unclear. Gemini 3.1 Pro, the current flagship model, dates back to February 2026.
What Google says
In a statement, Google said it is "currently testing 3.5 Pro, an upgraded Flash model, and other models with partners." The company added: "We're shipping quickly across a wide range of models while keeping them highly cost-effective for customers."
No new timeline for Gemini 3.5 Pro's release was provided.
Internal coding context
As of April 2026, 75% of all new code at Google is AI-generated and approved by engineers, up from 50% last fall, according to Bloomberg. However, efforts face resistance from engineers who believe important code should be human-written to adhere to Google standards, according to ex-employees.
Internally, engineers are facing AI capacity restraints with coding tools. Google is working to "unite the company's internal artificial intelligence coding tools."
Google DeepMind (AI Studio), Cloud (Vertex), and the Android team (Android Studio) each maintain separate AI coding tool efforts.
What this means
The delay signals Google is prioritizing quality over speed for its flagship Pro model, particularly in coding—a domain where OpenAI and Anthropic have set high benchmarks. The fact that a late June retraining attempt failed suggests deeper architectural or data challenges rather than simple fine-tuning issues. With 75% of Google's own code now AI-generated, the company faces pressure to ship coding models that match internal standards while competing externally.
Related Articles
OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold
OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.
Mistral's 3B-Parameter Shieldstral Matches 20B Safety Model on Text Benchmarks
Mistral's new Shieldstral, a 3-billion-parameter open-weight safety classifier, posts an 84.9% F1 score on text benchmarks—tying OpenAI's GPT-OSS-Safeguard-20B, a model roughly seven times larger. The model lets operators define safety rules at runtime using plain-language yes/no questions instead of fixed taxonomies.
Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier
Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.
Black Forest Labs Launches FLUX 3 Video, Claims It Beats Seedance 2.0 on Elo Rankings
Black Forest Labs has made FLUX 3 Video generally available via its API, offering up to 20-second HD/Full HD clips with native audio and lip-sync in 14+ languages. The company claims its internal Elo benchmarks put the model ahead of Seedance 2.0, Gemini Omni Flash, and Minimax H3.
Comments
Loading...