Google DeepMind Launches Lyria 3.5 Music Generation Model in Flow Music
Google DeepMind has released Lyria 3.5, an updated music generation model now live in Google Flow Music. The company claims improvements in melodic complexity, lyric quality, vocal expressiveness, and creative controls like tempo and duration.
Google DeepMind released Lyria 3.5, the newest version of its music generation model, on July 29, 2026, rolling it out immediately inside Google Flow Music.
The update targets four areas: musicality, lyrics, vocals, and creative control. Google DeepMind did not disclose parameter count, training data cutoff, benchmark scores, or pricing for the model.
What's changed
According to Google, Lyria 3.5 introduces the following improvements over prior versions:
- Musicality: The model claims to generate "richer, more complex melodic structures" with more natural-sounding composition.
- Lyrics: Google states the model produces higher-quality lyrics with improved adherence to user prompts and better awareness of song structure (verse, chorus, bridge).
- Vocals: The company claims more realistic and emotionally nuanced vocal output, along with improved pronunciation.
- Creative control: Users get more direct control over tempo and output duration, according to Google.
No technical paper, benchmark comparison, or third-party evaluation accompanied the announcement. All quality claims come directly from Google's blog post and have not been independently verified.
Availability
Lyria 3.5 is available now inside Google Flow Music, Google's music creation product. Google did not specify whether the model is accessible via API, whether it will be offered to enterprise customers, or how it fits into pricing tiers for Flow. Access appears limited to the Flow Music interface at launch.
Google also did not clarify whether Lyria 3.5 replaces the prior Lyria model entirely within Flow or runs alongside it, nor did it specify geographic availability or language support for lyric generation.
What this means
Lyria 3.5 is a narrow, product-focused update rather than a foundational model release with public benchmarks. Google is positioning this squarely as a consumer creative tool improvement inside Flow Music, competing with music generation products from Suno and other AI audio startups that have pushed rapidly on vocal realism and songwriting quality over the past two years.
The lack of published benchmarks or technical detail is consistent with how Google has handled prior Lyria releases — treating the model as an internal capability layer for consumer products rather than a developer-facing API with quantifiable specs. That makes independent comparison to competitors difficult until users can test outputs directly in Flow Music.
The emphasis on lyric structural awareness and vocal pronunciation suggests Google is specifically targeting complaints common to AI music tools: garbled vocals and lyrics that ignore song structure. Whether these improvements hold up against Suno's latest models or open competitors will depend on user testing now that the model is live, rather than any claims in Google's announcement.
Related Articles
Microsoft Releases Mage-VL, a 4B-Parameter Codec-Native Streaming Vision-Language Model
Microsoft has released Mage-VL, a codec-native multimodal foundation model built on a from-scratch 4B-parameter visual encoder paired with Qwen3-4B-Instruct-2507. The model claims up to 3.5x inference speedup over uniform frame sampling and outperforms Qwen3-VL-4B on video and temporal-grounding benchmarks, according to Microsoft.
OpenAI's GPT Transcribe Cuts Word Error Rate to 3.31% but Trails ElevenLabs, Google, and Mistral
OpenAI released GPT Transcribe and GPT Live Transcribe, improving word error rate to 3.31 percent and cutting prices 25 percent to $0.0045 per minute. Independent benchmarks still place OpenAI behind ElevenLabs, Google, and Mistral on transcription accuracy.
Unsloth Releases GGUF Quantizations of Kimi K3, a 2.8T-Parameter Open-Weight MoE Model
Unsloth has released GGUF quantizations of Kimi K3, a 2.8-trillion-parameter open-weight Mixture-of-Experts model from Moonshot AI with a 1-million-token context window and native vision support. The largest lossless quantization (Q8) weighs in at 1.56TB.
Microsoft Releases VibeVoice-ASR-BitNet: 1.58GB Speech Recognition Model Runs Real-Time on CPU, No GPU Needed
Microsoft Research released VibeVoice-ASR-BitNet, a quantized 1.58GB version of its VibeVoice-ASR speech recognition model that achieves real-time inference (RTF < 1) on as few as 3 CPU threads. The model runs 1.6-2.3x faster than Whisper.cpp on commodity x86 and ARM hardware, with a modest accuracy tradeoff.
Comments
Loading...