audio generation
3 articles tagged with audio generation
Google Launches Gemini 3.8 Flash TTS and Flash-Lite TTS with Voice Creation from Text Prompts
Google DeepMind has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models that generate custom voices from natural language prompts and support line-by-line performance direction. The models top Hume AI's Voice Design Benchmark at 71.4 and claim first and second place on its Overall Quality Index.
MiniMax Releases H3, a 33B-Parameter Omni-Modal Model That Generates 2K Video With Native Stereo Audio
MiniMax has published MiniMax-H3, a 33-billion-parameter omni-modal generative model capable of producing up to 15 seconds of 2K video with native stereo audio. The model accepts text, image, video, and audio inputs, though its full 2K pipeline depends on a hosted preprocessing component not included in the open-source release.
Stability AI Releases Stable Audio 3 Medium: 2B-Parameter Audio Generation Model with 180-Second Output in Under 2 Secon
Stability AI has released Stable Audio 3 Medium, a 2 billion parameter latent diffusion model capable of generating variable-length audio up to 380 seconds. The model generates music and sound effects in less than 2 seconds on an H200 GPU, trained on 1.28 million licensed and Creative Commons audio recordings.