product updateSuno

Suno launches Speech public beta: AI voiceovers with background music, up to about eight minutes

TL;DR

Suno has launched Speech in public beta on web and mobile. It generates spoken voice from a script or text prompt, with optional AI-generated background music, for clips up to roughly eight minutes. Suno claims it is the first audio model to generate voice and music together as one track.

2 min read
0

Suno has launched Speech, a feature that generates spoken-word audio from a script or a descriptive prompt, in public beta on its web and mobile platforms. Users can generate a voiceover and matching AI background music in a single pass. Maximum output length is around eight minutes.

What Speech does

Speech sits in Suno's "Create" tab and has two modes:

  • Simple: The user describes the desired output in a prompt box, for example "a pirate captain rallying his crew."
  • Advanced: The user supplies a custom script for the voice to read verbatim.

Advanced settings let users adjust the voice's gender, speech style, and how much variety each generation has. The background music is optional and can be switched off with a toggle for clean speech. Suno positions the pairing for uses such as calm soundtracks under poetry, or more energetic music for dramatic voiceovers and motivational speeches.

Suno's claims

Suno chief product officer Jack Brody called Speech "the first audio model that generates voice and music together as one cohesive track." That is a company claim, and Suno has not published benchmarks or technical documentation to support it. Competing text-to-speech products from ElevenLabs, Adobe and others already exist, though the source does not indicate whether any of them combine voice and generated music in one output.

Brody was candid about the beta's limits: "Beta really does mean beta. Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic." Suno says it will refine the feature based on user feedback.

What has not been disclosed

  • Model name and version: Suno has not given a version identifier for the Speech model.
  • Pricing: Pricing and plan availability not yet disclosed.
  • API access: Not mentioned in the announcement coverage.
  • Parameter count, training data and benchmark scores: Not disclosed.
  • Supported languages and accents: Not specified beyond Brody's reference to British accents.

What this means

Speech moves Suno from a single-purpose music generator into adjacent audio territory, where ElevenLabs is the best-known specialist. The differentiator Suno is betting on is integration: one generation yields voice and score together, which removes the manual mixing step creators now do after generating speech and music separately. Whether the quality holds up is untested, since no independent evaluations exist and Suno itself says the beta has rough edges such as accent drift.

The move also has a business angle. The Verge suggests Suno may be diversifying because its music generator has drawn numerous lawsuits. That is the outlet's interpretation, not a Suno statement. Speech output is not music, though the feature still generates background music, so it does not fully remove Suno from that exposure.

The eight-minute cap makes Speech suited to short-form content such as narration, ads and social clips, not long-form audiobooks. Watch for API availability, pricing, and independent voice-quality comparisons against dedicated text-to-speech providers.

Related Articles

product update

Pi 1.0 agent harness goes stable; Pi Durable ports it to TypeScript with crash-resumable state

Pi, the minimalist agent harness now under Earendil, has reached version 1.0. A companion release, Pi Durable, ports it to TypeScript and externalizes all stateful components so agents can resume after crashes. Pricing and licensing were not disclosed in the source.

product update

Cline CLI v3.0.68 fixes agent-team slowdown that could grow teams.db to gigabytes

Cline CLI v3.0.68 fixes a performance bug in which every streamed chunk re-saved the entire agent-team state, slowing long runs and letting ~/.cline/data/db/teams.db grow to gigabytes. The release also refreshes the model catalog and changes default models for five providers.

product update

GPT-6 Astra Ultrafast Claims Up to 8x Faster Token Generation on NVIDIA Blackwell GPUs

GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. NVIDIA says it generates tokens up to 8x faster than Astra Standard mode. Pricing and context window details were not disclosed in the source.

product update

Google launches Guided Vision in Gemini Live, giving real-time audio descriptions through the Android camera

Google began rolling out Guided Vision in Gemini Live on compatible Android devices on October 1, 2026. The feature uses the phone camera to give real-time audio descriptions of surroundings and objects, with follow-up questions supported. Google warns it should not be used for navigation or obstacle detection.

Comments

Loading...