Suno launches Speech, generating voiceovers and background music together

By EnkiEdited by VK, Editor

Published

Reporting from The Verge, The Decoder

Suno's new Speech beta turns a prompt or a script into spoken audio with a matching music bed in one track, up to about eight minutes, on web and mobile. Suno has not said what it trained the model on.

What it means for founders

  • If you produce podcasts, app onboarding audio, meditation content or ads, Speech is a cheap way to prototype a narrated track with a score in one pass. Test it against your current voice vendor on a real script before switching anything.
  • Check the license before shipping output commercially. With training data undisclosed and litigation ongoing, a court outcome against Suno could change what customers are allowed to do with generated audio.
  • For voice startups, the bar just moved: bundled music is now a feature from a well funded consumer brand. Differentiation will have to come from voice quality, control, latency or clear rights.
  • Watch for Suno's pricing announcement and any ruling in the US label cases. Those two signals decide whether Speech is a toy or a production tool.

The story

Suno, the AI music generator, has launched Speech, a public beta that produces a spoken voice and accompanying background music as a single track. It is available now in Suno's web and mobile apps, under the Create tab. Suno describes it as the first audio model to generate voice and music together, a claim about its own approach rather than the result of any outside benchmark.

How it works

There are two modes. Simple mode takes a short description, such as a pirate captain rallying a crew, and writes and performs it. Advanced mode accepts a full script and adds controls for the voice's gender, the speaking style and how much each generation varies. The music bed is optional and can be switched off with a toggle for clean narration. A single generation runs to roughly eight minutes.

Suno's suggested uses are poems, guided meditations, bedtime stories and dramatic voiceovers, where a fitting soundtrack does real work. Product chief Jack Brody said the feature spent a month with a small test group before release, and the company is candid about rough edges: accents can drift, with a British voice sometimes sliding toward Australian, and dramatic pauses can run very long. "Beta really does mean beta," Brody said in the announcement, as reported by The Verge.

Why Suno is doing this

Text to speech is a crowded field. ElevenLabs has become the best known specialist since 2023, Adobe ships its own tool, and Google DeepMind has worked on speech synthesis for a decade. Suno's angle is the pairing with music, which none of those lead with.

There is also a business reason to diversify. Suno's core music product faces lawsuits from major record labels, and a Munich court recently ruled against the company, rejecting fair use as a defense for training on copyrighted recordings. Suno has not said what data trained Speech, which leaves the same question hanging over the new feature.

What we don't know yet

Suno has not published pricing for Speech, said how generations count against existing plan credits, or clarified the commercial rights attached to speech output. It has not said whether users can clone or upload their own voice.

Sources

Enki Daily

Get stories like this every weekday morning.

The day's AI stories for founders, each with what it means for your company. Free.

Tools in this story

We may earn a commission if you sign up through our links. It never affects our ratings or which stories we cover.

Suno logo
Suno

Full songs from a single prompt

8.2Editor’s scoreVisit Suno
ElevenLabs logo
ElevenLabs

The most lifelike AI voices

9.0Editor’s scoreVisit ElevenLabs
Udio

AI song generation, now licensed and stream-only

6.8Editor’s scoreVisit Udio

More in Products & Launches

How Enki covers newsCorrectionsReport an error

Search Enki

Search AI tools, categories and news