Suno launches Speech, generating voiceovers and background music together
By EnkiEdited by VK, Editor
Published
Reporting from The Verge, The Decoder

Suno's new Speech beta turns a prompt or a script into spoken audio with a matching music bed in one track, up to about eight minutes, on web and mobile. Suno has not said what it trained the model on.
What it means for founders
- If you produce podcasts, app onboarding audio, meditation content or ads, Speech is a cheap way to prototype a narrated track with a score in one pass. Test it against your current voice vendor on a real script before switching anything.
- Check the license before shipping output commercially. With training data undisclosed and litigation ongoing, a court outcome against Suno could change what customers are allowed to do with generated audio.
- For voice startups, the bar just moved: bundled music is now a feature from a well funded consumer brand. Differentiation will have to come from voice quality, control, latency or clear rights.
- Watch for Suno's pricing announcement and any ruling in the US label cases. Those two signals decide whether Speech is a toy or a production tool.
The story
Suno, the AI music generator, has launched Speech, a public beta that produces a spoken voice and accompanying background music as a single track. It is available now in Suno's web and mobile apps, under the Create tab. Suno describes it as the first audio model to generate voice and music together, a claim about its own approach rather than the result of any outside benchmark.
How it works
There are two modes. Simple mode takes a short description, such as a pirate captain rallying a crew, and writes and performs it. Advanced mode accepts a full script and adds controls for the voice's gender, the speaking style and how much each generation varies. The music bed is optional and can be switched off with a toggle for clean narration. A single generation runs to roughly eight minutes.
Suno's suggested uses are poems, guided meditations, bedtime stories and dramatic voiceovers, where a fitting soundtrack does real work. Product chief Jack Brody said the feature spent a month with a small test group before release, and the company is candid about rough edges: accents can drift, with a British voice sometimes sliding toward Australian, and dramatic pauses can run very long. "Beta really does mean beta," Brody said in the announcement, as reported by The Verge.
Why Suno is doing this
Text to speech is a crowded field. ElevenLabs has become the best known specialist since 2023, Adobe ships its own tool, and Google DeepMind has worked on speech synthesis for a decade. Suno's angle is the pairing with music, which none of those lead with.
There is also a business reason to diversify. Suno's core music product faces lawsuits from major record labels, and a Munich court recently ruled against the company, rejecting fair use as a defense for training on copyrighted recordings. Suno has not said what data trained Speech, which leaves the same question hanging over the new feature.
What we don't know yet
Suno has not published pricing for Speech, said how generations count against existing plan credits, or clarified the commercial rights attached to speech output. It has not said whether users can clone or upload their own voice.
Sources
Enki Daily
Get stories like this every weekday morning.
The day's AI stories for founders, each with what it means for your company. Free.
Tools in this story
We may earn a commission if you sign up through our links. It never affects our ratings or which stories we cover.

Full songs from a single prompt

The most lifelike AI voices
AI song generation, now licensed and stream-only
More in Products & Launches
- Meta open sources firmware and an SDK so anyone can build Muse gadgets

For founders: Hardware startups get a free agent with Meta's distribution behind it to build on.
TechCrunch · 13h ago - OpenAI launches Dots, always-on agents for ChatGPT Pro and business plans

For founders: Budget for usage, not seats. The chat is included, but anything a dot builds in Codex draws on the same limits your team already spends, so a busy dot can…
TechCrunch · 3d ago - OpenAI adds Codex cloud environments, a Decisions API and an Ultrafast tier

For founders: Speed is now a line item. Ultrafast turns latency into something you buy per token at a steep premium, so reserve it for user facing paths where waiting loses…
TechCrunch · 3d ago - OpenAI turns ChatGPT into an app platform, with no revenue share yet

For founders: Distribution is real, monetization is not. A spot in the sidebar of a product with a billion weekly users is a serious channel, but you still need your own…
TechCrunch · 3d ago - Google will turn Gemini Gems into skills starting November 17

For founders: Audit anything built on Gems now. If your team or customers rely on a Gem that leans on Canvas, Deep Research or GitHub files, it may lose that behavior in…
TechCrunch · 4d ago