Skip to content
News

Suno AI Adds Speech Feature to Generate Spoken Voices

Suno AI Adds Speech Feature to Generate Spoken Voices - Suno Speech
Suno launches Speech, a public beta feature generating spoken voices with optional background music across its web and mobile platforms.

Suno is moving beyond AI music with a new feature that generates spoken voices from scripts or prompted descriptions. The tool, called Speech, is now available in public beta across Suno’s web and mobile platforms, and can simultaneously produce voiceovers alongside background music to accompany them.

According to Suno chief product officer Jack Brody, music remains central to the platform, but the company’s vision has always extended to other forms of human expression. He described Speech as the first audio model that generates voice and music together as a single cohesive track.

How the Speech feature works

Pairing AI music with generated voices is Suno’s approach to text-to-speech technology. The background music is optional and can be switched off with a toggle if clean speech is preferred, but the intention is to complement particular use cases. Examples include a calming soundtrack for poems or a more energetic backing for dramatic voiceovers and motivational speeches.

To use the feature, users select the “Create” tab and navigate to the Speech option. There are two modes available. Simple mode allows a description of the desired output via a prompt box, such as a pirate captain rallying his crew. Advanced mode lets users add a custom script when they already know exactly what the voice should say. Advanced settings also allow adjustment of the AI voice’s gender, speech style, and the degree of variety in each generation. Speech has a maximum duration of around eight minutes.

A crowded market for generated voices

AI-generated speech is not new. DeepMind has experimented with deep learning speech synthesis for a decade, Adobe offers a text-to-speech tool, and ElevenLabs has become one of the most recognisable platforms in the field since launching in 2023. Suno is entering the space in what appears to be an effort to diversify the platform, given that its music generator has attracted numerous lawsuits.

Suno acknowledges that the feature is far from perfect and says it will continue to refine Speech based on user feedback. Brody noted that beta genuinely means beta, adding that British accents can occasionally wander off to Australia and back, and that dramatic pauses may become very dramatic. He also suggested users would find applications for the tool that the company had not anticipated.

The Speech feature is live now in public beta on both Suno’s web and mobile applications.

Source
Image: theverge.com

The UK tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your UK list is ready.

Shop on Amazon UK — Discover deals Shop on Amazon UK — Discover deals