Suno Speech Generates Narration and Music Together in One Pass
Suno's new Speech beta generates spoken audio and matching background score as a single track, extending the music app into narrated storytelling.
- Suno launched Speech (beta), generating spoken audio with matching original background music in one pass.
- Pitched as the first model to produce voice and score as a single cohesive track.
- Open to all users inside the Suno app after a one-month closed beta.
- Demos include the Odyssey Sirens passage in Gen Z slang and Victorian-English dish-washing pleas.
- Known rough edges: drifting accents, over-long dramatic pauses, inconsistent pacing.
- Builds on Suno's earlier Bark TTS work and ships alongside the new v6 music model.
Suno Speech generates narration and music in one pass
Suno has opened the beta of Speech, a text-to-audio model that generates spoken narration and an original score as one continuous track. Users provide a script, describe the voice, and specify a musical direction.
Suno describes Speech as the first model to produce narration and a matching score within a single generation. The approach allows vocal delivery, pauses, tempo, and musical cues to develop together, reducing the manual synchronization required when speech and music come from separate tools.
| Status | Open beta |
|---|---|
| Access | Suno’s main app |
| Inputs | Script, voice description, and musical direction |
| Output | A continuous track containing narration and an original score |
| Pricing | Uses Suno’s existing credit-based plans; no separate Speech pricing has been published |
| Developer access | No API, SDK, or programmatic access announced |
Build a scored reading in three steps
- Enter the text to be spoken, such as a poem, toast, story, or message.
- Describe the voice, including qualities such as accent, era, tone, or register.
- Prompt the accompanying music with a genre, mood, or dramatic direction.
The feature uses the same app workflow as Suno’s song generator. Before opening the beta broadly, the company tested Speech for a month with a smaller user group.
Suno’s demonstrations include a passage from the Odyssey rewritten in Gen Z slang and a request for a roommate to wash the dishes delivered in Victorian English. Each example uses music that follows the pacing and dramatic shape of the reading.
The timing advantage
A conventional narrated-audio workflow starts with a text-to-speech model, then adds music in a digital audio workstation. Editors must adjust volume ducking, pauses, tempo, transitions, and emotional cues after both tracks have been created.
Speech generates those elements jointly. A pause in the narration can coincide with a musical transition, while a change in vocal intensity can align with a swell or rhythmic hit. Successful generations could reduce editing work for short pieces that do not require precise, repeatable timing.
Where the beta fits
- Dramatized readings: Poems, scripture, speeches, bedtime stories, and wedding toasts.
- Scored messages: Voice notes or personalized recordings with music matched to the script’s tone.
- Character performances: Prompted accents, periods, and vocal registers paired with suitable musical styles.
- Guided audio: Meditations, pep talks, and other spoken formats that commonly use background music.
- Short-form media: Narrated clips and story segments that can tolerate variation between generations.
Beta rough edges
Suno warns that Speech remains inconsistent. Accents can drift during a recording, including British voices shifting toward Australian pronunciation, and pauses can run longer than the prompt suggests. Pacing and interpretation may also vary across repeated generations.
The beta does not advertise controls for exact duration, word-level timing, placement of musical cues, deterministic regeneration, or cloning a voice from a reference recording. Suno has also not published technical details covering latency, output formats, rate limits, model versioning, or production guarantees.
From Bark to Speech
Speech extends Suno’s work beyond its core music generator. The company released Bark, an open-source speech and audio model, in 2023 before expanding further into generated music. Speech applies that earlier speech expertise inside Suno’s main product and couples it directly to the company’s music stack.
Suno’s separate Voices feature accepts 15 seconds to four minutes of vocal audio and lets users reuse the resulting voice across generations. It requires a spoken verification phrase intended to deter unauthorized cloning. The company has also released its flagship Suno v6 music model.
A gap between voice and music tools
Voice platforms such as ElevenLabs and OpenAI focus on expressive speech, while Suno, Udio, and tools derived from Riffusion specialize in generated music. Most production workflows still create those layers separately. Speech targets short narrated pieces whose score needs to follow the delivery within the same generation.
That capability could support story apps, meditation products, personalized messages, games, and short-form video tools. Its immediate use is limited by app-only access and the absence of documented timing controls.
The missing piece for developers
Suno has not announced an API for Speech. Developers therefore cannot yet submit scripts programmatically, request structured parameters, automate retries, or integrate generated tracks into a production pipeline.
A production release would need documentation for authentication, pricing, quotas, latency, file formats, content limits, licensing, and version stability. Until Suno supplies those details, Speech remains a consumer-facing beta and a preview of how jointly generated narration and music could simplify audio workflows.