Suno, known for its artificial-intelligence music generator, has introduced a new Speech capability that produces synthetic voiceovers. The feature is currently in public beta and can be accessed through both the company’s website and its mobile application. Users can generate spoken content while the system simultaneously composes accompanying music, creating a unified audio output.
The addition places Suno among a growing roster of firms offering AI-driven speech synthesis. DeepMind has explored deep-learning voice generation for more than a decade, Adobe provides a text-to-speech service, and ElevenLabs has become a prominent platform since its 2023 launch. Suno’s move appears aimed at broadening its appeal beyond music after its audio generator attracted legal challenges.
A distinctive aspect of Suno’s Speech tool is the optional background music layer, which can be toggled on or off. This design supports scenarios such as soothing soundscapes for poetry readings, energetic accompaniments for dramatic monologues, or motivational tracks for speeches. Users can therefore tailor the auditory atmosphere to match the intended tone of the spoken piece.
To create content, users navigate to the Create tab and select the Speech option. The interface offers two modes: Simple, where a brief description such as “a pirate captain rallying his crew” guides the generation, and Advanced, which accepts a full script and provides controls for voice gender, speaking style, and variability of the output. The system limits each generated clip to roughly eight minutes of audio.
The generated speech is entirely synthetic, produced by Suno’s underlying language model and paired with its music synthesis engine. This combination enables creators to produce complete audio pieces without external editing tools, potentially streamlining workflows for podcasts, advertisements, and educational content that require both narration and background scoring.
As a beta offering, the Speech feature may receive updates that expand voice options, improve realism, or adjust the integration of musical accompaniment. Suno’s entry into the text-to-speech arena reflects a broader industry trend toward multimodal generative AI, where audio, text, and music are increasingly produced by single, adaptable platforms.