Suno, a company known for its AI music generation tool, has launched a new feature called Speech that creates spoken voices from scripts or prompted descriptions. The public beta is available now on Suno's web and mobile platforms, according to an announcement from the company. The feature allows users to generate voiceovers and background music at the same time.
Suno chief product officer Jack Brody said in the announcement that music will always remain central to the company and its products, but that Suno's vision has always reached into other areas of human expression. He described Speech as the first audio model that generates voice and music together as a single cohesive track. The company is positioning the tool as a way to diversify beyond its core music generator.
AI-generated speech is not a new category. DeepMind has been experimenting with deep learning speech synthesis for a decade, Adobe offers a text-to-speech tool, and ElevenLabs has become one of the most recognizable platforms for synthetic speech since it launched in 2023. Suno's entry into the space follows a period in which its music generator has attracted numerous lawsuits, which the company may be hoping to offset by expanding its offerings.
Speech is optional in the sense that users can toggle off the background music if they want clean speech alone. Suno says the music is meant to complement particular uses for generative spoken word, such as a calming soundtrack for poems or something more energetic for dramatic voiceovers and encouraging speeches. Users access the feature by selecting the Create tab and navigating to the Speech option.
Two modes are offered. Simple mode lets users describe what they want to create through a prompt box, with the example of a pirate captain rallying his crew. Advanced mode lets users add a custom script if they already know what they want it to say. Advanced settings also allow adjustment of the AI voice's gender, speech style, and how much variety each voice generation will have. Speech has a maximum duration of around eight minutes.
Suno acknowledges that the feature is far from perfect and says it will keep improving Speech based on user feedback. Brody said beta really does mean beta, noting that British accents can occasionally wander off to Australia and back, that dramatic pauses may be very dramatic, and that users will almost certainly find uses for the tool that never occurred to the company.
More AI news from TechManNews.






