Suno adds Speech, which generates spoken audio with matching music

Suno, best known as an AI music generator, is adding a new feature called Speech. It produces spoken text and matching background music together as a single audio track.
The workflow is simple. Users type in an idea or written text and describe the voice and the music style they want. The model then generates both the voice and the sound.
Suno product chief Jack Brody says the company tested Speech with a small group for a month. Suno says the feature can be used for poems, meditations, and bedtime stories. The beta still has bugs, though: for example, a British accent can sometimes sound Australian.
Suno has not said how it trained the model. That matters because AI music generators face criticism over potential copyright infringement. Major record labels have already sued Suno, and a Munich court recently ruled against the startup, rejecting fair use as a justification for using copyrighted data.
Key facts
- Suno's new Speech feature produces spoken text and matching background music together as a single audio track.
- Users type an idea or written text and describe the voice and music style they want; the model generates both the voice and the sound.
- Suno product chief Jack Brody says the company tested Speech with a small group for a month.
- Suno suggests poems, meditations and bedtime stories as uses; the beta still has bugs, such as a British accent sometimes sounding Australian.
- Suno has not said how it trained the model, while major record labels have sued it and a Munich court recently rejected fair use as a justification.
Why it matters
Suno is known for music generation. Speech extends it to spoken audio, with the voice and the backing music produced in one pass as a single track rather than assembled separately. It is an incremental product step, but it moves the company from songs toward other kinds of audio content.
Who it affects
Suno says the feature suits poems, meditations and bedtime stories, so the obvious audience is people making short spoken pieces with a musical backdrop. The record labels suing Suno and the court that ruled against it are also part of the picture, since the training question applies to any new model the company ships.
How to use it
You type in an idea or written text and describe the voice and the music style you want. The model then generates both the voice and the sound. The source gives no release date, pricing or plan availability for Speech.
How solid is it
The description comes from a single report that cites Suno product chief Jack Brody on the month-long test with a small group. The source does not give the size of the group, and it does not list supported languages or voices. Suno itself acknowledges the beta has bugs.
Risks and caveats
Suno has not said how it trained the model. Major record labels have already sued the company, and a Munich court recently ruled against it, rejecting fair use as a justification for using copyrighted data. The source does not say whether that ruling concerns Speech specifically. On the product side, the beta is buggy: a British accent can sometimes sound Australian.