How do I generate music and voiceovers?
Stable Audio makes tracks up to six minutes for a flat 25 credits. Text-to-speech is 1 credit per 1,000 characters.
Music
Describe the genre, mood and instrumentation in the Music Generator and it produces an original track, with AI prompt enhancement and auto-generated cover art included.
| Model | Length | Credits |
|---|---|---|
| Stable Audio (newest) | 30s to 6 minutes, selectable | 25 flat |
| Stable Audio 2.5 | 30s to 3 minutes, selectable | 25 flat |
| Google Lyria 2 | About 30 seconds, fixed | 15 |
Stable Audio costs the same 25 credits whether you generate 30 seconds or six minutes, so there is no reason to under-request length. Album mode generates 3 to 12 tracks at once with shared cover art.
The Music Generator takes a text prompt only — there is no lyrics field, vocals toggle or reference audio upload.
Text-to-speech
1 credit per 1,000 characters, up to 5,000 characters per generation. There are 50+ ElevenLabs stock voices plus any voice you create yourself, with favourites and category filters. There is no model dropdown — the right provider is chosen automatically from the voice you pick.
Custom voices
- Voice Design (100 credits) — describe a voice in words and get a brand new one. Needs a name, a description and 100 to 500 characters of sample text.
- Voice Clone (100 credits) — clone from an audio sample of at least 10 seconds, up to 25 MB, with a consent confirmation.
Custom voices support seven emotions — neutral, happy, sad, angry, fearful, disgusted and surprised — and are saved to your voice library so they become selectable in Text-to-Speech, Lipsync, UGC Builder and elsewhere.
Cloning a voice requires documented consent from the person it belongs to. See our policy on likeness and voice.
Last reviewed against the live product in August 2026.