StepAudio 3 Music
StepAudio 3 Music (https://huggingface.co/spaces/stepfun-ai/StepAudio-3-Music and documentation: https://platform.stepfun.ai/docs/en/guides/models/stepaudio-3-music) generates songs or instrumental pieces from a description, with or without lyrics.
Three modes offered:
- In “Song” mode, the mandatory description, optional lyrics and title (if empty lyrics the model writes them. “Instrumental” option.
- “Vocal to Song” builds an arrangement around an imported a cappella voice (wav, flac, mp3 or opus).
- “Song Cover” changes the style of an imported song (same formats).
Songs up to 5 min 30 in MP3 (format and bitrate adjustable in “Advanced”).
Note: interface in Chinese and English, French accepted for lyrics.
Free and open-source, usage limits related to Hugging Face (vary depending on demand on the site).




Sources: