Lolaby
Lolaby (https://huggingface.co/spaces/build-small-hackathon/lolaby and documentation: https://huggingface.co/build-small-hackathon/lolaby-llama-3b) generates personalised lullabies based on an image (drawing or photo), a child’s first name, age and interests.
The interface offers two steps. The first offers to draw directly on the screen or import a photo. The second collects the first name, the age, what the child likes, something that scares them, an atmosphere and one or more instruments among six (music box, guitar, keyboard, ocarina, harp, xylophone). Advanced settings allow you to choose the key and tempo.
Models used: a vision model (MiniCPM-V 4.6, 1.3 billion parameters) analyzes the visual and extracts a description and a Llama 3.2 3B model generates the lyrics by integrating all these elements. Music synthesized by signal processing. Playback provided by Kokoro TTS 82M.
On-premises generation without calling a cloud service.
Note: lyrics are generated in English by default but an instruction that asks for French in the field “What do they love?” produces lyrics in French, but the text-to-speech voice (Kokoro TTS) remains with an English accent. The “Just the words” option allows you not to generate the music and audio.
Open source, free and account-free.



