One day, one generative AI tool

Focus of the letter 24

Sesame and Kugel Audio made me want to talk to you this week about voice generation.

I test dedicated applications quite regularly (85 are present in the “voice” category of the site), this was the case recently last week with Voxtral or Scribe three weeks ago.

More than the enormous technological advances made in recent months (a bit like distorted faces for images, metallic robotic voices are a long way off), it is rather on these low-noise advances that I would like to dwell today.

Perhaps because these models and applications are less visible and less publicized, it seems to me that we have followed their evolution less, whereas in my opinion they certainly have more consequences in our lives. We have already talked about them here: direct consequences on professions dedicated to voice as well as for dubbing, of course, but also impacts on the voices we hear on a daily basis. On the one hand, there are positive aspects: speech synthesis accessible to all, many open-source models, free applications for accessibility, or the possibility of listening to texts when you can’t read them. On the other hand, there are more negative aspects, such as the exponential multiplication of audio content or the excesses of voice cloning, which is now possible with only a few seconds of recording.

Kyutai with its Unmute model had already introduced the analysis and generation in double “flow” (a simplification perhaps excessive…), taken up here by Sesame : unlike a classic conversational tool that receives audio, processes it and then generates its response without being able to perform any other task, the incoming stream and the outgoing flow are processed simultaneously. This allows the AI to produce audio and analyze the received audio at the same time, as in a human conversation.
Our relationship with the machine seems to me to reach a new level, that of perception and emotion. The mediation of writing and its necessary prior reflection no longer exists and the direct and constant flow installs a new, almost more “intimate” relationship that could make us forget the machine…

The automation of phone calls, for example, is now possible and increasingly likely. Perhaps, like me, you recently received a troubling cold calling call generated by AI…
Audio deepfakes are more difficult to spot, clues are erased little by little, there are still silences that are sometimes absent, too short or too long, misplaced intonations or sometimes short background noises. What to do then? Exercising critical thinking and applying it seems more and more necessary…