DocsToAudio
DocsToAudio (https://docstoaudio.online/fr) converts PDF, EPUB, DOCX, and TXT documents into audio files. Export in MP3 format (one file per chapter,...
Proseable
Proseable (https://www.proseable.com/) has students practice a foreign language orally by conversation with an AI, on topics chosen according to their...
Voice to Text
Voice to Text (https://voicetotext.pro/) is a Chrome extension that transcribes voice to text using Whisper AI, either directly recorded or...
CastReader
CastReader (https://castreader.com/) reads web pages, PDFs, Kindle books aloud and explains their content with synchronized highlighting. Extension for Chrome, Edge...
TextaVoice
TextaVoice (https://www.textavoice.com/) converts text to an audio file. The text (2000 characters maximum) is associated with a language and a...
Vibes
Vibes (https://vibes.ai/) generates images and videos from an instruction.Integrated with the Meta AI app (browse, remix, publish) and accessible from...
GPT-Live
GPT-Live (https://openai.com/index/introducing-gpt-live/) is OpenAI’s new voice model that powers ChatGPT Voice, real-time conversations with simultaneous listening and speech. GPT-Live processes...
Longscribe
Longscribe (https://longscribe.com) transcribes videos and long recordings (YouTube, Vimeo, TikTok, Instagram, Facebook, Twitch, X, Dailymotion, SoundCloud or imported audio/video files)....
Kyutai Pocket TTS
Kyutai Pocket TTS (https://kyutai.org/tts/ — documentation: https://github.com/kyutai-labs/pocket-tts) is an open-source text-to-speech model. Choice of language (French, English, German, Spanish, Italian,...
Lolaby
Lolaby (https://huggingface.co/spaces/build-small-hackathon/lolaby and documentation: https://huggingface.co/build-small-hackathon/lolaby-llama-3b) generates personalised lullabies based on an image (drawing or photo), a child’s first name, age...
MAI Playground
MAI Playground (https://playground.microsoft.ai/) provides access to Microsoft’s AI models: image generation, audio transcription, text-to-speech, and multimodal conversation. 4 models available:...
DocsToAudio
DocsToAudio (https://docstoaudio.online/) converts pdf, epub, docx and txt files to MP3 or M4B audio with 300 AI voices in 50...
Our stories
Our Stories (https://readourstories.org/) generates children’s stories from an instruction. The generated stories can be translated into multiple languages (10 optimized,...
Linto Studio
LinTO Studio (https://studio.linto.ai/ and documentation: https://github.com/linto-ai) transcribes and generates subtitles from audio or video files. Developed by Linagora 🇫🇷 and...
Honen
Honen (https://honen.com/) generates courses from an instruction and/or a supporting document. Several activities are offered, they can be selected or...
Reka
Reka (https://reka.ai/ and models: https://huggingface.co/RekaAI) offers several multimodal models for image and video analysis, text and voice.Two first versions were...
Podcasty
Podcasty (https://podcasty.fm/) generates podcast episodes from online articles. Two entries possible: via an RSS feed for automation or article by...
Gemini 3.1 TTS Preview
Gemini 3.1 TTS Preview (demo on Google Ai studio https://aistudio.google.com/generate-speech?model=gemini-3.1-flash-tts-preview) is Google’s latest text-to-voice model. Two modes of presentation: text...
Holotab
Holotab (https://chromewebstore.google.com/detail/holotab/hlaoiikljjgcjdhkakedfngifaopbcop) is a Chrome extension launched by HCompany (a French company already seen here with Runner H) and which...
Voice Generator
Voice Generator (https://voice-generator.com/) transforms texts into voices from written or pasted text, pdf or epub. 12 languages and 70 voices...
Seedance 2
Seedance 2 (available in Capcut and in Dreamina) is the latest version of ByteDance’s video generation model. On CapCut, two...
Coherent transcribed
Cohere Transcribe (documentation: https://huggingface.co/CohereLabs/cohere-transcribe-03-2026 and demonstration: https://huggingface.co/spaces/CohereLabs/cohere-transcribe-03-2026) transcribes recorded or imported audios. 12 languages available including French.Very faithful and fast...
Cartesia Sonic
Cartesia Sonic (https://cartesia.ai/sonic) is a “Text-to-Voice” application deployed with two versions: a text generation version and a searchable audio chatbot...
Runable
Runable (https://runable.com) brings together several recent templates and agents in a single application to generate different types of text, image,...
Slides orator
Slides orator (https://www.slidesorator.com/) generates an explainer video with an audio 3D avatar from a pdf. From the pdf, a knowledge...
Inkr
Inkr (https://inkr.app/) transcribes audio from audio or video files, a URL, or a live recording. Transcription is offered in a...
Kugel Audio
Kugel Audio (documentation: https://github.com/Saganaki22/ComfyUI-KugelAudio and demonstration: https://huggingface.co/spaces/multimodalart/kugelaudio) is a text-to-speech (text-to-voice) model with 24 European voices. Two possibilities: generation from...
Sesame
Sesame (https://www.sesame.com/) is an oral conversation model currently in research version, based on the open-source csm model (https://github.com/SesameAILabs/csm).Current research is...
Voxtral
Voxtral (documentation: https://mistral.ai/fr/news/voxtral-transcribe-2 and demo: in Mistral AI studio or for direct transcription voxtral mini realtime – out of order...
Leadde
Leadde (https://leadde.ai) generates an explainer video with avatar and voice from a document or presentation. Choice of avatar and voice,...
Anam
Anam (https://anam.ai) creates chatbots based on video avatars from an instruction. 8 language models, 42 avatars and 499 voices available...
Supertonic
Supertonic (documentation: https://huggingface.co/Supertone/supertonic-2 and demo: https://huggingface.co/spaces/Supertone/supertonic-2) is a “text-to-voice” model that can be installed and run locally on your computer...
Scribe V2
Eleven Labs’ Scribe v2 (https://elevenlabs.io/speech-to-text) is a voice-to-text transcription model with two versions: v2 real time for on-the-fly transcriptions and...
MindVideo
MindVideo (https://www.mindvideo.ai) has 40 image generation templates and 30 video generation templates, including the most popular (including Flux unlimited and...
Google AI Edge Gallery
Google AI Edge Gallery (Github: https://github.com/google-ai-edge/gallery, direct IOS and Android Appstore links via the Play Store) is an experimental mobile...
Chatterbox
Chatterbox (documentation: https://github.com/resemble-ai/chatterbox and demonstration: https://huggingface.co/spaces/ResembleAI/Chatterbox-Multilingual-TTS) is a “text to speech” model open-source available in 23 languages for its Multilingual...
Audio converter
Audio Converter (https://audioconverter.ai/) transcribes audio from a file, a YouTube video or a url and then offers export, a text...
Design Arena
Design Arena (https://www.designarena.ai) is a platform that offers a ranking of multimedia generative AI models (image, video, presentations, applications, 3D,...
Youmind
Youmind (https://youmind.com) groups relevant sources from an instruction in a table in the form of materials for other generations or...
Gemini TTS
Gemini TTS (https://aistudio.google.com/generate-speech) is Gemini’s text-to-voice model, developed in two versions, Gemini-2.5-pro-preview-tts and gemini-2.5-flash-preview-tts, and available on Google ai studio....
SAM Audio
Meta’s Sam Audio (description: https://ai.meta.com/samaudio/ and demo: https://aidemos.meta.com/segment-anything) searches for and isolates elements of a video or audio file from...
YapperBot
YapperBot (https://www.yapperbot.com/) generates audios and videos from a persona and an instruction. After creating the persona with the choice of...
Fobizz
Fobizz (app.fobizz.com) brings together several generative AI tools for education: chatbot, document analysis, templates for generating course materials, personalized assistants,...
Time AI
Time AI (https://time.com/timeai/) is Time magazine’s conversational agent based on the content of its content and archives (102 years announced)....
Unrav
Unrav (https://unrav.io/) supports the understanding and communication of web pages and YouTube videos by generating 15 different “interpretations” – the...
Handy
Handy (https://handy.computer/ and documentation https://github.com/cjpais/Handy) is a speech to text tool, available on MacOS, Windows, and Linux and that works...
Blobu
Blobu (https://blobu.ai) generates summaries of books in its database. It also offers “mash-ups”, a story based on several books in...
Anpo
Anpo (https://anpo.ai) generates images, videos and realities from images, texts, and audio generated or imported on an “infinite” board. Each...
Oboe
Oboe (https://oboe.fyi/) generates a course from an instruction and/or a document. Several contents are generated: podcast (in English), deep dive,...
Parlai
Parlai (https://www.parlai.app/) uses a WhatsApp conversation for language learning with the help of AI via text or voice messages, oral...
Underlord of Descript
Underlord by Descript (https://www.descript.com/ – video and audio editing and creation tools) is a chatbot that allows the generation of...
HunyuanVideo-Foley
HunyuanVideo-Foley (description: https://huggingface.co/tencent/HunyuanVideo-Foley and demonstration: https://huggingface.co/spaces/tencent/HunyuanVideo-Foley) sounds videos from the original and an instruction (optional). Up to 6 audios can...
Gemini Storybook
Storybook (https://gemini.google.com/gem/storybook/) is a Gemini “gem” that generates stories of about ten illustrated pages (only one image style possible) with...
Tldraw Computer
Tldraw Computer (https://computer.tldraw.com) automates tasks organized by logical blocks on a whiteboard. Generates text, image, and voice from given sources,...
Gladia
Gladia (https://www.gladia.io/) transcribes videos or audios into text. Speaker recognition (“diarization”), translation possible. Export JSON, SRT, VTT and TXT. Live...
Minimax audio
Minimax audio (https://www.minimax.io/audio): “text to speech” and music generation grouped together in the same tool. “Text to speech” brings together...
Kyutai STT
Kyutai STT (https://kyutai.org/next/stt) is Kyutai’s “speech to text” model, available in two versions open-source: STT-1B-en_fr (Documentation: https://huggingface.co/kyutai/stt-1b-en_fr): , which includes...
Tila
Tila (https://tila.ai/) brings together all possible generations in a single tool in the form of an infinite table: text, image,...
RoboTeach
RoboTeach (https://roboteach.us) generates a course from a given topic. The course is offered with the possibility of editing, its audio...
Unmute
Kyutai’s 🇫🇷 Unmute (https://unmute.sh/), based on Gemma 3 and Moshi, is an audio chatbot with very low latency of very...
Gemini Deep Research
Gemini Deep Research available in Gemini (https://gemini.google.com) generates answers with “reasoning” from an instruction and/or documents and then, after establishing...
Teach me anything
Teach me anything (https://tma.live) generates an explainer video from an instruction. A video generated as an animated and commented presentation...
Research Bunny
Research Bunny (https://www.researchbunny.com) is a search engine for scientific articles. Searches are carried out via 25 categories and keywords with...
Monica’s Tools
Monica’s Toolkit (https://monica.im/tools) – the extension of which has already been described here – brings together a series of generative...
OpenAI FM
OpenAI FM (https://www.openai.fm/ and documentation: https://platform.openai.com/docs/guides/audio) is the demo space for OpenAI’s new voice generation models, based on GPT 4o....
Imagine explainers
Imagine explainers (https://imagineexplainers.com) generates image, text and voice videos from an instruction with internet search and possible images, or from...
Conversational AI
Conversational AI from Eleven Labs (https://elevenlabs.io/app/conversational-ai) creates specialized audio chatbots from instructions and a knowledge base. Models for Claude, Gemini,...
Hedra’s Character 3
Character 3 (https://www.hedra.com) is Hedra’s new video generation model (tested in June 2024): from an instruction describing a scene with...
Uniscribe
Uniscribe (https://www.uniscribe.co) transcribes audio, video (mp3, mp4, mpeg, mpga, m4a, wav, and webm) or YouTube documents to produce a summary,...
Zonos
Zonos (description: https://www.zyphra.com/post/beta-release-of-zonos-v0-1#zonos_2, test space: https://playground.zyphra.com/audio) voiced a pasted text. 4 expressive voices and 6 languages available. Cloning voices from...
Rapport
Rapport (http://rapport.cloud) allows you to create chatbots with avatar and voice from a pre-prompt. Main models available (LLM, voice recognition...
Zenmic
Zenmic (http://zenmic.com) generates a podcast episode from a topic or pasted text. Editable script. FR if specified, set the number...
Humva
Humva (https://humva.ai) uses avatars to speak a text up to 2000 characters in 10 min. About 200 avatars available, 5...
Vera
Vera (https://askvera.org) is an audio/phone and text/Whatsapp chatbot that offers to verify the information offered. Based on ChatGPT4 and 350...
GenFM
GenFM by @elevenlabsio (http://elevenlabs.io/genfm) generates podcasts from documents, text, url, YouTube videos or scanned document. Choice of voices FR automatic,...
Ellipsis News
Ellipsis News (https://ellipsisnews.co) generates a daily audio synthesis of a news topic with global sources cited. 5 accessible themes (change...
Chat Camera
Chat Camera (https://talkycamera.com/chat-camera) allows analysis and real-time audio or text conversation via the camera (smartphone or tablet) and chatGPT 4o....
Mumble
Mumble (http://mumbleapp.com): a note-taking application that transcribes audio and organizes with the addition of a To-Do list. Possible addition of...
Playcast
Playcast (https://playcast.ai) generates summaries of texts, documents, URLs or YouTube videos as podcast episodes available in dedicated apps. Example in...
Swift
Swift (https://swift-ai.vercel.app): a voice assistant based on LLama developed by Groq, Cartesia and Vercel. Quick reaction allows conversations (FR if...
Cleft
Cleft (https://cleftnotes.com): iOS note-taking application based on an audio request. The transcribed note (to be improved) is increased and formatted...
Turboscribe
Turboscribe (https://turboscribe.ai) transcribes audio files into text with timing, 11 import formats, 5 export formats. Quick to copy/paste for text...
Dicte
Dicte (https://dicte.ai): a tool dedicated to the transcription of oral information (meetings, conferences, etc.) offered by@livdeo » 👏🇫🇷 Transcription, reports,...
Text 2 multimedia
Text 2 multimedia (https://text2multimedia.com) offers image generation (up to 10 images for a prompt, seems to be based on Stable...
Once upon a bot
Once upon a bot (https://onceuponabot.com) generates children’s stories with text, images and audio. FR ok if specified. Sharing url and...
Q by Vemo
Q by Vemo (https://apps.apple.com/us/app/q-by-vemo-ai/id6497067239…): a streamlined chatbot based on orally expressed queries and textual responses. LLM GPT3 (base fixed in...
Speechmatics
Speechmatics (http://speechmatics.com) transcribes audio into text, from a loaded audio file or in real time. Translates, summarizes, chapters, transforms into...
Zerobot
Zerobot (https://zerobot.ai) allows interaction with agents and their creation using chatGPT 3.5 or Groq in free version (GPT 4 in...
Notes GPT
Notes GPT (http://usenotesgpt.com) transcribes audio notes into items in a checklist and offers a summary. Generation via Mistral ai. Specify...
Morpheeus
Morpheeus (http://morpheeus.com) generates stories and offers them to read from a cloned voice after choosing the theme, age and a...
QuickTakes
QuickTakes (http://quicktakes.io): 1h30 of free recording/week of transcribed audio notes, summarized with reading guide and glossary, questions, videos to go...
Reading coach
Microsoft’s Reading Coach (https://coach.microsoft.com) supports reading, pronunciation and fluency skills. “Preview” in EN: stories generated (animal and place), library or...
Reelcraft
Reelcraft (https://reelcraft.ai) generates animated stories – video, voice, and music – from a prompt or text. Multiple templates, voice EN,...
Bedtime Ai
Bedtime Ai (https://bedtime-ai.nokk.io/stories/create) generates an audio story from a very simple prompt. Base: OpenAI API. Accepts French if specified. Sharing...
Hi Santa
A call to Santa Claus or his friends before he arrives tomorrow night while working on his English pronunciation? Hi...
Hinotes
Hinotes (https://hinotes.hidock.com) transcribes audio from a microphone or sound file. Summary, todo list or important points depending on the content....
Story Studio
Story Studio from Artflow (https://app.artflow.ai/story-studio) produces stories with editorial assistance and fully generated media. In beta integrated into Artflow (https://x.com/bertrandformet/status/1664338670193614849?s=20…)....
Educates
Educates (https://educates-ai.com) generates a text and audio course from a prompt title. 12 languages, choice of level, quiz, script, everything...
Trellis
Trellis (https://readtrellis.com) allows dialogue with works in pdf. Tools of an ebook reader, library (EN). Generates the audio of the...