Generating everything from a single place: a godsend or the disappearance of the creative gesture?
Focus of the letter 26
Before getting on my bike, I would like to point out a little novelty on the site: I have grouped all the focuses of our letters on a dedicated page. This will allow us to find them more easily and for me to quote them if necessary as I will do today ;).
A little break on our road this week with two Google apps, Gemini Music and Flow.
Even if I chose this focus first and foremost for the quality of the models used and the final renderings, it is also for their resonances with other focuses: last week with the quick analysis of the site’s database and the categories over three years, in letter 8 on the plausibility of video generations or letter 13 on the rapid evolution of models.
Gemini Music and Flow condense all these questions a little. The “wow” effect is still there, model after model, in an almost incessant flow of new products. The race of the big companies is such that every two or three weeks brings its share of new steps. Some are more marketing than novelty, but Google, by the reach that Gemini has, has in my opinion crossed a threshold that had been emerging for several months: that of almost complete multimodal integration, that of the centralization of possible generations in the same place.
We are no longer just “generating an image” or “producing a text”: in a few instructions, a video, a music, a voice, a coherent visual style, all in the same ecosystem. The question is no longer “does it work?” but “what does it change for us, users?”. And this is where it gets a bit dizzying: the tools seem to me to go beyond the uses even before these uses have had time to stabilize.
We could be happy about it, but I have the impression that over the course of my tests something is moving. At the beginning, each tool required learning as a form of taming, you had to understand the logics and the limits and these steps had, I think, a value: they forced you to think about what you were doing.
Today, these steps seem to me to be disappearing little by little. We describe, we generate, we adjust at the margin.
The technical barrier becomes almost non-existent, which is certainly good news for accessibility, but isn’t this also the disappearance of the “friction”, of that moment when when the tool does not offer to do what you want you have to circumvent, adapt, invent.
The complete multimodal integration therefore raises a question that I find more interesting than the “wow” effect, than the “bluffing” result, it may finally question creativity. Do we create or do we delegate more by starting a production and reproduction machine?
As usual, I don’t have a clear-cut answer to give, I’m just sharing with you my thoughts of the moment. I still observe, in the times spent testing or using on a daily basis, that the richest are not the ones where the application has succeeded the first time. These are the ones where something didn’t give the expected result, where I had to understand why, where the unexpected result opened up a path that I wouldn’t have looked for. I invite you to take a little detour through Arthur Sarazin’s post in the shared readings: the challenge is not to discuss the power of these tools, we already know that they are, it’s to know if we remain, in this increasingly fluid ecosystem, tireless learners with the possible support of the machine, the authors or the publishers of what the machine offers.
This is certainly the skill we need to acquire and maintain: not to master the tools but to know precisely what we want to recognize when the machine takes us elsewhere and to decide in full knowledge of the facts whether we follow the path it will take us…
At the end of the break, we get back on our saddles and we go to analyze and maybe explore the proposed paths!