One day, one generative AI tool

Focus

Video generation: a geography that is taking shape?

Focus of the letter 35

Happyhorse is Alibaba’s latest video generation model. You enter an instruction, add a reference image, the video arrives, with sound, faithful to the original request. Five formats, up to fifteen seconds, possible editing by instruction afterwards, it’s efficient.

Six weeks ago, OpenAI’s closure of Sora was the subject of a reading shared in letter 30. The announcement was sober, almost brutal: a “we say goodbye to the Sora application”, posted on X without detailed explanation. Six months after a launch with great fanfare, with a teaser, a waiting list and the first spectacular visuals, the platform closed its doors. Not because the technology was bad, but because the business model did not hold.

To understand why, we need to look at what video generation really costs. A video of a few seconds mobilizes considerably more computing resources (and energy…) compared to text or image, at least 25 frames per second, so it is at least 25 times more computation than the generation of a still image: each image must be generated, consistent with the previous one.

The losses related to Sora would have reached 15 million dollars per day for OpenAI. This figure should of course be taken with caution, without official confirmation from OpenAI, but the order of magnitude put forward is enormous. Integrating Sora without a consumption limit in a monthly subscription was tantamount to offering a service whose real cost was not covered by the price paid. OpenAI has decided: to withdraw Sora and refocus on code and productivity tools for companies, where the perceived value justifies the rates charged and certainly also where its direct competitors are advancing. With a wave of video content generated in a few months on social networks and the media…

Meanwhile, Alibaba is releasing Happyhorse. ByteDance has Seedance. Kuaishou owns Kling. These Chinese publishers apparently do not ask themselves the same questions of short-term profitability, or at least not in the same way. These companies continue to invest in a segment that others choose to leave, not because of technological weakness, the models are comparable and sometimes superior, but because of market orientation.

We could see companies less subject to pressure from short-term investors and growth logics that do not have to demonstrate their immediate profitability. But it may also be simpler than that: these platforms may have understood that the value is not in the technological demonstration but in the installation of habit, perhaps to ensure that the video generated becomes a reflex. We would almost still find here the passage from the “wow” to the “ah yes” that we were talking about in letter 28

This is where the divergence becomes interesting, beyond the companies themselves. If generative video for the general public continues to develop on the Chinese platform side while OpenAI and others are withdrawing from it, it will not be without effect on who defines uses, formats, and expectations. Kling and Seedance already have millions of users. These are the platforms that teach creators what can be done and what works. They are the ones who install the reflexes. They are also the ones who collect the usage data that makes it possible to improve the models, and this advantage will be difficult to make up.

The question is not whether it’s worrying or not. I think it’s more precise: will OpenAI return to this field once computing costs have fallen sufficiently? Will Western players close the gap, Google with Veo, Runway, Adobe with Firefly? Or are we seeing an economic geography of who produces what and for whom, with creators who will adopt the tools available, regardless of their origin?

The video generated was presented as the next big breakthrough in creation, it has mainly produced mass content for the moment. It may be, but the breakthroughs also reveal who has the means to stay in the field and who chooses to leave and with what objectives.