From the text to the world
Last week, I introduced you to Echo, the game I created that features a shift, that of the human user who delegates little by little to the machine, wanted in a simple mechanism to make agency visible.
In the zoom that I then proposed in the publications of this letter 30, Make it asked a similar question from the other direction: a generative AI application that generates the complete blueprint of a physical object to make yourself, with a list of materials, programming, step-by-step tutorial and even generated photos of the finished object. I did not physically manufacture or test the two projects presented in the article, but it was one of the first times that a tangible world was presented to me with a convincing and documented physical rendering. The representation was therefore there, the object perhaps, its construction remaining to be tested.
This week, Omma extends the thread from a different direction. In a natural language instruction, the application generated an Earth-Moon-Sun system with animated and configurable orbits, a Flappy Bird-like game playable from the first generation. And for the first time, a navigable open world was roposed to me from an instruction (admittedly at the beginning with a bicycle with perpendicular wheels but corrected in the following iterations): trees, buildings, roads and paths, coherent physics and a character who pedals by stumbling over objects… and all this made in a browser with accessible code.
I described “a piece” of the world and it appeared: what struck me was not the technical performance itself, but that the simulation had become plausible enough for us to walk around and find our way around.
Three years ago, when an AI was launched, the models produced text with string responses, without anchoring in a physical reality, the evolution of which was traced in letter 25, from the site’s categories over three years. Text first, then images, then voice, then video, then interactive 3D environments. Each step seemed dizzying (the “wow” effect of letter 28), then normalized, then outdated. Omma and Make it are no exceptions in this trajectory, these applications are, I believe, concrete and accessible expressions of it today, within the reach of anyone with a browser and an idea.
What I have observed on a small scale in a browser, Yann LeCun aims at the scale of real physical systems. Three weeks ago, he announced AMI Labs (see this article from Euronews) and raised a billion dollars on a conviction he has been holding for several years: text-based language models have or will reach their limit. What he is aiming for with “world models” are systems capable of understanding the physical world from data from videos or sensors, understanding the world as animals and humans do, not simulating it in a browser, really understanding it, in order to be able to act in it. A bit like the next milestone in a trajectory that is already underway…
Like three levels of the same progression, on very different scales. Not a hierarchy, a trajectory that continues and of which we can observe here two milestones that are already accessible and a third that is financed at a billion.
So we went from text to an increasingly plausible representation of the world in three years. The sprinkler I designed in Make it will have to work in the rain with real wires, Omma’s bike only rides (for now?) in a browser. What if the distance between the two worlds was shrinking faster than we think?