Mercury 2: when texts emerge as images
For the past three years, and we’ve already discussed it several times in our focuses, we’ve been used to seeing image or language models progress with ever more parameters, more data, more features… An evolution, often linked to a commercial race for proprietary models, whose stages we follow as we follow a sports race. I think that this week Mercury 2 is not quite competing in this same competition.
Where models generate token-by-token text, word by word, and each choice conditioning the next, Mercury uses a new diffusion approach: text emerges globally from the given instruction, by successive “improvements” rather than by step-by-step accumulation (a video capture is available on the site).
The researchers and engineers of “Inception”, the company that develops Mercury (one can also wonder about the choice of this name: the beginning, the incipio, and/or a reference to Nolan’s film?), have managed to apply the same principle to text generation as image generation.
To understand more precisely what Mercury changes, the article in the Financial Times spotted by Hervé Allesant (present in the shared readings) which explains the functioning of a classic model followed by the description of Mercury on the Inception website seems to me to be a good entry.
After my first tests this week, which certainly revealed a speed and an anchoring to the instruction that seemed to me more important, it is rather what this approach says implicitly about the limits of current language models that seems interesting to me. If generating token by token mechanically creates biases and drifts, then perhaps a lot of what we attribute to the performance or “errors” of current models is only the consequence of a technical choice, and in the end perhaps not a fatality. The speed deported to the model and more to the hardware opens up new perspectives, both for applications on less powerful hardware and for energy consumption…
It may not make diffusion models like Mercury 2 the models of tomorrow, it reminds us that we are still at the beginning, that the foundations are not set in stone, and that the next interesting step may not necessarily follow from the previous ones. Arthur Sarazin says it in his own way in this week’s readings with the analogy of the cyclist and his bike: understanding how the machine works will certainly remain the best and only way to decide if and how to use it…