Molmo 2, the small vision model that sees (almost) everything
Focus of the letter 16
One of you could have said that this week there are only number 10s in the team, and I would have fully agreed with him! : )
Another complicated choice for this focus… but since I have to make one, it will be Molmo 2 even if I could have chosen each of the publications with Incredible and Everyday which make accessible the sometimes complex automations in dedicated applications, the isolation of the sounds of SAM audio, the bank of web pages with which you can converse from Browsewiki or the duo-biiiiip like of Google translate with its generation of situations… Really all of them but only one in the end.
So why Molmo 2?
First of all, because even if only two tests are published in the article, I spent a lot of time there between a certain fascination with the model’s performance and the game trying to “trap” it… which I almost didn’t manage to do, apart from a die account two of a flow of cars that I had to recount on my side several times ;).
Then for the evolution that this model seems to me to bring: I have already published several applications that analyze images, videos or a stream from a screenshot in real time, Molmo 2 brings in my opinion an additional precision, especially in the identification, counting and tracking of objects.
Finally, Molmo 2 is available as open-source and it is lightweight, so we can imagine that it will be included in future applications and that it can be found combined with other features. Maybe the ambition to create models of “perception” of the world, put forward by some teams of researchers, is not so far away?