Molmo, every mention
16 scenes, the whole family · ← back to Molmo
tap a year for its mentions
every year anyone Vibhu Sapra 21Luca Soldani 2Vik Korupati 1Eugene Xia 1
Verbatim, from the transcripts: the passages where Molmo comes up
Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
- ▶ 15:54 Luca Soldani Um, we released a multimodal model called MolMo. 2 times in the scene
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 50:10 Vik Korupati I was originally trying to do this with bounding boxes, but then Malmo came out with pointing capabilities, and I was like, pointing is a much better paradigm to, ah, to represent this.
- ▶ 54:58 unnamed speaker This is the year that vision language models became mainstream, with every model from GPT-Forty to One, to Claude Three, to Gemini One, and Two, to Llama 3.2, to Mistral's Pix-Trol, to AI-Two's Pixmo, going multimodal.
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 0:08 Vibhu Sapra Okay, so uh, Momo and Pixmo. 4 times in the scene
- ▶ 7:31 Vibhu Sapra So there's Momo one B, which is based on their open source language model, which is a MOE one billion active parameter model that performs on par with GPT four V.
- ▶ 7:47 Vibhu Sapra Seven B, which is based on their seven B language model and Quinn seven B that outperforms GPT four V and four.
- ▶ 7:54 Vibhu Sapra Oh, then there's a large one, which is mobile.
- ▶ 14:03 Vibhu Sapra Like the Quen's 72 versus Momo's 72?
- ▶ 19:31 Eugene Xia It wasn't urgent for me, because there was a lot of commentary regarding text performance, and, and the, and I guess the presumption that the Momos one is inferior in text performance, that's why I kind of understood.
- ▶ 25:40 Vibhu Sapra Image to QA pairs that you train Momo on, but the, the annotation, that's what I, so like, I think that's what you were saying. 2 times in the scene
- ▶ 30:52 Vibhu Sapra but yeah, I mean, across the board, the seven B, the one B, the 72 B very on par with four O four V 1.5 thoughts on it and whatnot. 2 times in the scene
- ▶ 30:52 Vibhu Sapra but yeah, I mean, across the board, the seven B, the one B, the 72 B very on par with four O four V 1.5 thoughts on it and whatnot. 3 times in the scene
- ▶ 30:52 Vibhu Sapra but yeah, I mean, across the board, the seven B, the one B, the 72 B very on par with four O four V 1.5 thoughts on it and whatnot. 6 times in the scene
- ▶ 32:44 unnamed speaker Um, and also, one point on the MOMO one, right, if you scroll up all the way to the top, yeah, over here, the only reason why MOMO seven B and seven BD is not open data and code, it's because it's based off the Quen data. 2 times in the scene
- ▶ 41:22 unnamed speaker So like, I've seen some really cool demos on the blog posts for Momo.
- ▶ 50:36 unnamed speaker Well, thank you so much for covering the, uh, Malmo paper alongside Nathan.