Molmo, every mention
9 scenes across 1 show · ← back to Molmo
tap a year for its mentions
Latent Space 15
every year every show
Latent Space 15
Verbatim, from the transcripts: passages where Molmo comes up on Latent Space
Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
- ▶ 15:54 Luca Soldani Um, we released a multimodal model called MolMo. 2 times in the scene
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 50:10 Vik Korupati I was originally trying to do this with bounding boxes, but then Malmo came out with pointing capabilities, and I was like, pointing is a much better paradigm to, ah, to represent this.
- ▶ 54:58 unnamed speaker This is the year that vision language models became mainstream, with every model from GPT-Forty to One, to Claude Three, to Gemini One, and Two, to Llama 3.2, to Mistral's Pix-Trol, to AI-Two's Pixmo, going multimodal.
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 0:08 Vibhu Sapra Okay, so uh, Momo and Pixmo. 4 times in the scene
- ▶ 19:31 Eugene Xia It wasn't urgent for me, because there was a lot of commentary regarding text performance, and, and the, and I guess the presumption that the Momos one is inferior in text performance, that's why I kind of understood.
- ▶ 25:40 Vibhu Sapra Image to QA pairs that you train Momo on, but the, the annotation, that's what I, so like, I think that's what you were saying. 2 times in the scene
- ▶ 32:44 unnamed speaker Um, and also, one point on the MOMO one, right, if you scroll up all the way to the top, yeah, over here, the only reason why MOMO seven B and seven BD is not open data and code, it's because it's based off the Quen data. 2 times in the scene
- ▶ 41:22 unnamed speaker So like, I've seen some really cool demos on the blog posts for Momo.
- ▶ 50:36 unnamed speaker Well, thank you so much for covering the, uh, Malmo paper alongside Nathan.