OLMo, every mention
19 scenes, the whole family · ← back to OLMo
tap a year for its mentions
every year anyone Pratyush Maini 15Luca Soldani 9Nathan Lambert 3Vibhu Sapra 2Shawn Wang 2George Cameron 1
Verbatim, from the transcripts: the passages where OLMo comes up
⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
- ▶ 12:27 Pratyush Maini And something interesting, uh, that I was able to validate this was with the Olmo series. 3 times in the scene
- ▶ 12:39 Pratyush Maini They release the data, the models, um, and they also have a very fantastic Olmo trace now. 2 times in the scene
- ▶ 12:59 Pratyush Maini Oh yeah, I think like, honestly, Olmo Trace or the original paper which produced the Infinigram, like, I think it came out last year at Cullum, or, uh, that was one of my favorite papers that year. 2 times in the scene
- ▶ 13:39 Pratyush Maini And so you will see for the All More Two model, which was released a year ago, like say there exists an emoji, it's the wrong answer, but does not really try to self-correct itself, right? 2 times in the scene
- ▶ 13:48 Pratyush Maini On the other hand, in the 3.1, 32 B series, it has the same self correction behavior. 2 times in the scene
- ▶ 14:33 Pratyush Maini And then the more interesting thing is now we can actually trace the exact difference in the data sets for all mode two and all mode three. 2 times in the scene
- ▶ 14:48 Pratyush Maini So, with the instruct variant, they do not have any thinking data in the post-training phase, but they do mention that we are going to put, like this is the line that they write, there is some intentional addition of thinking traces in the… 2 times in the scene
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 53:49 George Cameron AI two with their extremely open OMO three, 32 B.
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 3:43 Nathan Lambert I mentioned the origin of the RLVR thing, which is, like, realistically, when you work in the open, a lot of it is trying to match what industry has done, and we're on a different path because our infrastructure is different, so some…
- ▶ 26:52 Nathan Lambert But then once it's like, once it's like this, a lot of things in the middle feel obvious, which is why I would describe one of the things that we work on for Olmo.
- ▶ 1:05:32 Shawn Wang Is this somewhere where you, like, as speaking as AI to OMO, you, you want to win? 2 times in the scene
- ▶ 1:16:23 Nathan Lambert A lot of it is just more resources, but it's like, like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.
Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
- ▶ 1:30 Luca Soldani Um, I added my own, uh, Olmo at the bottom.
- ▶ 4:55 Luca Soldani We see a lot of these even in our own work of like, you know, as we iterate in the various version of Olmo, um, it's not just like every time we collect from scratch all the data.
- ▶ 12:59 Luca Soldani Um, so Olmo, uh, the model that we built at AIT being one of them.
- ▶ 14:42 Luca Soldani So, for example, for the Olmo II model, um, a lot of our pre-trained data for the first stage of pre-training, um, was from this DCLM, uh, initiative, uh, that was led by folks, uh, ooh, a variety of institutions.
- ▶ 16:02 Luca Soldani A text-only model to a multimodal model, and we applied this recipe on top of Quen checkpoints, on top of Olmo checkpoints, as well as on top of Olmoe. 3 times in the scene
- ▶ 16:50 Luca Soldani Um, and finally, the last thing we released this year was Olmo II, um, which so far is the best state-of-the-art, uh, fully open language model. 2 times in the scene
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 3:30 Vibhu Sapra From that, they train it into a model that's based on their open source OMO models. 2 times in the scene