AI2
company on 2 shows · 19 statements across 6 episodes
19 statements about AI2, every show
AI2's OLMo 3 32B leads Artificial Analysis's 18-point Openness Index
“It's out of 18 currently. And so we've got an openness index page, but essentially these are points. You get points for being more open across these different categories and the maximum you can achieve is 18. So AI two with their extremely open OMO three, 32 B…”
Ai2 samples 6T tokens from 10T pool for OLMo 3
“There's like a pool of about 10 trillion tokens from which we have like an algorithm also fully open source. To like sample about six trillion tokens that we use during training.”
Lambert: OLMo 3 models are the best open models outside Qwen 3
“I would say in post training where The best models that don't start with Quinn three and we're like reasonable to say that they are comparable to Quinn three, like on some benchmarks would beat them on some benchmarks. They're way ahead.”
Soldaini: Most open AI models are open weights, not open source
“Majority of models that get release I think the best term to describe them is open weights. Your Quinn, your Gemma, your Lama you know, Kimi it's what gets release is a set of weights that correspond either to the final state of model, that's the most common, …”
Lambert: AI2 coined 'reinforcement learning with verifiable rewards' replicating Llama 3
“We spent a long time to try to replicate what we thought was close to Lama three post training with multiple stages and optimizers, which is the project that like came up with the name reinforcement learning with verifiable rewards with a bunch of people.”
Lambert: Hugging Face outcompeted AI2's AllenNLP library
“It was the main competitor to Hugging Face Transformers. And they ultimately outcompeted AI two as the thing that people use for that because they had very different model and amount of support.”
Ai2 filtered OLMo 3's pre-training dataset from 300 trillion tokens
“Our initial pool was closer to 300 trillion tokens. You shrink it down till you reach your target number, and hopefully as you shrink, you only keep the best part of this.”
Ai2 fine-tuned OLMo 3 using Chinese teacher models DeepSeek-R1 and Qwen
“So in our case, we took a mix of existing data sets like Open Thoughts three and modified it, which is from Bespoke AI labs, a startup. And then we also generated a whole bunch of new data. So we ended up using a mix of teachers from like Deep Seek R one, oh f…”
Lambert: Ai2 generated billions of DeepSeek completions over a weekend
“We had a bunch of cloud credits and I, they were running out and we're behind and I just generated like as many completions as possible. So it was like a few billion completions from deep seek over the weekend.”
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Lambert: AI academics must build datasets and evals rather than papers
“If you're trying to have impact in AI right now, it's as an academic, you have to like level up out of papers to artifacts, which is models, datasets, evals. Datasets and evals are easier for people to have impact on.”
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Lambert: AI2 received early industry tip on reinforcement fine-tuning
“We got a tip from a industry lab member to do this a few months early. So we got a head start”
Soldani: Ai2 Trains OLMo Using Community Datasets and Other Models' Outputs
“We see a lot of these even in our own work of like, you know, as we iterate in the various version of Olmo it's not just like every time we collect from scratch all the data. No, the first step is like, okay, what are the cool data sources and datasets people …”
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Molmo Uses Base Model Without Instruction Tuning or Chat Template
“This is just, like, straight base model, no real instruction tuning. There's literally, like, no chat template for multi-turn. It just concatenates the messages together and, like, there's, like, go, look, good luck.”
AI2 Did Not Log Raw Audio for Pixmo Annotations
“Yeah, it was like Whisper, and I don't remember the details, and I had asked, and it's like, I don't think they actually logged the audio. I was like, oh, this would be super cool for, like, other types of multimodal, but I don't think it was in the terms of t…”
Pixmo Dataset Contains Approximately 1M Captions Across 700K Images
“They got about a million captions for 700,000 images, which I'm like, okay, that's kind of expensive.”
Molmo Reads Clocks but Fails to Generalize to Dials
“The model didn't work on clocks and then the lead was really on clocks and no models work on clocks. So they're like, we've got to make it work on clocks. One of the interesting things is that it doesn't work on dials, even though it works on clocks.”