Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Ok, so that's, uh, stage three, long context. Maybe just to bring this to life, um, what's the difference between before and after? Like, if you have a, uh, 40 page PDF, uh, that you fit into the window, it will just, uh, get faster results or better results. What happens?
A At the beginning, you just can't do it. Like, you know, do you pre-train as something like, Four, 8000 tokens. That's what we use for OMO three. That's what Lama is. That's about maybe eight pages. If you use like, you know, double spacing Uline kind of thing. Um, and after that we extend to about 65, um, in industry you have extension of a million token. I think Gemini recently announced like over a million token. At that point, a million token is like 10 books. Um, so you can work with extremely long amount of information. It's nice. You don't have to think about, you know, if you're building an application with this language model, you don't have to think about like, oh, of this amount of information, how the heck I'm going to extract the ones that I need to show the model. You can just give it all and the model will figure it out. So it's, it's really unlocks a lot of opportunities.
AI assessment note: “You can just give it all and the model will figure it out.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Allen, right? AI-II stands for Allen Institute for Artificial Intelligence. Uh, this, uh, you, uh, you mentioned some grants, Luca. I, uh, and I think earlier in the conversation we talked about a recent grant as well. I saw, uh, that it was a hundred and fifty two million from NSF and NVIDIA. So, uh, what is AI-II? Uh, how did it start? Has it founded at a high level?
A It was founded around 20 14 by the late Paul Allen. Initial, um, AI-II was very focused on, um, building machines that can do science, can understand science, solve science problems. Uh, that's when Semantic Scholar started as a repository of science paper. Slowly, one of the initiatives that started forming was more fundamental research around, like, how language model works, how At the time it was called natural language processing, uh, was working. Um, you had, um, teams like LNLP, uh, doing great work. Uh, since the very early, uh, since very early we worked on, we always had this idea of like, uh, not just releasing artifacts or research, but releasing the tool. Uh, back in the day we had this very, um, Very widely used for a library called LNLP, uh, that would allow you to build and customize these models.
AI assessment note: “It was founded around 20 14 by the late Paul Allen.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Thanks for that. So, uh, let's take those, uh, six modules turn by turn. So let's talk about pre-training. Uh, what did you guys do specifically for this model?
A Pre-training is very interesting. Um, the way, um, we sort of planned. So, uh, a good background is to have is that, um, pre-training, all that happens to be pre-training, we have to be very methodical in how we do it. Uh, Because, uh, first of all, it takes a long time to pre-train. I think it's standard practice among the frontier labs to try to cap your big final pre-training run to two months, uh, not more than that. Uh, but to get to like something that will not crush and burn in two months, uh, during these two months, uh, you have to do a lot of preparation around it. So we're really, um, everyone who works on pre-training is fairly methodical, um, and Just to sketch out how that works is usually you have a sense of, okay, the duration of this running is fixed. Um, the number of GPUs I have available is, will be fixed. Um, and therefore, you know, you write the fastest possible code to train this model. Uh, you have this three, you can figure out, okay, how much data can I show my model? Um, in our case, the number was like six trillion tokens. Um, Given that number, um, then we go back and we figure out, okay, where, what are the best six trillion tokens out there? Um, and the way you figure out is a combination of, like, what data you have access to. Um, you know, we want to do this with, we want to eventually release the data, so we limit ourselves to data that is, um…
AI assessment note: “in our case, the number was like six trillion tokens.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Specifically this, uh, Olmo three base, uh, seven B and 32 B. So what, what are those?
A Probably we have say five, you know, flagship, uh, checkpoints that we're putting out. Two of them are base models. That means they're, these are models before they get trained to respond to user instruction. Um, so these are really good for folks who want to take, sorry, the bulk of our compute that we spend in pre-training these models, and then they want to customize them for their use cases. So these are two base models. There's a smaller one. There's more efficient, uh, that takes about one GPU. To fine-tune for use case, and then there's a larger, that takes about one box of eight GPUs to fine-tune. And then on top of that, we have our fine-tune, our post-train models for various use cases. So there's models, there are a couple of models that are thinking models. Um, so there's almost seven B think and almost three B think. Um, these are models that, um, you know, just like a lot of the reasoner or like Pro models out there. Uh, they can spend, um, compute power, uh, inference time, um, to sort of think through a problem and solve it, and then give you an answer at the end. Um, and also we are releasing a seven B instruct model. This is a more like immediate model that gives you faster, um, responses. So it's really good for like bulk data processing or, or use cases for like, you want to have like low latency responses.
AI assessment note: “Two of them are base models. That means they're, these are models before they get trained”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q like this, uh, truly understand, uh, how the model works versus, you know, other conversations with commercial players. So before we go into the pipeline, I'd love to talk a little bit about, uh, you guys, uh, your, your backgrounds and AI too, which is a very important player in the ecosystem that, um, people may or may not have heard about. Uh, so, uh, who wants to go first?
A I sort of stumbled into this role, uh, by just picking problems that are interesting. Uh, so my background originally from Italy, moved to US for PhD. My PhD is in information retrieval. How do you build search engine to just simplify a lot? I slowly got into more and more, uh, the sort of natural language. Um, first joined after grad school, I joined Amazon. I was working on Alexa at the beginning of working on, like, The search part of Alexa, and then I got, wait, the, the actual part where the users taught me Alexa is the interesting part. So slowly move it towards that. Initially joined, um, AI to working on, um, a project called Semantic Scholar. It's still active. Um, it's a search engine for academic paper. Um, and there the interesting, interesting bits were actually, like, the interacting with users and less so the actual, uh, text of, of the papers that you were searching on. And then the, the, The way I got into LLM is, uh, in building language model is really intertwined of how AI two got into building language models. Um, that was, it all started around, um, was it November of 2022? This is around the same time, um, chat GPT got released, a bunch of, um, researchers at AI two. This is like individual contributors, um, not, uh, the There's no direction from the top as a bunch, like a very grassroots initiative to a bunch of researchers got really interested in both …
AI assessment note: “my background originally from Italy, moved to US for PhD. My PhD is in information retrieval.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And that just, that just came out, right? Like August of this year?
A Yep. The, the team has been cooking since middle of last year. But finally, we had our first release this year. Um, there's actually two releases. There was the main Asta release, and then recently we announced a partnership with Kaya, the Cancer AI Alliance, um, using some of the components in Asta to help researchers make progress on cancer research. And then there is a third branch on AI for the environment, building models that can understand, um, So I can model earth and, and can work with different signals, uh, to do prediction around the environment, uh, and so on. I'm being a little bit vague on this one because I don't know if it has been announced yet.
AI assessment note: “Yep. The, the team has been cooking since middle of last year.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q in terms of implementation of it, what's the right way to think about it, uh, for somebody that's, uh, trying to learn about the space? Is that, is, is, is one part better than the other? They need So, or do they need to exist, uh, together? Is, is, is, is, is are currently delivering more game than pre-training? What's, what's the overall kind of, um, uh, high level take?
A I think the, the way I like to think of it is, um, the pre-training phase, um, it's really like, um, A very expensive, expensive initialization of the model, right? You wanna, like, when I think of, like, oh, what do we wanna, um, what is a good final set of weights that I can pass to Nathan and the rest of the post training team is, well, I want a model that has great knowledge about the world, um, and it also can sort of, you can start seeing sparks of, of, of capabilities, uh, that you will want to model that then You know, you want to chat with, have, have great capabilities. So, um, it is a very expensive and very compute intensive, uh, way to, like, create an initial models out of, like, what is essentially, like, random parameters. Uh, but it's all about, like, yeah, let's, let's have this model have, like, a lot of knowledge of, of world facts and information. Um, and let's have it so that It can start behaving a little bit, um, like a chat model so that when we pass it to post training and you have this reinforcement learning, uh, there is some, some behavior to reinforce, uh, and to give rewards on so the model can pick it up.
AI assessment note: “pre-training phase, um, it's really like, um, A very expensive, expensive initialization”
Answered raw tape
D 3 · C 3 · P 3 · Cm 3 3.00
Q So you give it more, like, code data, for example, or math data, that kind of stuff?
A If, um, you know, you want, the model maybe cannot reason about certain math problems, uh, you do it. That's like when, when Nathan, Nathan mentioned early, um, or, you know, uh, sometimes there is, like, some leakage of, of, Uh, things that look like the test during this phase, um, you know, there is, um, uh, an uncharitable way to describe which is like, oh, they, someone is trying to cheat there by adding this data, but it's also like, it's so easy for like accidentally leaking your test data in there. We spend a lot of time making sure that doesn't happen because you want to add, um, it's really tricky balance because, uh, you want the model to start being able to solve problems like the ones that you see during tests, But you really don't want that test data to, like, accidentally leak there, otherwise you can't measure how well your model does.
AI assessment note: “the model maybe cannot reason about certain math problems, uh, you do it.”