Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I'm always curious, as an AI researcher who's like so deep into the very heart of all of this, if you zoom out, are you, are you still surprised by where we are? Like, from your perspective, are we well ahead of where you thought we would be a few years ago? Are we on track? Are we behind, possibly?
A I think it's easy to say we're on track, in hindsight. I think, if I'm being honest with myself, I think we're ahead of where I thought we could go. Um, starting work on LLMs in, in 2019 or 20 20, it's, it's kind of hard to believe, uh, the scale of everything we're doing, but also just what the models are capable of, of doing today. If you just, if you kind of looked at scaling laws back then, they were definitely pointing, uh, towards that direction. And, uh, some people really believe those deeply. I, I'm not sure if I would have bet a lot on, on that actually materializing and, and being where we are today. So, One interesting question that follows from this is where, where does that take us? If we assume the same, or if we assume the same kind of progress we've seen in the last five years, I think, yeah, this is going to be very, very cool what's going to happen in the next few years as well.
AI assessment note: “I think, if I'm being honest with myself, I think we're ahead of where”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q You use the word, um, research taste, uh, which I think is, is super interesting. What, what, what does that mean? How would you Define that, and how important is that for a researcher?
A Yeah, it's, it's very important these days, and it's quite hard to quantify, but the few things that matter is, the first one maybe is your research is not standalone. This is what I was mentioning before, but your research has to play well with everyone else's research and has to integrate, right? So let's say I have some improvement on the model, but it makes the model five percent harder to use for everyone else. This is probably not a good trade off, right? Because you're going to slow down everyone else and then their, and their research, which would then summatively slow down the durable research progress. That's the first thing. Um, the second thing is being allergic to complexity. Um, but complexity is quite subjective is in terms of what people are familiar, but still this, we have those certain, I think, budget of complexity we can use in a certain amount of like almost research risk we can accumulate before things go bad. And so being aware of that and managing that is very important. So oftentimes we don't necessarily want to use the best performance version of a research idea, but we'd rather trade off some of the performance for a slightly lower complexity version because we think that will allow us to do more and more progress in the future. So these are kind of the main two things I think around research taste.
AI assessment note: “these are kind of the main two things I think around research taste.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Since we're on the, on this topic of research and how to organize research team to be successful, let's double click on, on some of this. So you mentioned trade-offs. Presumably one kind of trade-off is short-term versus long-term. How does that work? How do you all think about that?
A This is part of what I spend a lot of time thinking about as well. Um, There's always critical path things to be done, or like this part of the model needs improving, or we know this part of the model is suboptimal. So we invest quite a lot in just fixing those immediate things. There's a few reasons for that. The first one is we know this will make the model better, so it's a fairly safe bet. But also we know that things that don't look quite good or quite perfect often tend to have issues later. Either when you scale up or when the model just becomes more and more powerful. And so actually really having, being very diligent about tackling those and fixing those is, is really important. So, so that's kind of the, the first part. The second part is slightly more exploratory research. So ideas that could land in the next version of Gemini or, or the version after that, that had maybe a bit, a bigger effect on, on the model performance, but aren't quite validated. How we balance these is, I don't think I have a very clear answer. It's also a bit periodical. So When we're doing a scale up, for example, there's often more, slightly more exploratory research because there's nothing right now that needs to be fixed in parallel. But just before we are ready to scale up a new architecture or a new model, it's very much like let's de-risk the last pieces. It's very execution focused.
AI assessment note: “How we balance these is, I don't think I have a very clear answer.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What does, what does it tell us in terms of where we are in AI progress? What sounds from afar as in sort of turning some knobs gives us such a leap? What does that mean in terms of what we can expect going forward?
A There's two things. The first one is, it's still remarkable how much progress we're able to achieve in this way, and it's not really slowing down. There's so many of these knobs and so many improvements that we find on a day-to-day basis. Yeah, almost on a day-to-day basis that make the model better. So that's the first point. The second point is, We're not really building a model anymore. I think we're really building a system at this point. Um, people have sometimes this view that we're just training a neural network architecture and that's it. But it's, it's really the entire system around the network as well that, that we're building collectively. And so that's the second part.
AI assessment note: “it's still remarkable how much progress we're able to achieve in this way”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q people. In the case of Google, I guess the vertical integration, uh, that tweet from Aurel that I was mentioning got, uh, uh, retweeted, a quote tweeted by, uh, Demi Sasabes, and he was saying that the, The, the real, real secret was a combination of research and engineering and infra. So is that, is that the secret sauce at Google, the fact that you guys do the whole stack?
A It definitely helps. I think it's, it's an important part. Research versus engineering is also interesting. I think over time that boundary has blurred quite a lot because we're working on these very large systems now. Research really looks like engineering and, and vice versa. And I think that's, that's a mindset that has really evolved over the last few years at DeepMind, especially where maybe there was a bit more of the traditional research mindset before, and now with Gemini, it's really more about research engineering. The infrastructure part is also very important. We are building this super complex system, so having infrastructure that's reliable, that works, that's scalable, is key in terms of not slowing the research engineering down.
AI assessment note: “It definitely helps. I think it's, it's an important part.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Presumably there is a cost aspect to this. Does being natively multimodal mean you're more expensive from a token perspective?
A Yeah, this is a really good question. There's kind of two costs to this. I would say that the benefits largely outweigh the cost here, and this is why we train these models. But the first cost is maybe less obvious to people, but it's this complexity cost and this research Uh, but I was talking about because you're doing a lot of more things, and especially different modalities interact in some ways, this, this can interact with different parts of the research and has, and it has a complexity cost. So we have to spend time thinking about these things. The second cost is yes, uh, images are often, uh, larger, uh, in terms of input size than, than, than pure text. And so the, the actual computational cost, um, If you do it naively is higher, but of course, then there's interesting research to be done on how you make these things efficient.
AI assessment note: “images are often, uh, larger, uh, in terms of input size than, than, than pure text”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q For any, uh, student or like, yeah, PhD student listening to this, uh, if they want to become you in a, in a few years, what problems do you think they should think about or focus on that's not, you know, like a year or two out, but like more interesting sort of a few years out?
A One thing that's, that's becoming increasingly important is being able to do research, but being aware of the system side of things. So we're building these fairly complicated systems now. So being able to understand how the stack works all the way down from TPUs to research is kind of a superpower, because then you're able to kind of find these gaps in between different layers that other people weren't necessarily able to see, but also to reason through the implication of your research idea All the way down to, to the, to the TPU stack. And people that can do that well, I think have, uh, have a lot of impact in general. So in terms of like specialization, it's really thinking about this research engineering and systems aspects of the model research and not just the pure model architecture research. That's one. I think personally, I still have a lot of interest in kind of this retrieval research as well that we started with retro And I think it wasn't quite ripe until now, but the things are changing and then I don't, I just think it's, it's not unreasonable to think in the next few years, something like that might actually become viable for, for a leading model like general.
AI assessment note: “thinking about this research engineering and systems aspects of the model research”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q just came out, uh, at the same time, like two hours ago, literally before we were recording this, GBT, 5.2 came out. What do you make Uh, of that from your perspective and how do you think that plays out? Is anybody going to break out or, uh, effectively the, uh, industry is going to continue with like the handful of top labs plus some new labs that are appearing?
A For the first question, there's definitely similarities between what the different labs work on. I think the, the, the base technologies are kind of similar. I might be surprised if, if we weren't all training. Transformer-like models, for example, in terms of the architecture side. But then there's definitely specialization, I think, happening on top of that and different, like, maybe tree or branches in the tree of research that are being explored and exploited by the, by the different companies. I think historically, for example, DeepMind has, and still, I think on the vision and multimodal side, we've been actually really, really strong. And that continues to be the case today and then shows in both how people use the model, but also in the benchmarks. Of course. And then, yeah, there's things like reasoning, et cetera. OpenAI came up with the first model, but we also had a strand of research on that. So there's similarities, but it's not exactly, exactly the same, I would say. For the second question, I don't know if I have a good answer. One thing that's clear is to make progress on a model like Gemini today, you do need a very large team and a lot of, a lot of resources. Now, that doesn't necessarily mean that what we're doing today is Optimal in any form and some disruptive research could definitely come along and allow a smaller team to actually take over in some form.…
AI assessment note: “For the first question, there's definitely similarities between what the different labs work on.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q And what did you do at, uh, first and how did that evolve, uh, to, uh, being one of the pre-training leads on Gemini III?
A Yes, uh, it's, it's, uh, at the beginning, having joined DeepMind and DeepMind being known for RL, um, the first project I, I managed to, to work on, or decided to work on was, uh, something on, on the RL side. So, um, specifically we're training, uh, some, Unsupervised network to, to learn key points on, on Atari environments and try to, to get the agent to, to play Atari, right? Um, so I did this for about six months maybe. It wasn't enough for, in the sense, I didn't like the, the synthetic aspect of this. Um, I, I always wanted to, to work more on, on, on real world data and have more of a real world effect. I think in general, I like to build things and, and build things that work. I don't really like the, the academic pure research part, and so that kind of drove me, um, to, to start working on, on, uh, representation. So creating these, or training these neural networks that have good representations to do different tasks. And, and one funny anecdote here is something I, I tell a lot of people on my team, but, um, the first effort I joined on this was called a representation learning from real world data. And, uh, at the time we had to add this from real world data Uh, to the, to the name of the project, because people would assume otherwise it would be synthetic environments or, or synthetic data. Um, and, and that definitely has shifted, uh, completely, uh, since then.…
AI assessment note: “the first project I, I managed to, to work on... was, uh, something on, on the RL side.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q conversation, but, uh, like another big question and theme, uh, seems to be indeed, um, how can models learn from less data, which, uh, I think what is what you were alluding to talking about a data limited regime, again, whether at, uh, DeepMind or, or in general, uh, are you seeing interesting approaches to use the, the famous analogy a model can learn like a, like a child does?
A Just to maybe clarify what I said earlier, um, in a data limited regime, I didn't necessarily mean with less data, but rather with a finite amount of data. So the paradigm shift is more from, like, we have infinite data to we have a finite amount of data. The second point is, in some sense, model architecture research is exactly what you mentioned. So when, when you make an improvement on the model architecture side, what it typically means is you get a better results, a better result if you use the same amount of data to train the model. But equivalently, you could get the same result as the previous model by training on less data. So that's kind of the first aspect of that. But it is true in terms of the volume of data needed today. We're still orders of magnitude higher than what the human has available to. Of course, there's the whole evolution process as well, which I find these high level discussions quite hard to, to understand or follow because you have to make so many assumptions to, to, to convert that amount of data into what is today's pre-training data. But At least at first order, it does seem like we're using a lot more data than humans do.
AI assessment note: “in some sense, model architecture research is exactly what you mentioned.”
Partly raw tape
D 3 · C 5 · P 4 · Cm 4 4.00
Q Gemini three was trained on TPUs, right? Not on NVIDIA chips. So it's truly, truly integrated. Okay. So I'd love to, uh, do a deep dive on, uh, Gemini III, but before we do that, uh, let, let's talk about you a little bit. So you are the pre-training lead, uh, on Gemini III. What does that mean? And then, uh, let's go into your, your background and your story.
A I'm one of the, the Gemini pre-training leads. So what this entails It's a mix of different things. So part of my job is, is actual research. So trying to make the models better, but these days it's, it's less running experiments myself, but to help design experiments and then review results with people on the team. So that's the first part. The second part, which is quite fun, is more of the coordination and integration. So it's a fairly large team at this point. Um, it's a bit hard to quantify exactly, but maybe a 152 hundred people I work on a day-to-day on the pre-training side between data, model, infrastructure, evals, and so coordinating the work of all of these people into something that we can build together is actually quite complicated and takes quite a bit of time, especially time to do well. To me, this is super important because actually being able to get progress out of everyone is, is really what makes us make the most progress rather than enabling maybe one or two or a small group of 10 people to To run ahead of everyone else. That might work for a short period of time, but over longer periods of time, what's really been successful for us is, is being able to integrate the work from, from many, many people.
AI assessment note: “I'm one of the, the Gemini pre-training leads. So what this entails”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q And in, uh, anti-gravity, there's a whole, um, vibe coding aspect. Truly vibes in that, like, you don't even really see what happens when you ask. Is vibes, same question, is that a pre-training thing? Is that just a post-training thing? How do you build vibes into a model?
A Yeah, this is interesting. I think you can probably ask five different researchers and you'll get five different answers. There's also this notion of large model feel. People call this, especially, I think, GPT-Four.five historically had some of this, presumably, where larger models maybe feel differently. I wouldn't actually just... I wouldn't put it in these terms specifically, but I think Vibes comes down to this. And actually, pre-training probably plays a larger role today in some of that and how the model feels and then generally then post-training. I think this is, yeah, this is in general for, for vibe coding specifically, I think that's, that's maybe more of an RL scaling and post training thing where, where you can actually get quite a lot of data and train them all to do that really well.
AI assessment note: “for vibe coding specifically, I think that's maybe more of an RL scaling and post training”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q Great. And then you worked on a retro, right? Do you want to talk about that?
A Yeah. So after that, uh, we started working on, on scaling up LLMs and LLMs in general. So, so we started this work first on, on Gopher, which is, I think the, the first deep mind, uh, LLM paper that was published. So already at that point, it was a team maybe of 10, 12 people. So already at that point, it was pretty clear you needed, uh, you couldn't just do that research. Uh, on your own. And this is really where, where I started doing pre-training and pre-training at scale and, and yeah, develop my research taste, but also, uh, yeah, what I enjoy about this. Um, so we, we, we trained the first, uh, dense, uh, transformer model. I think there was two hundred eighty billion, uh, parameters. I think three hundred billion tokens at that time and, and, and trained that. And, uh, we were, Definitely, we would definitely not do things like we were doing them back in the day, but it was, it was great and a very fun learning experience. After that, there were kind of two, two projects that emerged. The first one was Chinchilla and the second one, uh, Retro. So in Chinchilla, we were kind of, uh, we, we were re-examining how you should scale the model size and how you should scale the data, um, especially from a computer. Training compute optimal perspective. So, so the question is, you have a fixed amount of training compute. How do you train the best possible model? Should you incre…
AI assessment note: “After that, there were kind of two, two projects that emerged. The first one was Chinchilla and the second one, uh, Retro.”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q you are in the world of Gemini three, which is like massive amounts of data and training in a very long, Context windows. Do you think that, uh, this paradigm of having, again, larger models, large context windows, uh, effectively obviates the need for kind of drag and search, uh, and that everything gets folded into the model? I mean, obviously there's a corporate data part, but, uh, in general?
A There's some, some interesting questions here. So, so first of all, I think Retro was really about retrieving information rather than storing it, not necessarily about making models smaller. So it's about how we can use the model to do more reasoning already in a pre-training sense of reasoning rather than just, just store, store the knowledge. So, so this is still very much, um, the aspect today. The interesting part is, um, the, the, the iteration cycle maybe of pre-training, uh, used to be a lot slower than, than that of post-training until, until fairly recently. And so making these large changes on the pre-training side is quite costly in terms of risk and how long it takes. And then you have approaches like RAG or SEARCH, which you can do during post-training and iterate much more quickly on, which gives very strong performance as well. I think deep down, I do believe that the long-term answer is to learn this differentiable end-to-end way, which means probably doing pre-training or whatever that looks like in the future, Learn to retrieve as part of the training and learn how to do search as part of the large part of training. And I think that that's kind of RL scaling maybe starts that process, but I think there's a lot more to do also on the architecture side. But this is something that we'll see in the next few years and not immediately, I would say. The one thing I w…
AI assessment note: “I do believe that the long-term answer is to learn this differentiable end-to-end way”