Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Great. I'd love to talk about you and, and your background. Um, tell us your, your story in a few minutes. Like, how did you come to do this What's work and what was your journey to AI and then your journey to, to Google DeepMind?
A So I, I did my PhD at University of Amsterdam on machine learning and, and mostly on the, on the, like the language model side and text and, and search and retrieval. And then, uh, I think what kind of like pushed me toward trying really hard to be on the, uh, like on the, on the mainstream and be part of this group that are like, you know, hustling to, to make like really good progress. I did a few internships, like back in 2016 and 2017. And, and, and the funny story is, um, I, I did an internship at Google Brain in 20, like early 20 17. And then it was amazing. It was just like, you know, I, I went to, to, to this team. They were working on like LSTMs for, you know, like a summarization. Summarization was actually one of the most interesting problems at that time. I was like amazed. I was like, so this is so good. I, I really, I just want to keep doing this for the rest of my life. You know, this is it. And then I got, uh, I got a return offer to go back and then do another internship at the end of the, the same year. The recruiter told me that, oh, you know, there's this team that they just published a paper. Maybe you've heard about it, like Transformer. And then they're looking for an intern. And I haven't had a chat with, I remembered I had, I had a chat with Lukash Kaiser. And then Lukash was talking to me and saying like, oh yeah, like we have this idea of building lik…
AI assessment note: “So I, I did my PhD at University of Amsterdam on machine learning”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q So the same problem as reinforcement learning, right? Like the second you start veering away from math and code, uh, you start getting into like a very messy territory. Is model collapse, uh, one of the issues to think about or is that orthogonal?
A Model collapse is definitely a risk, right? And, um, I would say model collapse mainly happens when you have a loop that is Completely closed. Right. And, um, if you don't have any outside signal and just the model, for example, talking to itself or, or operating in a, in a very, uh, like a, like a restricted environment, um, there's a good chance that your model can access, but, uh, but if you have a strong verifier or some sort of a, like a real reward signal that anchor this, this kind of like, um, like signals that is coming from like AI generated data, for example, it can be quite powerful. I think like the key Here is to stay grounded, uh, uh, to, to something real. And then you can most likely avoid like, you know, things like model collapse. And, um, I, but yeah, I mean, again, it's, it's a risk, but it's not definitely like a major rocket.
AI assessment note: “Model collapse is definitely a risk, right?”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q A couple of weeks ago, we had a fun conversation with Karina Hong of Axiom Math, and we talked about formal verification. Is that a promising area from your perspective? Is something like formal verification, what would enable you to make sure that the, you know, the improvement loop keeps continuing?
A In my opinion, formal verification is, uh, one of the most powerful, uh, like keys to To, to enable like self-improvement, but it's not beaky. And if you think about it, like for mass code, logic, um, it's great. You can, you can run a proof. It either checks out or not. If you go to other domains that are a little bit messier, uh, like for example, you cannot write a formal proof that if a doctor's recommendation is good, right? Like, so, so it's not hard. It's not easy to, to have, like to extend this formal verification to all the domains, uh, in real world. But one question that is actually an interesting question. Which is, um, like, uh, very relevant to formal verification is how can we, uh, look at these methods and, and like formal verification and build that kind of tight and honest feedback loop for the messy part of the world. I think that that's like very inspiring to, to build like, like on top of these, like formal verification methods to, to extend to domains that like not easy to verify easily, uh, uh, but. Uh, but, but you need some sort of clean and tight feedback loop to be able to make progress.
AI assessment note: “formal verification is, uh, one of the most powerful, uh, like keys to To, to enable like self-improvement”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q for, uh, seconds. The big theme of the last year has been, uh, the, um, acceleration of post training in addition to Pre-training. So the, the whole, um, you know, reinforcement learning aspect, um, of, of things. Where do you expect gains to come from in the next, uh, few months or, or, or year? Is that more post-training? Is that more pre-training? Is that both? Is that something else?
A The answer to this question really depends on when you actually ask this question. And like, it's obvious that like, you know, we're going to be having a bit of a swing back and forth between pre-training and post-training. At the end of the day, I want to say that, you know, pre-training is still the foundation and like, you can never post-train your way out of a week-based model. But right now the current, like the return on post-training is really strong. And I started working on post-training myself, like, like a few months ago, like Gemini post-training, like mostly coding and agentic. I can see how a brilliant small idea can make a model like 10 X better, for example, in terms of behavior. Uh, at a fraction of the cost of the pre-training, right? This is again, like, you know, we can, we can see how post-training is like, like the place to make a lot of impact and improve these models. But on the other hand, um, like I, I know at different companies, it's also the case, but at GDM, a lot of exciting recent work is going into the pre-training side and like new, new recipe, new ideas. And, um, uh, and, and I would say like, you know, the work that we're doing on the pre-training is, um, Going to unlock a lot of downstream possibilities. Post training is just like a different mode of operation. It's like also super interesting for me because I'm, again, like a little bit lik…
AI assessment note: “we're going to be having a bit of a swing back and forth between pre-training and post-training”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q All right. Last couple of hot takes. What do you think people are too confident about?
A So people think that, you know, the pushing the technical side is, uh, is sufficient. That if we just like get a model that is a smarter, um, everything is going to follow. And in my opinion, like a version of AI that is like, you know, really, really brilliant at, um, like technical problems, but it has a, like a blind spot about, you know, everything else. And that version is not going to be able to actually create meaningful progress in, in the world. And the, the fact that, you know, people kind of assume and, and confident about like, like they're confident about this, that, that kind of like, you know, everything is going to, everything else is going to follow or, uh, or, or just everything else is just like a small list. Um, I think it's wrong. Like, you know, we have governance, we have like, you know, regulation, we have social trust, we have like, for example, distribution of access and the benefit and like in the world from, for, for this technology. And even the, like, you know, the institutional capacity to, to kind of like absorb and adapt this technology is, is just like, this is something that maybe we don't have enough, um, like attention to. And, uh, and, and these are not really soft problems. If not harder than the technical part, they're, they're really hard. And like the pace of technical progress is definitely like, like, um, currently running ahead of th…
AI assessment note: “people think that, you know, the pushing the technical side is, uh, is sufficient.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q Great. What is one idea in AI research right now that is underrated?
A Something that is underrated, like you mentioned, continued learning. I think this is, this is definitely underrated. As I said, you know, like sometimes the problem stays in the exploration mode, uh, until we are confident about something and then it goes to the exploitation mode. I, I think we are past the time that we really had to push this to the exploitation. So maybe like foundation models are, are essentially right now, like frozen in time and like when the, like the training ends. Right. And then everything is like built on top of this frozen model, like in a rag pipeline and, and fine tuning workflows and retrieval system. And all these like elaborate, like infrastructure is like, um, all based on this assumption that these models are frozen. And, uh, it's a bit of a, like a too much of a strong problem. There's an assumption to make. And, and, and I think like we are going to like get to the point that we need to change these assumptions and like, maybe like we need to think about it a little bit more actively and have, um, pushing it Toward like, you know, something that we actually appreciate to productionization. And maybe it's a little bit like underrated right now, like the continued learning.
AI assessment note: “Something that is underrated, like you mentioned, continued learning.”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q So people may have seen or heard about, uh, Karpathy's auto research project a few weeks ago. Is that an example Presumably, reasonably narrow to make it work. Is that an example of a self recursive loop?
A That is definitely. And, uh, I think that was one of the early examples of, uh, like seeing these models actually doing something super sensible on the research side. So we've been seeing them like doing a lot of good work on improving the engineering part of the development loop. Uh, but on the research side, which like, you know, you think about, okay, no, maybe Some sort of, you know, gut feeling or intuition is needed. And like a researcher with like a long time of like, you know, playing with these models and experience can do this, but, but not necessarily, you know, like a model. Uh, I think we're seeing the sign that, okay, you know, maybe that basically that kind of like golden, um, part of the recipe, a successful recipe that mostly coming from, um, like, um, intuition of a good researcher is coming to To kind of these development loops by, by these models. And, uh, it's a bit hard to think about, okay, you know, does it mean that we can replace like every genius researcher with these bottles? Like, like very soon, uh, maybe, uh, and I don't know like how, how soon, but this, this is definitely a sign of, um, something that we, we kind of doubt that like, you know, a few years ago, you know, we couldn't believe that, well, this is gonna happen that early, uh, which is very exciting.
AI assessment note: “That is definitely.”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q Uh, all right. Maybe that's more of a macro question as I think about, um, you know, where, where the value lands in this ecosystem, but if AI just keeps creating itself, then does, is data still needed in that equation or is that all compute?
A Concept of data is a little bit like broader than just like, you know, tokens, right? Like, and, and if you think about data as, um, whatever that the model can Get signal from either it is like predicting the next token in raw text, which we kind of like, you know, using print training or super complex environment that the model interacts with and then get signal. Um, this is something that basically like we can, we can like refer to it as like data. Right. And it's not like data or, or, uh, like the value of like having good data or working on data is going to disappear and compute. Is going to become, like, the only things. At the end of the day, I think, like, the work that we're doing on the data side most likely is going to shift toward building environments or, or making sure that these models can interact with, with physical words, and then it becomes more of a problem of, okay, how can we, like, provide more grounding for these models? They are good at, like, improving themselves, but as long as, like, you know, I, I have, I exposed them to real work data, right? Like, and, and like, you know, real work environment. So, so providing data becomes more about, okay, how can I give, uh, access to this specific, um, like model to something that, you know, we never had, for example, like I, again, like something came to my mind, which is like, again, like a little bit sci-fi…
AI assessment note: “it's not like data or... the value of like having good data or working on data is going to disappear”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q Is that a concept of world model that's built into the image representation?
A Exactly, exactly. So, so basically you, you want a world model basically, you know, like these models to be also like a world model. So you, you want these models to know about the world. It's a good chance that, that you can actually teach your model about the world just by presenting text to it, but it's just not efficient. And, and, and a good shortcut would be to bring multimodality into this. And the best way of learning about Uh, modality is learning how to generate that, right? Like, so we, we got to this point that, okay, you know, um, uh, we've been, we've been having Gemini generating like images, uh, from Gemini one. So we like, like basically like Gemini was multimodal from day one. And, and the reason that we kind of first like released the image generation at like 2.5, instead of, you know, like Gemini one, Gemini 1.5, Gemini two, uh, was that it was not great. And then like, it really needed a push. And then we figured out that, okay, you know how to push this without like, you know, introducing any regression to other capabilities that the model has and, you know, like bring all of these natively into the, into like this, this model. And that was like one side that was like super interesting for me, like not sad views, but, but, but it's really hard to see positive transfer. So it turned out to be a really, really good model, but it was like really hard to see t…
AI assessment note: “Exactly, exactly. So, so basically you, you want a world model basically”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q science fiction, except it seems to be becoming reality very quickly, which is this concept of, um, recursive self-improvement or, or RSI. That seems to be, um, what, what a lot of people are talking about. I think, uh, ICLR is coming up in a, in a few weeks and there's a bunch of papers. Focus on that. So what, what is that? What is recursive self-improvement as a concept?
A It's actually interesting because you, you referred to that as something that like looked like a bit of a sci-fi situation where these models are actually improving themselves. And, and that's true because a few years ago when you were, you wanted to talk about this, you could just write a perspective paper at a, at a conference and like, you know, talk about it at super high level. But if we go, uh, And, and check out what is happening right now. Like to a really good extent happening. Like most of the, the, the, like, and, and, and it's somehow like most of the people don't realize that this is like already, uh, happening, especially over the past few months. In almost every lab, um, the new generation of the models are built heavily using the previous generation of the models. I think that's basically like, like the case, um, again, everywhere. Uh, and, uh, it's not fully automatic yet. Uh, but the direction is like super clear and it's like easy to imagine that, you know, like we're going to get to a situation with full automation. These models are going to improve themselves and keep learning from the world. And again, it has like relation with other concepts, like continual learning and other, other concepts that we are, um, still not yet to, to the, to the most advanced, uh, like point of it. But if someone comes and say that, oh, you know, uh, like, uh, I, I have an ide…
AI assessment note: “new generation of the models are built heavily using the previous generation of the models”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Fascinating. Does part of this contribute to efficiency? So, especially Nanobanana II, you have the, the flash aspect of this. So you're able to create amazing images very fast and, uh, apparently, I mean, seemingly very Efficiently. So what's, what's behind the scenes that, is that what you described? Is that MOE? How are you able to do that?
A First of all, I was, I was involved in the original Nano Banana, Nano Banana Pro, and then the last version I get, like, you know, because I, I jumped on the post-training encoding and agent, and I find it exciting. This one, like this, the team actually shipped it. But if I want to say, like, like super high level that, you know, what is exactly the things that makes the model, like, uh, like faster and more efficient, Part of it is just the size of the model. So like Nona Banana was pro size and this one is just like flash. So definitely the, the, the parameter size, like, you know, configuration of the MOE and stuff. The other one was, uh, people actually spent quite a lot of time on figuring out, like, you know, uh, um, like nailing down, like distillation recipe, uh, both on the side of like knowledge and like, you know, other like things that basically you, you kind of like need to distill to something like a, a process that is like, you know, lighter than the full process. Surprisingly, a lot of infra work for serving. So we have like really, really, really brilliant people that they're like, you know, serving engineers. And it's kind of impressive that like you sit on your desk and then Uh, like they, they come and they say, oh, by the way, but casually I make a model 10 X faster. I was like, you know, like, it's just like, and, and they just kind of saying it like, you…
AI assessment note: “what is exactly the things that makes the model, like, uh, like faster and more efficient”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Fascinating. Another fundamentally important contribution to the, to the field that, uh, you did was the visual transformer paper in 2022. So the paper was called, and the image is worse, a 16 by 16 words, transformers for image recognition at scale. Do you want to walk us through what that was?
A That's also, there's also a funny story for that. I got into vision and multimodality with that paper. So I've, I've never worked on, on, on like any vision problem. It was mostly because I was sitting next to people who are working on vision. So like my desk was like next to people who were working on vision. And that was the reason that I got interested because I was just talking to them. I was like, oh, this is actually interesting. And, um, and then, uh, and then I remember that at that time I was like working on, um, uh, on, uh, like externally, we call it palm, palm paper with like, you know, uh, and other folks. And I was like, why we have Uh, four hundred billion parameter language models, but the biggest model that we have on the vision side is just like maybe a hundred million, like a rest that, like what, why, like, like why there's no benefit of a scale. I started looking into this with folks on like, okay, like maybe, you know, like there's something in, in transformer that actually kind of like, you know, make it a scalable and then maybe, you know, like we, we can move away from convolution to try this. And, and at the end of the day, I don't want to say that, you know, like That's the only way of scaling. Maybe, you know, if it could actually spend like enough time on convolution, they can also make it a scalable and like, you know, like as good, but there was a…
AI assessment note: “why we have Uh four hundred billion parameter language models, but the biggest model”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q Your comments on, on pre-training are, uh, sort of against like that narrative that, that appeared a few months ago that pre-training was dead. That's not your take at all, right?
A I think everyone has ideas on pre-training side. At the end of the day, like going for that idea is a function of complexity and the expected gain, right? And sometimes you feel that, okay, you know, there are low hanging fruits and, and it like, you know, instead of bringing this complex, like, you know, recipe to the pre-training, the one that I have, like, which is simple, elegant, super scalable, I'm going to push this and then move the effort to the post-training. And then at some point, like the base model becomes the bottleneck, and then you're happy to take the complex recipe and bring it to the pre-chaining and then like hit pushing it. I think pre-chaining is dead. I would say like, maybe like, you know, the old, it's also like a little bit like difficult to talk about old and new, because like the, the, the timeframe is like very different. So when I say, oh, maybe I'm referring to like, you know, two weeks ago or something, but, uh, but But, uh, the way that we used to do pre-training, maybe, like, you know, like two, a year ago or two years ago, uh, maybe, like, you know, diminishing return is, like, obvious, but I can see how new ideas are, are bringing, like, you know, fresh, fresh, um, uh, energy into the pre-training and suddenly just open a door toward, like, like something exotic that might actually drastically change, um, the, the base model capability over …
AI assessment note: “new ideas are, are bringing, like, you know, fresh, fresh, um, uh, energy”
Answered raw tape
D 4 · C 3 · P 3 · Cm 3 3.30
Q And, uh, what's the reality of continual learning as of now? Is that, is that built into existing systems? Not at all. About two.
A There are two sides of it. Like one side is, I think like the research is not like yet to a very, like, uh, to, to a point that you, you think that, oh, you know, this is, this is the recipe, you know, I just need to kind of like, you know, exploit it and push productionization. Right. But basically, you know, like every time that you have, uh, a new problem that is like key, you have this phase of exploration where people like try to kind of like, you know, try different ideas and, you know, like, um, go Jump over this like idea to another idea, which could be like so different. And then when you're confident about this kind of working to some extent, you go to the exploitation mode and say that, oh, you know, let me just make it as good as it can be. And, you know, this is the way to, to kind of push it and, and let's scale it. Let's just like, you know, develop infra for it, make it like super fast, like, you know, productionize it and see what happens. I think that that is not yet there. The other one is also, again, like, as I said, because we, we've never had like, Like, super confident, um, uh, recipe for continual learning. Like, building infra for investing in, in something that is, like, fast is hard. Given that, like, I've, I've seen, like, very impressive progress on this, uh, of the Sweden Gemma and GDM. It's, it's kind of interesting because it is one of the thing…
AI assessment note: “research is not like yet to a very... point that... this is the recipe”