Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And, uh, quickly, what was your path to OpenAI? So how did you go from studying physics to being where you are today?
A I did a PhD in theoretical physics, uh, from, from MIT, thinking about the intersection of quantum gravity and quantum information. Thought a lot about black holes and quantum chaos kind of thing of what if you throw something into a black hole? What happens to the information? Does it, does it come out? How, if we think about black holes as computers, how fast are they? I was very interested in this fundamental question in theoretical physics, which is how do you find a quantum theory of gravity? I also got very interested in this interplay between computation and the laws of physics. You know, any computer exists in the universe in, in, you know, behaves according to physical law. So the sort of computations you can do are bounded by the laws of physics, and there's some sort of interesting relationship there. Black holes are pretty interesting because they sort of saturate some conjectured bounds around processing of, of information. And from there, uh, I did a postdoc at the Institute for Advanced Study. And around that time, I'm pretty old now for, at least for this field, so that was about 20 16, uh, was when the DQN Atari paper from DeepMind happened in 2015, and then AlphaGo was in 2016, and I got very excited about these, um, about the possibility of, uh, machine learning, and then, and then deep learning was statistical Science that lived in a similar framework to the…
AI assessment note: “I did a PhD in theoretical physics, uh, from, from MIT”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Since you mentioned, uh, test time compute, I think there's, um, something that, that still puzzles people, which is the The whole chain of thought thing which is so magical from a user perspective, whatever you can see, what actually happens during test time compute that creates those artifacts? What does the model actually do?
A I think it does what you see it do. We lightly rewrite it or summarize it, but it just, it just produces tokens, and those tokens are like a running thought process, just like You might have, or, or maybe it's more akin to, if you're solving a math problem, the, the scratch pad, the collection of notes that, that you have, but it, it just keeps generating. The cool thing about generating is that, you know, it, it, it, it's, uh, a forward pass to the model. So we're using a bunch of computation. So we're, we're, you know, it's a way of leveraging a lot more computation on a problem. Then, then you would before. So my colleague, Noam Brown likes to talk about the Riemann hypothesis a lot. And, you know, wouldn't you want to have a model that runs for years that, that, that can, um, resolve, resolve that, prove that if you present it and you want it to produce an answer, then it only has the number of flops in a, in a single forward pass to produce one, one token if it's forced to answer right away. But if it gets to answer after thing, you know, after a long time, it can, it can Re, reuse its weights, you know, produce, um, a final answer that is a function of a much, much larger amount of computation, and the, like, the natural way it thinks is in language. It's a language model, and so that's sort of this key insight that, that you can, um, cause it to do better just by produci…
AI assessment note: “it just produces tokens, and those tokens are like a running thought process”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you think there could be an equivalent in AI to thermodynamics, meaning, uh, you know, a compact theory that, uh, predicts behavior without tracking every individual bit?
A Yeah. Kaplan McCandlish scaling, open AI, uh, scaling laws work originally is, is a version of this where you throw away, you know, all you know about the network is how many parameters it is and how, how much you've, how much data you've trained it on it. And you, you can predict like the, the final loss. I think the, the missing piece Is going from all the individual weights and biases and, and how does that add up to the scaling law? I have some very like initial work and there's some other initial work about like trying to bridge that connection. But like, I think that's, that's the missing piece, like the sort of statistical mechanics to thermodynamics of how do we like, how do these things emerge? But there's definitely a lot of useful effective descriptions of how these systems behave. I think the other part of your question is like, Is it, is enough to characterize everything that we care about, right? There's probably a lot, there's a lot that we care about other than just the final loss function, and so there's, there's more thermodynamics to be worked out. In addition to like, how does the thermodynamics arise from the microscopic description?
AI assessment note: “Yeah. Kaplan McCandlish scaling, open AI, uh, scaling laws work originally is, is a version of this”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Thank you for that. Where do you think we are in the evolution of AI being increasingly able to solve difficult scientific problems? I mean, certainly something that we've been talking as an industry about for a while now, but it seems to be accelerating, perhaps just like everything else in AI. But where do you think we are?
A I think one of the interesting things is that this process is smooth. The, there's no sharp point, or I don't think there will be a sharp point where we'll say that systems didn't, weren't able to be useful for scientific, the scientific process to their fully fledged scientists. There'll be sort of a gradual shift. If you had to point to one moment, maybe it would be the release of O-one and by OpenAI and, and the sort of paradigm of test time compute and, and, and reasoning. But I'm sure if I tried to make that claim, You could go and look at GPT-IV and, and see that there's, um, glimpses of that sort of useful behavior for, for the scientific process were, were already present. As a general point, you know, the, the models are very good at certain types of things that clearly are amenable to, to making progress in math. They're not open loop, fully fledged scientists in, in any domain, although, you know, neither am I. It seems like it's just this really nice gradual process.
AI assessment note: “There'll be sort of a gradual shift.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q The OpenAI approach and the DeepMind approach were very different. Do you want to compare and contrast the two approaches?
A One of the approaches that GDM takes is to Take problems, present them in a formal language called lean, and then used methods to search for proofs in, in that language and some problems for problems to be representable. There's this process called auto formalization where you take English version of the problem and you translate it into rigorous formal statements, and then you, you conduct your proofs there. And it's, it's, it's designed so that the proofs can be airtight. No one has to go and check for, for some hidden assumption or some Weird thing, or I guess it's usually hidden assumptions or definitions that are not airtight, but in that setting, which is a setting that DeepMind has, has, has cared a lot about, they were able to formalize some, some problems and, and use their, use their system to prove them. So that's, that's one approach. Uh, another approach is to just take the problem in English with mathematical expressions as well, but just the English statement of it, which is Informal and understand what is meant by that and solve that in informal language, presenting a proof much like the way a human mathematician would, or human mathematician who's not using lean. And then you have to check it. It's the, the verification problem is, is, is harder because it's not something that auto checks.
AI assessment note: “One of the approaches that GDM takes is... another approach is to just take the problem”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Okay, fascinating. So just, uh, tying this back to the beginning of the conversation about Erdos problem and, and solving, uh, unsolved math problems. Presumably the instinct would be that you need a lot of exploration, not exploitation. So how does, how does that work in, uh, the context of, uh, novel scientific discovery?
A I think math research or scientific research in general has a lot of versions of both explore and, and, and exploit, um, to give, to give the recent example, the open AI unit distance Uh, proof I think is, is, is very much in the explore setting where the, the model was happy to be contrarian and, and, and try to disprove this thing that everyone believed. And it was, it was just looking for It has this huge repository of understanding all of human math. And so it was, it was spending a very long amount of time. I forgot how, how many hours, but I think we published a rewritten version of this chain of thought, but like hours and hours trying different things. So it's clearly in the, in the domain of exploration. Um, a lot of times though, you can ask these models to, uh, compute something that they understand very well. And then that, That has a different structure and might look a lot like exploit. There's a, a paper that came out recently after the open AI result where the, uh, an unrelated Ursh problem, it has something to do with, uh, if you have a set and you try to add it to a cell, add the set to itself, or you try to multiply the set with itself. So like take the elements and add them all together or take the element. Individualize or multiply them together into how many unique sums or products you get. There's some conjecture around that. And this, this one was also d…
AI assessment note: “math research or scientific research in general has a lot of versions of both explore”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q And then conversely, what's the catch, and how does RIL break?
A The setting where very difficult is, is the setting, um, that I alluded to before, where you don't get much feedback from the environment. You have to take many, many, many, many actions, and then you get maybe, yes, that whole set of actions was good, or no, it was bad. For instance, you're playing a game of chess, and you don't know what, you know, until you make all the moves. That has an opponent, so it's maybe complicated. Maybe it's you are trying to do a homework problem, and it's a research level, or, you know, like someone gives you a well-defined problem, like we give our language models, and it's, uh, you know, a problem that requires Days and days of thinking. There's so many choices that you can make along the way, and at the end, if you, if you don't get any feedback at all, if you're just hidden in the woods by yourself scribbling in notebooks, it's very hard to make progress that way, um, because you don't have any, you don't have any sense. If you get a yes at the end, or you get a no at the end, you have no sense for which of the, the actions that you took, which of the things you did were, were good or bad.
AI assessment note: “where you don't get much feedback from the environment”
Answered raw tape
D 4 · C 5 · P 4 · Cm 3 4.15
Q Okay, great. Alright, so let's get into reinforcement learning. To make this broadly accessible, let's start from the top. What is the one, two, three sentences definition for reinforcement learning and perhaps give us a simple non-technical analogy for people to understand?
A Maybe a simple thing to do would be to, to give you Two examples of, of how you could try to learn something you, you as an individual. Um, and, and maybe we can, we can take a game or even a, say, say a video game, right? I, I'm old enough that where I played the original eight bit Mario brothers, the super Mario brothers. And so here are two ways you could learn how to play. One way you could learn how to play is your dad takes it out, um, and plugs it, it in and he boots up the game and then he plays for a few hours and then you, you just watch him play. That's all you do. Um, so he's demonstrating how to play. And then at the end of that, he, you know, and he's not very nice, so he doesn't let you play. And, but then he like goes and runs outside and does something else. And, you know, you sneak into his room, you, you plug it in and you try to play. How, how good are you going to be? Well, all you've done is tried to memorize what he's done. You haven't gotten to push any of the buttons yourself. You haven't gotten to, to interact with, with the game yourself. Um, this is sometimes called, um, expert demonstrations and, and, um, and you know, you're, you're, you're just trying to memorize what someone else is doing. It's the version of supervised learning. The supervision being like, you, you just watch what he does and, and accept that that's the true way of doing the thi…
AI assessment note: “Reinforcement learning would be your dad's like, here, why don't you play?”
Partly raw tape
D 3 · C 4 · P 4 · Cm 4 3.70
Q questions in the field is, uh, whether you can expand and generalize the success, uh, that LLM systems have had in, particularly in coding and now math, uh, but like domains where you can sort of verify, uh, whether, uh, what the model comes up with is, is correct or not. What, what is your view on, on that? And, and perhaps start by explaining what a verifiable reward is.
A So a verifiable reward is, is, is In principle, a reward that, that, that can't be hacked. Uh, so it's a, if the, if it's a math problem, and the answer is an integer, you just string match the integer, and, and then you verify that it, that it did, it solved the problem correctly. Um, that, that abstraction has all sorts of problems with it, but, uh, unverified, uh, problem with an, that, Can't be verified. Is, is this a good piece of creative writing? That, that that's, there's not, uh, something you can sort of string match against, right? That involves questions of taste and maybe different people ask differently. So maybe it's a distributional kind of thing. And so there's, there's clearly a big gap between those, those two things.
AI assessment note: “there's clearly a big gap between those, those two things.”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q Why did RL start working well? It's not an entirely new concept. It's, uh, been tried for many years now. Uh, what is different now?
A Yeah, I'm not sure, to be honest, what, when people say it wasn't working, What that actually means. Like, there was this 2016, 27, maybe even to 2018 before the transform period where DeepMind was all in on RL and OpenAI had Dota and, um, Rubik's cube and some other exciting, exciting results as well. But a lot of people were all in on RL, and then there were language models, and the obvious thing to do was scale up the thing that worked, which was pre-training. And I don't know whether or what people tried for RL. As you pointed out, RLHF was a central thing that came pretty quickly. Originally, it was developed for in the, um, in the context of, of game environments, of, of like, Uh, trying to prevent reward hacking by using, I think the original paper was about using human feedback to like control, uh, um, uh, like a character for what, or some, something like that. But there's an interesting thing to, to point out here though, which is that, uh, there, there's this question of how do you get models to think in test time and, and reason. And There was a reasoning effort at OpenAI that was quite early and spent some time and, and came up with some, some algorithms. I think maybe the simple thing to say is that if you have a powerful enough pre-trained model, then it can start to do well at RL. It can start to like think, um, at, at use test time compute to, to for instance, …
AI assessment note: “if you have a powerful enough pre-trained model, then it can start to do well”
Answered raw tape
D 3 · C 3 · P 2 · Cm 2 2.60
Q he's, uh, claimed, uh, in my best attempt to paraphrase it was that LMs were not really intelligence. And, uh, therefore, uh, RL was the only way to do it, and pure RL, not LLM plus RLs. What is your take on this? I mean, obviously you're on an RL team at a company that does, uh, you know, both pre-training and RL combined. So what's your, what's your take?
A Let me tell another story. So when I, when I was, uh, before I did my PhD, I spent two years in the UK and I was at Oxford for one of those years. And I was at a pub as one does. And two of my close friends, one was the cognitive scientists and one was linguist. And so we had that sort of argument that you do in those situations when you're that age. And so something like physics is the most fundamental of all the sciences because it explains how the world works and everything is in the world. I said this earlier, my computer exists in the world. I exist in the world. We all follow the laws of physics. And then the cognitive scientists said something like, yes, but then you have to You know, you have to process it, so there's all sorts of cognitive biases about that, and, and, you know, the way you collect data and learn something, something, but then the linguist was, was like, blah, blah, blah, Wichtenstein, you know, the, everything goes through language, that's the method of communication, that's, you know, that's the way words mean things are, are, um, the central thing, and, and when we want to talk about the laws of physics, we have to use language, and I sort of feel like he and Wichtenstein, like, that, that was correct, right? That's, what's, what, Or at least the path through AI suggests that that is a correct path. I'm, I'm conceding now to Kyle. And, um, if he's, i…
AI assessment note: “To, to make things really work is through, is through language”
Redirected raw tape
D 2 · C 3 · P 3 · Cm 2 2.55
Q One of the famous things in the history of RL is, uh, Moves. 37. How do you, um, train a model to encourage the model to do that kind of things and come up with brand new ways while being efficient and exploit known path?
A Yeah. So the great thing about Go is that you can just train it. It's a, you know, zero sum two player game. You can Train, train in what's called self-play. It plays itself, and it can go from playing randomly to expert play, and it will find whatever the, the sort of best, best strategies are. So if that, if that means exploring, great. If that means exploiting, uh, I actually have a, I have a funny story about this. Um, so I, I met Noam Brown in grad school. Um, He went to a different grad school than me, but he wanted to enter MIT's PokerBot competition, and he had a PokerBot that was the, um, best in the world, but it wasn't something that would compete against humans yet. He just won in this research competition. He collaborated with me and, uh, and another friend to, uh, enter MIT's PokerBot competition. This is great actually for me because I learned some really exciting Work in, in AI, and I got very excited about this while I was doing, doing physics. We were playing essentially this, this kind of self-play equal equilibrium strategy. There's some, some nuances, but, you know, essentially we could not lose assuming we did not have any bugs in our code. The way this, this thing worked was that it was a tournament where you would be paired with, say, another, another person and play them. And if you, you know, depending on the amount of points you got in like some sort …
AI assessment note: “I actually have a, I have a funny story about this.”