Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Dota two moment? Because OpenAI, interestingly, in, in, in those early days of 2019 did a lot of reinforcement learning focus work, right? And then, then there was a whole like unsupervised learning, uh, GPT moment that happened afterwards, but like it started From roots in reinforcement learning, right? So did, did you work on that project specifically, or was it, uh, too advanced by the time you showed up?
A So the, the project that I worked was robotics project at OpenAI, which shared the same code and same methods as the Dota projects. Like in, in one hand, the Dota project was OpenAI's way to demonstrate the world, like what scaling up reinforcement learning can do. And in some way it was like taking the, the, the, the, the DQN agents and just, just, Doing all the hard work of making it bigger and bigger and solving harder, harder problems and opening. I generally from the very beginning was aware and really like, you know, it was a, it was simple, but, but genius inside that you need to have large scale system to, to learn really, really interesting, complex behaviors. And that was, that was like one, one way of what. Dota was trying to show that by scaling up reinforcement learning, we can solve pretty, pretty complex environments. And then, like, we have, there was another project. There were, I think, three reinforcement learning projects at OpenAI at that time. And then the second one was robotics, which is applying the same methods that we now knew or were proving that can solve pretty complex computer games. Can they solve all the practical problems? OpenAI was always optimistic and ambitious and trying to see If we can scale our own to solve data, can it load my dishwasher? Can it fold my clothes? Can it build a house? And, and this is, this is what we are doing. The pro…
AI assessment note: “the project that I worked was robotics project at OpenAI”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So let's, um, do, uh, if you will, a little bit of, uh, reinforcement learning one-on-one to make this, uh, really interesting to like a broader group of people listening to, to this. So in very, uh, simple terms, like explain it to me, like I'm 10, uh, what is reinforcement learning?
A Yeah. Yeah. I, I, I usually the, the metaphor and the analogy I have to a reinforcement learning is like training a dog. It's, it's very, very close, and I used to have a dog when I was a teenager, and even, even, I remember my parents did, I didn't know anything about raising a dog, but they, what they, they kind of invited through some friend of a friend, a fireman, who I think was working with like service dogs, and he came to me, and he basically told me a little bit about how do you train your dog, and what most dog owners that are like ambitious about training your dogs know, It is always extremely important to have a bag of treats in your, in your pocket. That's, that's, that's what you always do. And whenever you see your dog behave well, what you should be doing, you should, you should smile and you should give your dog a treat. Whenever you see your dog do something bad, you basically like give your attention away, turn away and become sad. And before the years of breeding, the dogs discover it's a, it's a, it's a, Like, you know, bad reward and bad behavior. And this is exactly doing that, but with models, we elicit a lot of different behaviors in the models, put them in challenging situations, and then we give them cookie if they do something we want, if they do a good thing, and give them some kind of punishment and negative reward if they do something like that we…
AI assessment note: “we give them cookie if they do something we want, if they do a good thing”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q world? But, like, taking a, a, a, a, a quick, um, uh, sort of, going down the rabbit hole a little bit about math. So, just in September, like, just a few weeks ago, uh, you guys did something, uh, unbelievable with the, uh, ISPC World Finals. Do you want to talk about, uh, what that was and, uh, what went on from, uh, Model technical perspective behind the scenes?
A There happens surprisingly little from, from our perspective, from the model perspective. We just, we just have a pretty smart model. And then when we ask them to solve programming problems, we, they, they, they, they, they are, they are correct. Uh, like what's, what's a little bit of a backstory in it is that I think we used like specifically programming puzzles for a while as a, Very nice research test bed of our, of our ideas. Those are, those are nice problems to experiment on them, and they, they weren't ever, like, considered part of the product, but it's, it's, those are pretty, like, complex problems, and they require a whole bunch of thinking, are very nice to give rewards to it, so all of researchers just, just, like, working on those problems as a way of trying out their, their oral ideas, like, you always need a data set, I will, I'll take a data set of programming puzzles and try it, and I think, I think, like, Because of that, a little bit, our moles just were always very, very good at competitive programming as a kind of byproduct. We never, we never tried to be good at it, but like researchers were trying their ideas on it, and because of it, like every, every training run, whatever, whatever we are doing, just, just end up being very, very good at those, those type of puzzles. And then like, you know, it was, it was a little bit of a, of a formality for us to …
AI assessment note: “There happens surprisingly little from, from our perspective, from the model perspective.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So fast forward to, um, today, still in the same vein of like the behind the scenes of, um, you know, you all at OpenAI and sort of life there. What's a day in the life of Jerry? Like what, what, what does somebody like you do? Like you, you, you, you, you read papers, you train models, you manage teams. Like what, what's your day?
A Like my days are surprisingly uniform, which is I come to the office early in the day after driving my kids to school. Then what do I do all day is basically talk to other researchers. I talk to other researchers all day, every day, and this is basically exclusively what I do. I take ideas from people, bounce with them, brainstorm with one partner, then move to another one and do the same thing over and over. And iterate. And in that K in, in, in, in, in that way, like keep refining our research program. Sometimes those are group meetings and group meetings are there as well and have their own team, team dynamics. But that is, that is basically exclusively what I do. The only thing that that changes is that the topics of research from, from, from, from meeting to meeting and from, from person to person.
AI assessment note: “what do I do all day is basically talk to other researchers”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q How are, uh, priorities in research determined? Uh, there's a range of possible projects. Is that top down? Is that bottoms up? Do people suggest ideas and others vet them? How does that work?
A Yeah, yeah, yeah. It's the, the art of, like, structuring, organizing, and leading a research project is, like, something that I generally, like, learned to appreciate very quickly in, in OpenAI's journey and in, in my career. There is something we are good. It's, it's, it's, Structuring research projects. And it's, I think it's a unique mix. Like you can't say it's top down. You can't say it's bottom up. It's a mix of those two, which is, which is balancing all the important aspects, which is one, one, one thing OpenAI embodies and determines is like, we all work on a very few projects total. There are not that many projects. OpenAI is not trying to do everything. We are not trying to, to, to like have portfolio. We are trying to have like multiple different bets. Always the, the idea is we do a few core things really, really well and put a lot of effort there, which means there are, there needs to be a lot of people working together on the same large scale, large ambition project. And we have, we have a few of those, one number, probably three or four, depending, depending on how do you, how do you call it? And that's, and that's it. And from that perspective, like people don't have ultimate freedom. It's not that people come to open AI and say, Hey, I want to do this. And they just, they just do this. Because you need to do something towards the goal of one of those four pro…
AI assessment note: “It's a mix of those two, which is, which is balancing all the important aspects”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q the other hand, like you guys seem to be just shipping and shipping and shipping and shipping across the organization, but including in terms of like core models, like again, to the point that you went from a one to three to GPT-Five in like a year. Um, how do you, how do you balance all of that? Like what, why are you guys able to ship so, uh, quickly?
A I, I, I think that the fundamental reasons for it is in general, OpenAI Kind of in my, my, at least worldview is a generational company in a way that we have incredible momentum behind us. We know that we were doing pretty great in the past, and we need to continue that. We have incredibly smart people, like literally the most talented people in the world are all coming and want to work at OpenAI right now, which means like people, every output per single person is incredibly high and everyone really Every, every single person does a whole lot. So, so we have, we have like a momentum that carries us forward. We have really great people that work together. We have like good operating, like way of structuring research and, and, and, and, and can borrow a lot from Silicon Valley, how, how to get things done quickly. And people are generally very excited about work. Everyone feels the weight and potential of what we are. Doing what we are trying to do. And because of that, people at OpenAI have a tendency to work pretty hard and, and then like, you know, having great people excited about what they are doing or working together reasonably well results in, in doing a lot of things. Um, like we like understand that there, there's only one time in history where AI is being built and deployed and developed and People want to do it in the best way that's, that is possible.
AI assessment note: “having great people excited about what they are doing or working together reasonably well results in”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q And that's part of, um, for people who may be curious, like the, the entire, uh, data labeling industry, so, uh, scale AI and, and, and a bunch of others, that's, uh, what they do, right?
A Yes, yes. Like, I think, like, in a way, I think it's getting more and more to be a thing of the past as the models are getting smarter and smarter. This is becoming less of a thing, but I think a few years back, and especially in GPT-IV days, this was the thing. The, the, the, the, the, the interesting bit about data labeling industry, I'm not sure how much we want to go on that tangent. It has to constantly reinvent itself because the AIs are getting smarter at some moment. Like certain things you don't want to label with humans if AI already can do it. So, so like you need to, you, you move the frontier and you, and you change the date type of data you are labeling as the, as, as, as, as you already are late to have the previous part.
AI assessment note: “Yes, yes. Like, I think, like, in a way, I think it's getting more”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q And, and since, uh, time and by that, I mean, the time spent thinking is so, uh, important to that concept of reasoning. How does the model decide Uh, how long to think when we're in ChatGPT five and we're in auto mode, and it says, uh, that is going to decide automatically, uh, how long to think? What, what happens there?
A It's basically part of our optimization process, partially for, for, for the, the happiness of the users and what they want to expect. Cause when, when, when you have a thinking process, you need to balance two things, which is like the quality of the result. There's, as we said, and there have been like those pretty great scaling laws that we demonstrated with the release of a one, the longer normal things, the better result you get. But also people don't like waiting, waiting, waiting is time lost that you could, that you could do something. Everyone wants to get results as quickly as possible. And like, you know, there is this, there is this saying you can get like, you know, cheap, fast or good, and you can take two. And that's, that, that applies to language models as well. There is a, there is a trade-off and it's, it's delicate. That's why like we also expose some of that trade-off to the users where you can have like a high rezoning model and the low rezoning models. And this is, this is like in the end, the same model. You just, we just tweak the parameter, which says we want you to think longer or shorter. We try to encode some heuristics of what we think the users will want when, when, when get, Thinking on an answer a little bit longer and getting to a better answer is worth it. Waiting versus not, but it's a, it's a bit of a trying to guess the anticipation of the …
AI assessment note: “We try to encode some heuristics of what we think the users will want”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Yeah. Uh, so agent is a model. Action is what the model does. Reward is how you, um, say whether that's good or bad. Environment. You, you hear a lot of things, um, these days about designing the right environment for RL. What, what does that mean?
A Environment is like, In some way it is everything that the model sees, but, but the interesting thing about difference about our environments and most other types of like what you can call supervised learning or unsupervised learning is that reinforcement learning environments, you want them to be interactive. You want like them to evolve as the model does things in general, like, like, Similarly, how, like, if you want to like learn how to play guitar, you kind of take a guitar and you strum it. And what happens is you hear sound of that, and then you hear it, and then you can do like learn to play like, like with, with actual like feedback of what is happening with the, of the guitar. And it's in that way, you know, the environment is like, how does the world react to your actions? And a lot of like, what drives your actions is the, It's what is happening in your, in your environment, and it's in your, in your world. And that's kind of, like, the only way how to, like, really teach agents to, like, learn to react to changes in the environment is through reinforcement learning.
AI assessment note: “the environment is like, how does the world react to your actions?”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q one, I guess a little over a year ago in September, uh, with the concept of, of chain of thought, which is, in layman's terms, the little messages that you see when you query, uh, chat GPT and it Tells you, it shows its work. It tells you what it does. What does that actually do? Is that a logical tree and it eliminates option after option? What actually happens?
A Language models do on their own, like fundamental level is they, they are often called as next token prediction machines. And that's not completely accurate in the age of reinforcement learning, but they still operate on mostly on tokens that are mostly text. The language models, again, are these days also multimodal, and they operate on mostly text, but to simplify a little bit for a second, language models generate text, and what chain of thought is, is their thinking process verbalized using human words and human concepts. So the, the magic that we are seeing why this is all possible is that while you are training on all of internet, on a lot of Human knowledge and human thinking process. The model starts learning in some ways to think how humans do, and in some ways get to the answers, how humans do from seeing humans do it a lot in the text that was, that was pre-generated and that, that, that was based in a training data on humans. And then, and then the chain of thought is basically eliciting that capability in language models of like, Thinking and getting to an answer, like, like, like, like humans do. Uh, a lot of what, what, what, like what early chain of thought work was doing was kind of solving math puzzles. And the first, like most famous prompts to elicit chain of thought in language model was so-called like, let's solve it step by step. There is this, this, this…
AI assessment note: “chain of thought is, is their thinking process verbalized using human words and human concepts”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q the right way to think about it as a combination of pre-training and RL? First of all, is that the right way to think about it? And second, if so, just at a high level, how does the articulation between both of those work? And then after that, I'd love to do a little bit of a, you know, deep dive on RL to make this very educational for folks.
A Today's language models basically can be thought as like, first they are pre-trained, then you do reinforcement learning on it. Like the reinforcement learning would not work without pre-training. And I think in a similar way, like pre-trained models have a lot of limitations that are very hard to like. Resolve without doing something that looks like reinforcement learning. So I think both of those bits are here to be and to stay. Like, I think. The way how they are, like, combined and, and do, that may, it probably will evolve in the future. Nothing should be treated as, as dogmatic and fixed, and we need to, we need to keep generally figuring out the way how to train better models, and this is what we are trying to do. The interesting thing, and, you know, I, I can credit that to, to Ilya, how much foresight that he had, but whenever I was, like, uh, started at OpenAI early, Right. And, and, and, and I remember there was like research all hands or something like that where India came on stage and talked about like, what is, what is OpenAI's research program? What are we trying to pursue? And what he said at the beginning of 2019 was to train large generative model on all data we can and then do reinforcement learning on it. That was, that was the OpenAI research Plan at the beginning of 2019. And this is exactly what we are doing today. Now the algorithms change, architecture…
AI assessment note: “first they are pre-trained, then you do reinforcement learning on it.”
Redirected raw tape
D 2 · C 4 · P 3 · Cm 3 3.00
Q You, you mentioned, um, working on, uh, such a GPT agent, like the agentic AI, like what, where does it all fit the, you know, the, the, the tool use, like the, all like, uh, agentic autonomy versus reasoning RL, like help us to reconcile. What does what, and what impacts what?
A I think what is the important thing is I, I believe, and I think that there can be a lot of positive impact of AI on our world and on our lives through automation, through problem solving, and through AI doing good things for us, the things, the things that we want. And for a long time, and for a long time again, it's not that long, but the last two years or so, or maybe, or maybe approaching three, We've been like living in this world where we kind of ask questions to AI and gave us an answer, uh, for at the beginning instantly. Now it can think for like a minute or two, which, which feels long, but in many ways, like what's, what can you do for two minutes? If you think of like how many problems humans solve, and AI is probably a little bit faster in the things that can, it can solve, but it's still a limit of what it's, what it can do. There are still a lot of tasks that you know that Would take AI to do much, much longer. If I, when I prompt codex, it works for a while. I get a few minutes. Like, there are a lot of, like, things we have internally and we are doing that, like, allow them all to work for much longer. We still didn't figure out the right product to deploy them, but the models I can think for, like, 30 minutes, hour, two hours these days on certain, certain types of Tasks and problems like even, even, even longer than that. Um, they, they, they generally are ca…
AI assessment note: “I think what is the important thing is I, I believe”
Partly raw tape
D 3 · C 3 · P 3 · Cm 2 2.85
Q Okay, great. All right. So going back to, uh, RL, you tweeted the other day, GRPO, the GRPO release has been, uh, in a large way has accelerated the, uh, research learning, uh, program of most U.S. Research labs. So what, what, what is, uh, GRPO?
A Uh, yeah, it was a, it was a little bit of a tongue in cheek moment. I am extrapolating here a little bit what exactly happens because I, I haven't been in most US research labs, but I have some mental model of, of what happened and how, and like, long story short, GRPO was the open source released from deep seek. And there was a, like everyone who like is terminally online follows AI discourse knows that deep seek moment. Uh, whatever it was when the, when the, the, the Chinese company that seems to be doing really, really great work, released new model. And it was, uh, it also pre-trained model, a reasoning model, the open source, the algorithm, the open source, a lot of things they did overall, like really, really great and technically excellent release. And like, there was a lot of discourse about, about, uh, like, you know, that they, they, they, they pre-trained their model particularly cheaply. And that was, that was part of the discussion about that deep seek moment. But the other part of the discussion was that they, that they kind of like released their reasoning process. It was like not very far after our O-one release. As far as I know, like our O-one release mostly caught a lot of us labs by surprise. They didn't have like similarly advanced RL research program to my knowledge, basically no one. And like, and I think like the only company in the world that they get…
AI assessment note: “GRPO was the open source released from deep seek”
Answered raw tape
D 3 · C 3 · P 2 · Cm 2 2.60
Q to close this conversation. Uh, you said the other day, um, you tweeted, we all collectively believe AGI should have been built yesterday. Uh, and the fact that it hasn't yet is mostly because of a simple mistake that needs to be fixed, uh, which is, um, super awesome as, as, as a tweet. Uh, do you think that the combination of, uh, pre-training and scaled RL Texas to AGI?
A There's always an interesting question of, like, what, what, what do we consider something that is not pre-training on RL? And like, where, where, where, where is the, where is the limit? I generally think something that we are doing, like, pre-training today is necessary. I think something that, like, we are doing RL today is necessary, and there will surely be a few things more, and like, we have a lot of, Very ambitious research programs on some of those things. And like, I, I, I, like, I didn't think like the question of distance in research space is hard to say, like with some, for some people, like what we, what we are, what we want to do and what we are planning to build is not very far from those things for someone will say, oh, it's completely different. And it's like a very much not that side. I don't want to go into, into like debates, whether it's the same or not, but like, We are and want to be constantly changing the way how we train the models to more represents what we think the, the, the, the right form of, of intelligence is and the most useful, useful form of, form of learning is, and constantly are, are we searching various things? And then all the, the, the distance from like, what is, what is the distance from AGI is also like a very complex question. I really like someone said it to me, but I think It is right that, you know, if you talk to someone from 1…
AI assessment note: “there will surely be a few things more”