The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Yann Dubois argument clarity score 4.2/5 from 15 exchanges on raw tape · average scores: directness 4.6 · coherence 4.3 · precision 3.9 · compression 3.3 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
15exchanges match
15on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And, uh, going back to, um, generalization as well, are there examples where, um, actually getting better at one domain makes the model worse at, uh, the rest? A little bit, uh, to what you were saying about, like, some people are very good at math. Some people are very good at English. Pretty often they're not the same people.

A In domains, usually not. What will happen though is, um, You will make decisions based on which domain we optimize for. And if you optimize for one domain, you will be able to optimize less for another one. So it's not necessarily that optimizing for one thing will make the other one worse. It's just that as a result, you can optimize less for the other one because your compute constraint, your data constraint, you have like, like your, your, uh, human bottleneck also in terms of that work. What does happen is, uh, you can have negative kind of generalization, like bad generalization or negative transfer. More for these horizontal aspects of the model. So I'll give you a very concrete example. Explicit instruction following versus implicit instruction following. If I, if I have a model, and this is, we often hear, for example, from OpenAI models, that they tend to be really good if you tell them exactly what you want. Um, but as a result, sometimes we hear also that they're, like, less good if you are not as, as specific about what you wanted. For example, if I make, if I make a typo, and I say, like, change this file, and I make a typo in this file, um, an extremely good model at, like, explicit instruction following will change the wrong file, the one that has a typo. But, like, humans would probably realize that you made a typo. Um, and, and like, as a result, there are case…

AI assessment note: “In domains, usually not.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Does reinforcement learning, uh, create new capabilities, uh, or does it make the model better at existing capabilities?

A It's really hard to say because pre-training, when it's trained on all of the internet, arguably already has all capabilities in it. Um, so it's, it would be even hard to answer this question scientifically, um, because arguably everything is, is already there. What I would say is that if you look, uh, at models that we were training or that we were post-training, like, two years ago in the open source world, uh, for example, I, I worked on one of them, Alpaca, where we used 50,000 examples for SFT, and, like, now when you look at reinforcement learning from From models like Kimi or, or from DeepSeq models, it seems that they are closer to one million data points. So definitely people scaled up a lot the reinforcement learning stage. Um, and from this, it seems that they've learned like new capability, like this reasoning aspect, this fact that you can check your answer and, and, uh, and try to improve it. Um, so you can, you can really think for longer to get, to get a more correct answer. So all this to say that Arguably everything is already in pre-training, but we were definitely able in the last one year and a half, even in the open source world, um, to have more capabilities after reinforcement than we used to before.

AI assessment note: “it seems that they've learned like new capability, like this reasoning aspect”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q and I'm, I'm curious, What reasoning means in 2026 that's any different from, you know, a conversation we could have had about oh one or three. Um, in particular, one of the claims, uh, of 5.5 and, and also my experience as a user is that it's particularly good with, with messy data, which seems to imply that, um, it needs to reason through ambiguity more. Um, what has changed?

A What I would say is that oh one and oh one preview, uh, we're really. Really breakthroughs, um, in, in the research community about having a model that can think, and the longer they thought for, the more, like, the higher likelihood they would be of being correct. Um, so that was really a breakthrough, but initially, and if you look at, like, old blog posts, you would mostly see, like, math, uh, math evals, and also, like, Maybe coding competitions, but things that are really easy to test whether you're correct or whether you're not. Um, and it also gives you like some suggestion about like how we were training some of these models. Um, and how I see maybe all of last year and especially the end of last year and the beginning of this year is that we were able to take these algorithms that work with, uh, verify rewards, like things where we can say you're correct or you're not, uh, to the messy real world. Um, and really optimize for the utility that we provide to users, and like making them more productive. Uh, so I think that's what really changed.

AI assessment note: “we were able to take these algorithms that work with, uh, verify rewards... to the messy real world”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q The 5.5 was particularly good, um, on genetic coding, computer use, knowledge work, and early scientific research. How does that work internally? Do different people focus on those different parts? How do you get to that result?

A Yeah, we definitely have different teams that are working on specific use cases and are pushing on these use cases. Uh, my team specifically is actually the one that is kind of taking all these vertical improvements and try to put them together in the final model. You could see it as a team that is doing both kind of the smoothing function. So you have all these improvements, but you need to make sure that the model doesn't feel too spiky, doesn't feel differently on different verticals. And also you need to have some teams that are working, and that's basically what my team is doing, on all the horizontal improvements. So there are many things like instruction following, function calling, or like thinking about how much should a model think for on different Uh, problems. Those are very horizontal and that kind of impacts all these use cases. So we have both these more vertical teams and these more horizontal ones. Um, and both are very important, uh, to, to, to improve the, to improve on the model. Um, and the good thing is that these things can kind of be improved orthogonally. So you might have like multiple different teams that are working on certain verticals and maybe for one model, there's only a Half of these teams that made integrations basically in the last run and like improve the model on these capabilities, and maybe for the next model, it'll be the other half. So …

AI assessment note: “we definitely have different teams that are working on specific use cases”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay, great. And, uh, what was your journey to OpenAI?

A Oh, it's a long story, but I'll try to keep it really short. Uh, basically I did my undergrad in biomedical engineering, um, in Switzerland. Um, I'm from Switzerland. And then I went on an exchange in Canada and I learned about what to VEC. So I don't know if you heard about this algorithm, but it basically takes words, which is like a, something discreet, uh, and puts it in a, in a vector space. Um, so puts it basically in a way to think about it as a plane where if words that are more civil to one another will be closer to one another. So it brings these, like, discrete words into, like, some continuous space that is semantically meaningful, and I was absolutely blown away by that algorithm, and that's when I decided that I wanted to work on natural language crossing and just, like, understanding language. Um, at that time, I was very wrong, but I thought that, uh, English, Uh, uh, NLP was basically solved. Well, like, close to being solved. That was in 2017. So that was, uh, uh, right when Transformers started. It was actually right before Transformers. So I was very wrong, but, uh, I decided that I wanted to work on under-researched languages, and basically, um, I wanted to improve, um, NLP on languages where we don't have that much data. Uh, so I went to, uh, work, uh, for Grab in Singapore. And I was basically building the natural language processing, uh, pipeline for the…

AI assessment note: “ended up at Stanford, did my PhD there... and then, uh, went to open it.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And still on the topic of, of reasoning, um, what, what's ultimately the difference between, uh, 5.5, uh, thinking versus 5.5 pro? Is that, is that just more test time compute, more tokens, and more time invested in solving a problem?

A Yes. Basically, it's just a question of, of, uh, how much test time compute we pour into the model, uh, or we pour into this entire, uh, system that we're shipping. Um, so, We, we've seen again and again, the longer the model think for, uh, the better answers we will get. The problem is that this, these curves that we're talking about, um, are not, are definitely not linear, and like they, there's some plateauing effect, and they kind of look, um, look logarithmic, uh, on some, in some sense, um, or depending on which evals. So You can pull, like, two times more compute and actually only get, like, small performance gains. Um, I personally don't use Pro that much because I really don't like waiting. I'm pretty impatient, so I don't like waiting for that long, and, uh, and I know that the probability of being correct definitely improves, but it doesn't improve, like, enough for, for me to use it. Um, but there are some people who use Pro and who really love it, especially actually for academic research. And, uh, I know especially a lot of mathematicians who are using it, and that's because they're kind of just have this in the background that is running for maybe one hour, uh, two hours, and they don't really need to like iterate really quickly with the model. Um, and pro is really good for that.

AI assessment note: “Yes. Basically, it's just a question of, of, uh, how much test time compute”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q What's the current, uh, frontier of reinforcement learning? It, it seems like there's a jungle of acronyms like GRPO and other techniques. What, uh, what are you using? What are you excited about? What do you think is promising?

A So I can't talk about what we're using, but like, for example, uh, in the open source world, GRPO seems to be working very well. Um, and People used to have different methods like PPO and DPO, and like, people seem to have really converged to this one. Uh, the big, the big difference with others, other methods is that, um, you, Again, you do this, like, simple method that I told you about, like, sampling as many answers as possible, and you say which one is correct. Uh, so in some way, GRPO is a very simplistic method. Uh, and in general, we saw over and over again in machine learning that the, the simplest method that where you can scale up in terms of compute usually is the one that ends up working the best, and that is kind of, uh, what is happening here, um, at least in the open source world.

AI assessment note: “in the open source world, GRPO seems to be working very well.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Okay. Fascinating. So we're going to unpack a lot of this, particularly on the, on the RL side. Uh, for the first thing that you mentioned, reliability, is that, uh, engineering? Is that models? Like what, what makes a model reliable in, in, in the way you meant it?

A It's a little bit of everything, but in general, given that these are agentic models, uh, the longer, if you just think about it as like every two minutes, there's like a certain probability that they are wrong. Uh, the longer that they run, The, the higher the probability that, like, the final answer is going to be wrong. Um, so it's just something inherent in, like, agentic models, and what we've been pushing a lot on is, like, making sure that the model, like, we decrease this probability of being wrong every, like, two minutes. So purely from a model point of view, of course, there's a lot of reliability that is also happening on the applied side, and the team at OpenAI has been doing an amazing job, um, on that. Um, but I'm, I'm even talking only about reliability of our models. And, like, making sure that, like, basically we decrease the probability of being wrong.

AI assessment note: “It's a little bit of everything, but in general, given that these are agentic models”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q And by embodied intelligence, so you mean, so potentially robotics. And so if you use a video, uh, that shows how gravity works and how a robot evolves in space, then presumably that would be more useful. Is that, is that the thought?

A Yes, the idea, the intuition that I think many people had, and I definitely thought for a long time, is that it's hard to understand the world only through text, and, and there will be, um, it's hard to understand what, like, what physics is without really seeing what, like, for example, you can't understand gravity without really seeing things falling, um, and When you look at our models, I mean, it's, they kind of understand gravity without having seen that, but it still seems not obvious. Like, it seems, still seems like they would get it more. And like, they are still kind of missing some common sense, uh, aspects. Um, so I do feel like we will improve the common sense of our model by having them interact in the real world. Um, but we are, we're still pretty far from that, I think. And by we, I mean just generally the academic community and, and the AI community seems pretty far from that.

AI assessment note: “Yes, the idea, the intuition that I think many people had”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Yeah. And while we're on the topic, as a quick detour that, that leads us to the concept of world models. So leaving your, taking your open AI hat off, are you, are you bullish on world models?

A World models in the sense that, um, yes, you can try to replicate or like simulate things, uh, Simulate, like, basically work in an environment that is simulated. Um, yes, the problem is simulations are always going to be really hard and are not going to be truthful. So I think there will always need to be a certain, a little bit of training that will need to happen in the real world to make sure that the model realizes kind of these mismatches between the simulated world and the real world. Um, and I think, uh, we as a field have a tendency Of, uh, optimizing something that is simulated or not quite realistic, um, past the point where this is useful. Uh, so that's like something that I think we should always be careful with, is we spend a lot of, of time and effort on optimizing something simulated or not quite, not quite realistic, and it's great at the beginning, but at some point, once you start optimizing too much for something, it's, it's not representative of the real world, and, uh, and people continue doing that just because that's what they've been doing for a long time. Um, so I just think people need to realize when to stop that. I don't work with, uh, um, With this type of synthetic environment as much, um, or just because I don't work on embodied AI. Um, so I don't know if we heard that yet.

AI assessment note: “yes, the problem is simulations are always going to be really hard”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q As you describe some of the challenges, question crossed my mind, uh, you know, you often hear that AI systems are not built, they're grown. How you'd characterize it as well? What part is science versus a craft or trying multiple things and then just keeping what works best in your day-to-day life?

A Yeah, that's, that's a great question. I think how it usually works is that it starts being craft. Uh, people just try out many things, and they start building a mental model of what works and what doesn't, and over time, we move to, like, from this, like, craft land, uh, to more science. Science, uh, is, well, like, more scientific approach is, are really the ones that, like, first end up working. It's hard, it's very rare that, uh, you, you take a really scientific approach, and, and you say, um, Uh, like, this is the optimal, the optimal thing to do, and you do it, and it just works. Like, people just, uh, there's some sense of alchemy. People just have, like, a good flair for something, and they make it work, and then other people, or that person, uh, starts trying to improve what we are doing by being very scientific. Um, and I would say this, this happens over and over, um, in, uh, in, in machine learning. Uh, so first craft, then science, and both are really important. Uh, but it's different stages of the pipeline. In terms of engineering, this is definitely something that is, uh, uh, always necessary. Uh, so I would say most researchers have moved to Being relatively, uh, good at, like, figuring, at least, I wouldn't say good engineers, but good at working in, like, complex systems, and, like, figuring out what they need to, to, to try out, and the systems, the, the, an…

AI assessment note: “it starts being craft... over time, we move to, like, from this, like, craft land, uh, to more science.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q And the jagged nature of those models, does that come from this approach of picking this problem and that problem, and therefore it's going to be excellent at those problems, but not as good as other problems? Or is that a more fundamental characteristic of AI models?

A There's definitely some of that. Uh, for sure, if you optimize more on specific types of problems, you will be better in that setting. Um, I would say, this is my intuition, is that it's less About the exact, like, problems that you're optimizing on, and it's more about the class of problems that you're optimizing on. So for example, uh, if you are really good at, like, math competitions, your model will probably be pretty good at, like, coding competitions. So it's not about the domain. It's more about, like, the skills that are necessary and the way to think, um, uh, and, and it's, like, horizontal, um, capabilities that you need for performing these tasks. And that's, that's, um, what I think you're usually seeing when some, when some model is really bad at something, it's actually bad at that in any domain, in any language. Uh, so, so you have to think about this domain and then, then this generalization of this domain, not necessarily per domain, uh, capability.

AI assessment note: “it's less About the exact, like, problems that you're optimizing on, and it's more about the class”

Answered raw tape D 4 · C 3 · P 4 · Cm 3 3.55

Q And is the pace of progress in model as a judge and AI evaluating AI, is that, is that moving as fast? Is that a distinct part of research or is that fundamentally the same idea or the same techniques?

A It's really fundamentally the same method. It's like nothing. Also, Most of the things that we do in evals, especially now that we have reinforcement learning, could just be applied nearly exactly as is during training. So that's another reason actually why evals are so complicated is that every time you build an eval, you actually build a way to build training data sets. Um, so now you're going to optimize that training data set. Well, not even if it's not that evil, it's going to be the same type of data. And now you're going to do super well because we have this generalization of, of, uh, Of, uh, capabilities that I was telling you about. You will learn that on that other data set, and now you'll become really good at that eval, and that eval will become, um, obsolete really quickly. Um, so, so that's also an issue with you. But yeah, to come back to your question, Um, the model as a judge, it's really important, and I think it's one, one probably of the most important things, because as we get, like, better models, uh, we have this self-reinforcing loop, and we have this, this, like, capability flywheel, where better models become better teachers for other models. Um, and this is really important for training, but then you can also do the same thing for evaluation. So I, A lot of my team works on that, and I think it's really critical is to work on this, uh, uh, model, mode…

AI assessment note: “It's really fundamentally the same method.”

Answered raw tape D 4 · C 3 · P 2 · Cm 2 2.90

Q And, uh, quickly in layman's terms, what is the fundamental difficulty?

A It's a good question. I actually don't quite know, to be completely honest with you. Uh, I don't quite know why it's taking us that long to figure it out. Uh, it's this type of Of domain that I, I think if we really put enough resources behind it, like, we would figure it out. Of course, there's, especially when we talk about, like, this memory inside of a company, there's, there's big questions about, like, permissions, and there's, like, a lot of, of, uh, questions about, like, um, uh, privacy and, like, what you can share and what you cannot, like, across models, across users, sorry, but for a single user, even for a single user, we're not quite there, and I, I don't quite know why. At least at the high level that I can talk about, I don't know why.

AI assessment note: “I actually don't quite know, to be completely honest with you.”

Redirected raw tape D 2 · C 4 · P 3 · Cm 2 2.85

Q of, um, 5.5 specifically, you know, big narrative of last year was that, uh, pre-training was hitting a wall, uh, and was not going to yield much progress. That seems to not be the case at all, uh, in 2026. Uh, can you walk us through some, some ideas for what is happening in pre-training and why it's progressing now in a way that people hadn't predicted, uh, last year?

A For pre-training, I, I can't talk, um, in a lot of details about what is happening internally. Besides that, um, the team has been really doing a lot of good work, um, and our models are really getting better and better. Um, one, one thing that I do want to, um, highlight when we're talking, for example, with efficiency, um, If you have larger models, uh, the amount of thinking time, so the amount of tokens they will think for, um, will usually decrease. And the way that you can think about it is that, um, metaphorically, the model already thinks through its weights when it generates a certain token. Um, so you can, you can decrease the number of, like, tokens that it needs to generate for thinking by kind of, like, increasing the size of the model. Uh, that you are training. Um, so, so oftentimes if you just increase the, the model size, if you basically train, uh, pre-train larger models, uh, you will get better efficiency. Um, and the good thing with larger models is that they can be paralyzed better on, on, uh, at inference time. So the, even though you might think, okay, you actually generated fewer tokens, But by a larger model, uh, so you actually might decrease the, uh, the overall efficiency of the system. This is not true, because the larger the model is, the, the more chances you have to actually, uh, optimize for, optimize BC for inference on, on, on GPUs. So you wi…

AI assessment note: “I can't talk, um, in a lot of details about what is happening internally.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.