Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q And going back to that increasing, uh, pre-training, increasing the scale of pre-training, delivering predictable improvements in model performance. Um, yes, now post-training is in the picture. It's making models better in really impressive ways. Um, but are you of the belief as opening? I have the belief now that there are diminishing returns from pre-training, um, given that we're now talking about different forms of training these models.
A Not at all. Um, our scaling laws still hold. Uh, empirically, there's no reason to believe that there's any kind of diminishing return, uh, on pre-training and on post-training. We're really just starting to scratch the surface of, of that new paradigm. Um, you know, the, the O series of models, which were kind of the previous reasoning models, um, were really just the beginning of, uh, us starting to explore what's possible in that post-training regime. And I think that's going to be kind of the dominant theme here for the next Year or two, um, is continuing to scale in that dimension, uh, and continuing to see the gains that you get there, um, simply because they're so significant. Uh, and so now we're pushing on two axes for how to improve models, and we think that's going to tighten and condense the rate of, uh, of, of innovation.
AI assessment note: “Not at all. Um, our scaling laws still hold.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q have seen is just a general increase in capabilities across the board. There were no caveats of like, um, and maybe there's a reason for those caveats, but there were no caveats of, you know, there's, uh, intelligence increases in this place and that place. It was We trained a bigger model, I'm pretty sure this is what it was, and it's better across the board. So, have things changed?
A They've changed, yeah, from a technical perspective. I think when you go from GPT two to GPT three, three to four, these were really just, uh, exploits of what was, uh, and is the scaling paradigm of training larger, pre-training bigger and bigger models, training larger models. Um, it's kind of one vector of training, uh, and you get a better model that, uh, as a, as a result. Um, and that continues to hold true, but we now have this kind of other category of, of, of training, which is post-training, Uh, and being able to use test time compute in more interesting ways than we used to as almost kind of a second stage of training. And so we think that that actually gives us a little bit of a boost, um, a force multiplier on our ability to push the model toward new intelligence levels, um, and also be able to train into it a lot of the things that you want an intelligent model to be able to do. Um, so using tools, for example, is something that really thinks really important, uh, for overall intelligence, GPT two and three, Um, couldn't really do that as well. GPT-IV could do it in a more nascent way. Um, and now GPT-V, you get that baked in, uh, with the benefit of, of these kind of multi, multi-step and, and longer horizon reasoning processes. So, um, yeah, we, we want to abstract that from users. Obviously, we don't think that you as a ChatGPT user should have to stop and thin…
AI assessment note: “They've changed, yeah, from a technical perspective.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah, it is interesting to know that you've been, you have been already working with many companies, uh, and letting them use GPT-V already. So has there been a sort of unified, we couldn't do this with the previous models, but we can do it now with GPT-V or is it sort of spread out in terms of the capabilities that it's now enabling?
A Um, I would say it's, it's been, uh, you know, rising tide across the board. So everyone who's kind of benchmarking and all the companies that we work with typically now are, are pretty accustomed to, to evaluating and benchmarking performance across all the models that they use. But, um, everyone has kind of reported, you know, much higher, kind of consistently higher performance on those evals. There are a few areas in particular we've seen spikes. So one is coding for sure. Um, I mentioned companies like Cursor, JetBrains, Windsurf, Uh, you know, Cognition and others that we work with who, um, anecdotally are all, uh, you know, have, have all said that GPT-V now feels like the most capable coding model, whether that's in an interactive coding environment or more of an agentic coding environment. Um, and then also one of the things that we see consistently now is its ability to reason and problem solve in very technical domains, uh, is significantly improved. And so, um, Harvey's a great example of that where, uh, you've got, you know, Harvey AI working with legal firms and law firms, Uh, is, you know, very, very reliant on its ability to, uh, reliably, accurately, um, and, uh, and consistently portray, uh, you know, uh, cases that, that, that it's looking at, legal analysis, um, to provide that kind of level of structured thinking you want when you're doing legal analysis, a…
AI assessment note: “I would say it's, it's been, uh, you know, rising tide across the board.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q five comes out yesterday or as the release happens, Sam says, I kind of hate the term AGI. Because everyone at this point uses it to mean a slightly different thing, but this is clearly a model that is generally intelligent. Help me understand what's going on, uh, because it seems like maybe, maybe he wants to call it AGI, but you're not yet, so why is this not AGI?
A Well, it is, it is a hard thing to define. Um, you know, you ask the, the joke here is you ask five people what AGI is, you'll get seven answers. Um, and I think the way we kind of look at it is it's a cumulative process, right? It's a system. Um, and I think you have to define kind of what is it that that system is and what do you expect it to be able to do? And for me, at least that's a system that is reliably able to learn new things that are kind of out of distribution by virtue of its ability to reason, to Think to solve problems, to use tools, to come up with new ideas. And so I do, I think we're at a system that I would call AGI. No. Um, but I think we see, we start to see the traces and the, um, the pieces of that overall system for, for generalized learning start to come together, uh, in models like GPT-V and I suspect suspect in its successors. Um, I don't know if we'll have a point where we are like, okay, we've crossed from a non-AGI world into an AGI world. Um, and even if there were, I'm not sure we'd actually realize it necessarily until after the fact, because one of the things we've learned working with the models that we have is the capability overhang is significant. Um, I think when Sam refers to the intelligence of the models and having a PhD in your pocket, we haven't yet really exploited that as, uh, as a thing. Um, you know, in some sense, like, I think …
AI assessment note: “And so I do, I think we're at a system that I would call AGI. No.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q know, I think Sam mentioned this on the media call, but GPT-III, he said, was high school level intelligence, GPT-IV, maybe the level of a college student, and GPT-V, an expert. So I guess, I wonder for OpenAI, is the quest to add more intelligence to the mix, Or is it to focus on capabilities other than smarts? Some of the things that you mentioned, like memory and continual learning.
A It's gonna be, I think, all of those things. Um, certainly there are some unsolved problems. Uh, you mentioned a few here, and I would agree with those, um, that, you know, you'd expect a really smart person to, you know, kind of comes by default that our models still struggle with. Um, and so there's open research there that we still have to do, I think, to be able to kind of close the loop on what I would call the full spectrum of intelligence. Um, but, you know, there's intelligence like we were talking about earlier in, in the podcast expresses in a lot of different ways. Um, and part of it is just your, you know, pure IQ. It's your knowledge of how things work and your ability to recall information, but then it's also your ability to reason about how to use other tools to solve problems. Uh, it's your ability to be reflective and to look back on your own chain of thought, your own line of thinking, and actually course correct when you feel like, you know, I actually went down the wrong path and maybe I didn't come up with the right strategy to solve this problem. And so, uh, that's one of the cool things we see is GPT-V on those vectors. Um, we can actually reliably measure as better than the previous systems we had. And for us, I think one of the real world things that we really want to understand is how do they actually perform, uh, in, you know, in, in the real world? H…
AI assessment note: “It's gonna be, I think, all of those things.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q level biology, um, the average chatbot user may not feel that even though it's, um, even though it's gotten much smarter. So I guess, I'm curious how, how you think this will be reflected, the increased smarts will be reflected. And the average users chat GPT experience and the plus users experience who've been using these reasoning models for a while, is it going to feel any different for them?
A Yeah. Um, I saw something on, on X that was akin to what you're describing, which someone basically kind of said, I think for the, you know, upper echelon of, of ChatGPT users who are probably in the paid tiers, who are very, you know, active on a daily basis and are really kind of expert level using these systems, it, it, it's gonna feel like an improvement, but maybe a, uh, you know, a more subtle improvement, but for the average user, for the free user, um, and we're, we're bringing GPT-V to our free tier, It will feel like a dramatic increase. Um, if you actually look at kind of the way free users have used ChatGPT, most of them have actually not experienced the power of the reasoning models. Um, they mostly are using GPT-IV-O, um, and, you know, they, they mostly are kind of using it for this very kind of, um, you know, turn-based kind of like very quick, uh, you know, back and forth, almost search-like, uh, that ways that I think don't actually kind of express the full capability of the model. And so for a lot of people, this will be the first time using a model that has reasoning capability. And not only will it be, you know, the first time using it, uh, with reasoning, but it'll be the first time that they're experiencing a model making a decision about how long to think about a problem and how good of an answer to give relative to how hard the question is. And so we ex…
AI assessment note: “for the average user, for the free user... It will feel like a dramatic increase.”