Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q eight hundred million. Reports say approaching nine hundred million. Um, but then on the other side, you have distribution advantages at places like Google. And so, I'm curious to hear your perspective. If the models, do you think the models are gonna commoditize? And if they do, what matters most? Is it distribution? Is it how well you build your applications? Is it something else that I'm Not thinking of.
A I don't think commoditization is quite the right framework to think about the models. There will be areas where different models excel at different things. For the kind of normal use cases of chatting with a model, maybe there will be a lot of great options. For scientific discovery, you will want the thing that's right at the edge that is optimized for science, perhaps. Um, so models will have different strengths, and the most economic value, I think, Will be created by models at the frontier and we plan to be ahead there. Um, and we're like very proud that five two is the best reasoning model in the world and the one that scientists are having the most progress with, but also, um, we're very proud that it's what enterprises are saying is the best at all of the tasks that a business needs to, to, you know, do its work. Um, so there will be, you know, times that we're ahead in some areas and behind in others, but the overall most intelligent model I expect to have Uh, significant value, even in a world where free models can do a lot of the stuff that people, that people need. The, the products will really matter. Distribution and brand, as you said, will really matter. Um, in ChatGPT, for example, personalization is extremely sticky. People love the fact that the model gets to know them over time, and you'll see us push on that, uh, much, much more. Um, people have experiences …
AI assessment note: “I don't think commoditization is quite the right framework to think about the models.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q planning elements for weeks now, and I can just come in in a new window and be like, alright, let's pick up on this trip, and it, it has the context, and it knows, it knows the guide I'm going with, it knows what I'm doing, uh, the fact that I've been, like, planning fitness for it, and can really synthesize all of those things. How good can memory get?
A I think we have no conception because the human limit, like even if you have the world's best personal assistant, They don't, they can't remember every word you've ever said in your life. They can't have read every email. They can't have read every document you've ever written. They can't be, you know, looking at all your work every day and remembering every little detail. They can't be a participant in your life to that degree, and no human has like infinite perfect memory. Um, And AI is definitely gonna be able to do that. And we actually talk a lot about this. Like right now, memory is still very crude, very early. We're in like the, you know, the GPT-II era of memory, but what it's gonna be like when it really does remember every detail of your entire life and personalized across all of that, and not just the facts, but like the little small preferences that you had that you maybe like didn't even think to indicate, but the AI can pick up on, uh, I think that's gonna be super powerful. That's one of the features that still, maybe not a twenty-twenty-six thing, but that's one of the parts of this I'm most excited for.
AI assessment note: “what it's gonna be like when it really does remember every detail”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q You've talked about building a cloud. Um, here's an email we got from a listener. At my company, we're moving off Azure and directly integrating with OpenAI to power our AI experiences in the product. The focus is to insert a stream of trillions of tokens powering AI experiences through the stack. Is that the plan to build a big, big cloud business in that, in that way?
A First of all, choice of tokens, a lot of tokens. And if, you know, you asked about the need for compute and our enterprise strategy, like enterprises have been clear with us about how many tokens they'd like to buy from us. And we are going to again fail in 2026 to meet demand. But the strategy is companies, most companies seem to want to come to A company like us, and say, I'd like to enable my company with AI. I need an API customized for my company. I need ChatGPT Enterprise customized for my company. I need a platform that can like run all these agents that I can trust my data on. I need the ability to get trillions of tokens into my product. I need the ability to have all my internal processes be more efficient, and we don't currently have like a great all-in-one offering for them, and we'd like to make that.
AI assessment note: “we don't currently have like a great all-in-one offering for them, and we'd like to make that.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q He's citing Yahshua Bengio. He's, he did his homework. You told him, this was right before GPT-V came out, that GPT-V is smarter than us in almost every way. Uh, I, I thought that that was the definition of AGI. Does, is that, isn't that AGI? And, and if not, has the term become somewhat meaningless?
A These models are clearly extremely smart on a sort of raw horsepower basis. You know, there's all this stuff on the last couple of days about GPT-FW as an IQ of 147 or 100 and 44 or 151 or whatever it is. It's like, you know, depending on whose test it's like, It's some high number, and you have, like, a lot of experts in their field saying it can do these amazing things, and it's, like, contributing, it's making it more effective, you have the GDP value, things we talked about. One thing you don't have is The ability for the model to not be able to do something today, realize it can't, go off and figure out how to learn to get good at that thing, learn to understand it, and when you come back the next day, it gets it right. And that kind of continuous learning, like, Toddlers can do it. It does seem to me like an important part of what we need to build. Now, can you have something that most people would consider an AGI without that? I would say clear. I mean, there's a lot of people that would say we're at AGI with our current models. Um, I think almost everyone would agree that if we were at the current level of intelligence and had that other thing, it would clearly be very AGI like, um, but maybe most of the world will say, Okay, fine, even without that, like, it's doing most knowledge tasks that matter, um, smarter than us, and most, most of us in most ways, we're at AGI. …
AI assessment note: “What I think this means is that the term... is very underdefined.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q compute? Like, potentially, like, so for instance, would we have surefire, like, scientific breakthroughs if, you know, OpenAI were to put double the compute towards science or, or with medicine? Like, would we have, you know, that clear ability to assist doctors? Like, how much of this is sort of a supposition of what's to happen versus Clear understanding based off of what you see today that it will happen.
A Everything based off what we see today is that it will happen. It does not mean some crazy thing can't happen in the future. Someone could discover some completely new architecture, and there could be a 10,000 times, you know, efficiency gain, and then we would have really probably overbuilt for a while. But everything we see right now about how quickly the models are getting better at each new level, how much more people want to use them, each time we can bring the cost down, How much more people really want to use them? Um, everything about that indicates to me that there will be increasing demand, and people using these for wonderful things, for silly things, um, but it, it just so seems like This is the shape of the future. Um, it's not just like, it's not just, you know, how many tokens we can do per day. It's how fast we can do them as these coding models have gotten better. They can think for a really long time, but you don't want to wait for a really long time. So there will be other dimensions. It will not just be the number of tokens that we can do. Um, but the demand for intelligence across a small number of axes And what we can do with those, you know, if you're like, if you have like a really difficult healthcare problem, do you want to use 5.2 or do you want to use 5.2 pro even if it takes dramatically more tokens?
AI assessment note: “Everything based off what we see today is that it will happen.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q of knowledge work tasks, and GPT 5.2 pro 74.1% of knowledge work tasks, and it passed a threshold of, um, Being expert level. It, it, it handled, it looks like something like 60% of expert tasks, uh, of tasks that would make it, you know, on par with an expert in the knowledge work. What are the implications of the fact that these models can do that much knowledge work?
A So, you know, you were asking about verticals, and I think that's a great question, but the thing that was going through my mind and why I kind of was stumbling a little bit is that eval, I think it's like 40 something different verticals that a business has to do. There's make a PowerPoint. Do this legal analysis, you know, write up this little web app, all this stuff. And, and the eval is, do experts prefer the output of the model relative to other experts? For a lot of the things that a business has to do. Now, these are small, well-scoped tasks. These don't get the kind of complicated, open-ended, creative work that, you know, figure out a new product. These don't get a lot of collaborative team things, but A coworker that you can assign an hour's worth of tasks to and get something you like better back 74 or 70% of time if you want to pay less is still pretty extraordinary. If you went back to the launch of ChatTBT three years ago and said we were going to have that in three years, most people would say absolutely not. And so as we think about how enterprises are going to integrate this, it's no longer like just that it can do code. It's all of these knowledge work tasks you can kind of farm out to the AI. Uh, and. That's gonna take a while to really kind of figure out how enterprises integrate with it, but should be quite substantial.
AI assessment note: “all of these knowledge work tasks you can kind of farm out to the AI”