Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q what's coming next, or, uh, maybe this is probably, it probably is the most impatient industry in the world and the most impatient users in the world, but Um, it seems to me like the expectations for GPT-V are built up pretty high, and so I'm curious from, like, your perspective, um, do you think it's gonna be hard to meet those expectations whenever that GPT-V model does come out?
A Well, I don't think so, and one of the fundamental reasons is because we now have two different axes on which we can scale, right? Um, so GPT-V, this is our latest scaling experiment along the axis of unsupervised learning, but there's also reasoning. Um, and when you ask about, kind of, like, uh, why there seems to be, you know, a little bit bigger of a gap in release time between four and 4.5, we've been really largely focused on developing the reasoning parallel paradigm as well. So, um, I think, you know, our research program is really an exploratory research program, right? Um, we're looking into all avenues of how we can scale our models, and over the last, you know, one and a half, two years, We've really found a new, very exciting paradigm through reasoning, which we're also scaling. Um, and, and so I think, like, uh, GPT-V really could be the culmination of a lot of these things coming together.
AI assessment note: “Well, I don't think so, and one of the fundamental reasons is”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q For me, the future is very much aligned with niche models existing in workflows and less so of these general purpose God models. Um, so clearly open AI has a different thesis here, and I am curious to hear your perspective on what we get with the big models versus the niche models. And do you see them in competition or as compliments? Help us think about, think through that.
A Yeah. Yeah. So I think one important thing is we also serve models that are smaller, right? Like we serve our flagship frontier models, but we also serve mini models, right? Which are cost efficient ways that you can access the capabilities or fairly close to frontier capabilities for much lower cost, right? And we think that's an important part of this comprehensive portfolio here. Um, fundamentally at OpenAI though, uh, we're in the business of advancing the frontier of intelligence. And that involves developing the best models that we can. Um, and I, I think really kind of what we're motivated by is really pushing that out as much as possible. Um, we think there's always going to be use cases at the frontiers of intelligence. Um, you know, we, we think that, you know, going from 99.9 percentile in, in mathematics to the best in the world in mathematics, right? Like that difference means something to us. Like I think, uh, what You know, the best human scientists can discover is tangibly different, right, from what you or I can, can discover. So, um, we're, we're motivated by pushing the intelligence frontier as far as possible. And at the same time, uh, we want to make these capabilities cheaper and more cost effective to serve for everyone. So we don't think the niche models will go away. We want to build these foundation models and also figure out how to deliver these capab…
AI assessment note: “we don't think the niche models will go away. We want to build these foundation models”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah, but, but again, just to hammer home on this, like, what does that improvement in model get you? Like, what do you think that it will enable?
A Yeah, yeah. So, I mean, I think, ah, I mean, agents of all forms, right? When you look at stuff like deep research, for instance, right? Um, it gives you the ability to essentially kind of Get a fully formed report on any single topic that you might be interested in, right? Um, I've used it to even put together, like, hour-long talks, um, and it goes and really kind of, like, synthesizes all the information out there and, and really organizes it, comes up with lessons, allows you to do deep discovery, um, it allows you to, uh, You know, like dig into almost any topic that you're interested in. So I feel like, um, just the amount of information and synthesis that's, that's available to you now is, is just really rapidly evolving.
AI assessment note: “I mean, agents of all forms, right? When you look at stuff like deep research”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q you could probably do more with a better model. But I have to be honest, I'm kind of at a loss for words sometimes about what that getting from that 99th percentile in math to the best in world in math will do. So actually I'm curious to hear your answer on this one. Um, what does, what does building the best model in the world do that? Yeah. Yeah.
A A hundred percent. And I think really, um, it signals a shift, right? Like I, I think if you just think about, Hey, you take the current models and you build the best surface for them, that's certainly something you should always be doing and exploring that exercise. I think three years ago that looked like chat, right? We, we launched chat GBT. Um, and today when it look, when you take the best models and the best capabilities, I think it looks a little bit more like agents, right? Um, and I think reasoning and agents, they're, they're very, very much coupled, right? Um, when you think about what makes a good agent, it's something that you can kind of sit back, let it do its own thing, and you're fairly confident it'll come back with something that you want, right? And I think reasoning is the engine that powers that, right? Like, uh, you, uh, have the model go and try something out, and if it, um, if it can't succeed on the first try, it should be able to be like, oh, well, why didn't I succeed, and what's a better approach for me to do? So, um, you know, I, I think very much kind of like, ah, the capabilities are always changing, and the surface is always changing as a, as a response, and we're always exploring what the best surface for the current capabilities looks like.
AI assessment note: “when you take the best models... it looks a little bit more like agents”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q So what are the examples, like you talk about everyday knowledge work, what are some of the examples that you would use GPT-FORO Or that you would prefer it over a reasoning model?
A Yeah, um, so I, I wouldn't say like, it's a, it's a different profile from, from a reasoning model, right? Um, so with a larger model, um, what you're doing is, it, it takes more time to kind of process and think through the query, but it's also giving you an immediate response back. So this is very similar to what a GPT-IV would have, would have done for you, right? Um, whereas I think, um, with something like O-one, You get a model where you give a query and it can think for several minutes. Um, and, and I think these are fundamentally kind of different trade-offs, right? You have a model that immediately comes back to you, doesn't do much thinking, um, uh, but comes up with a better answer versus a model that, you know, thinks for a while, um, and then comes comes up with you with an answer. And, you know, we find that in a lot of areas like creative writing, for instance, um, Again, this is stuff that we want to test over the next one or two months, um, but, uh, we find that there, there are areas like creative writing where this model outshines reasoning models.
AI assessment note: “we find that there are areas like creative writing where this model outshines”
Partly raw tape
D 3 · C 4 · P 3 · Cm 4 3.45
Q them helping them get more efficient. So I just want to turn it over to you, um, without commenting on what they did, or if you can, if you want, but I'm actually more curious what open AI is doing on that front and what sort of, whether you did similar optimizations with GPT, 4.5, and are you able to run these large models more efficiently? And if so, how?
A Yeah, so I would say, um, kind of the process of making a model efficient to serve, I often see as fairly decoupled from developing the core capability of the model, right? Um, and we see a lot of work being done on the inference stack, right? I think that's something that DeepSeek did very well, um, and it's also something that we push on a lot, right? Um, we care about serving these models at cheap cost to all users, um, and we push on that quite a bit. Um, so I think this is irrespective of, you know, GPT-IV reasoning models. We're always applying that pressure to be able to inference more cheaply, and, and I think we've done a good job of that over time, right? Like, ah, the costs have dropped, you know, many orders of magnitude since we first launched GPT-IV.
AI assessment note: “we see a lot of work being done on the inference stack”