Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I want to start with what's feeling like a barometer of progress in AI, especially in engineering. What percentage of your code, if you even write code anymore, and your team's code is written by AI at this point?
A I do write code occasionally now, still. I'd actually say for managers like myself, it's way easier to use these AI tools, uh, than to manually code at this point. And so I know for myself and some of the other EMs, Engineering managers at OpenAI. Uh, all of our code is written by, by Codex, uh, at this point. But more broadly, there's just been this, there's just so much energy. There's like a tangible energy internally around just how far these tools have gotten, how good Codex as a tool has gotten for us. And, uh, it's, it's a little hard for us to exactly measure how much of the code is, is written because the vast majority of it, I'd say like close to a hundred percent is, is usually generated by AI first. Uh, what we do track, though, is, is, you know, at this point, uh, the vast majority of engineers use Codex on a daily basis. So, 95% of engineers, um, use Codex. Um, 100% of our PRs are reviewed by Codex daily as well, so basically any code that goes into production that's merged in, Codex kind of has its eyes on and, uh, suggests improvements, suggests changes, uh, uh, in the PRs. And so, uh, that's kind of what we're seeing internally, but by and large, the most exciting is just the energy that, that there, that there is.
AI assessment note: “close to a hundred percent is, is usually generated by AI first.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So to follow that thread, where are, like in the next six to 12 months, where is the API heading? Where's the platform heading? Where are the models heading? As much as you can share, I know there's a lot of secrets here that maybe you're more excited about, or do you think that people should start to Prepare for it and however much you can share.
A I mean, so the obvious one is, um, how long of a task, uh, these models can do coherently. Um, so there's like the, the meter benchmark that, that I think tracks software engineering tasks and how long, you know, like how long of a task can these models do, uh, 50% of the time, 80% of the time. Uh, I think we're at something like multi-hour tasks being able to be done by, uh, software engineering tasks being able to be done by, um, uh, these frontier models. Uh, 50% of the time, and then I think 80% is something, like, just under an hour. But the, the, the sobering thing about that, that chart is they plot all the, uh, previous models on this chart as well, so you can really see the trend of this. That's something that I'm really excited about, which is, you know, I actually think products today really optimize for tasks that the model can do for, like, minutes at a time. Like, even codecs and, like, the coding tools, I'd say, like, you know, it's, it's in the CLI. You're kind of, like, seeing it be interactive. It's really, you know, Quite optimized well for, like, maybe at most 10 minute type tasks. I have seen people push codecs to the limit into, like, multi-hour long, uh, tasks. Uh, but again, I, I, I, I think that that's more of the exception. But I, uh, if you follow this trend, like, I think, like, in the next 12 to 18 months, we could see models that could do multi-hou…
AI assessment note: “the obvious one is, um, how long of a task, uh, these models can do coherently.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q dealing with that is hard to scale in my experience. So unless you have, in my opinion, unless you have a bunch of contractors, which I don't know, does that count as a single person company? I feel like it's very difficult to scale a billion dollar startup and not have someone helping you with at least the support work. And AI I think will only take you so far.
A So I, I think that's true. Uh, and actually, I think my view on it is, is, is slightly different, which is I think that your, you know, Lenny's podcast might end up becoming a billion dollar startup. But, um, what I think might happen is, uh, instead of you kind of being the one person who has to dispatch an AI to solve and fix those support tickets, I think what might end up happening is there might be a whole smattering of other startups. That are building software and super, and like super tailored towards what you might need. And so, you know, uh, there might be like 10 or 20 startups that build support software for podcasts and newsletters. And, uh, that might be a one-person startup. Like, it doesn't need to be a big one. And, uh, it's, it's, and, you know, they might be able to just code up this product very, very easily. They're able to kind of like build their own thing. And because it's so tailored and unique and hopefully, you know, useful for you, it might be something that you purchase Um, as the one person billion dollar startup. I would buy that.
AI assessment note: “instead of you kind of being the one person who has to dispatch an AI”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q Is, is that under your umbrella, by the way, or is that a different org and team?
A It's a, it's a different team. So it's under ChatGPT. We obviously collaborate very closely with them. And, uh, you know, they built like an apps SDK, uh, which is a built-in close collaboration with our team. Uh, but that is more within the ChatGPT umbrella. Uh, but that is also another, like, that's another example of this, right? It's like ChadCBT is, like, we, we, we, we kind of, like, have these eight hundred million weekly active users who are just coming over and over again, like, it's a great asset to have as a business, but, like, man, would it be better if we could somehow allow, you know, uh, other companies to come in and, and, and, uh, take advantage of this as well, and, and build for this, this audience as well. And, and then ultimately, we think it'll help us expand that, that, that group as well, right? And so, it's all, it all kind of comes back to the mission, and, Uh, we find that being a platform, being open tends to help here.
AI assessment note: “It's a, it's a different team. So it's under ChatGPT.”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q Codex writing the code, codex reviewing its own code. I'm curious if you are open to using other models to review your model's work. Is that, is that a path or is it just, it's good enough? We don't need anything else.
A So I will say there's, there's definitely a circular thing here, and like going back to Sorcerer's Apprentice, like you want to make sure you're not letting the brooms go crazy here. Um, and so, you know, we're very thoughtful, I'd say, around which PRs kind of are completely just codecs, uh, reviewed. Most people still obviously take a look at their PRs, uh, and so it's not like it's going to zero. It's more like going from, you know, a hundred percent attention to like 30% attention, which, which just helps Things push through. Uh, in terms of, like, multiple models, uh, so we, we obviously test a lot of models internally, and so we have a lot of those. Um, we use, uh, external models less. Um, it's, we, we think it's important to kind of dog food our own models and kind of, like, get feedback there, but, uh, you can also, you know, there are a lot of, like, internal variants of models that you can use to give you different perspectives, um, here as well, and, and we found that to, to work quite well.
AI assessment note: “internal variants of models that you can use to give you different perspectives”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Okay. Great tip. Okay. Uh, favorite recent movie or TV show you have really enjoyed?
A Yeah, that one's tough. Cause you know, with, I, I have two kids and, uh, Uh, a busy job, and so I really haven't had much time, um, to watch TV shows. Uh, I will say in the last couple weeks, I watched a couple episodes. I'm actually a big anime guy, and so, uh, I watched a couple episodes. There's a new season of this anime called Jujutsu Kaisen, uh, that's out. Uh, so season three of JJK, uh, was, was, was really good. Um, in general, uh, I'm a huge, uh, fan of, uh, Japanese anime. I think they create the most, uh, Novel and unique, uh, plots, uh, universes that, uh, Western media has shied away from. Um, and so, uh, generally a big fan of that, but yeah, haven't really watched much, but saw a couple upsets with JJK recently.
AI assessment note: “There's a new season of this anime called Jujutsu Kaisen, uh, that's out.”