The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Alexander Embiricos argument clarity score 4.3/5 from 24 exchanges on raw tape · average scores: directness 4.7 · coherence 4.6 · precision 4 · compression 3.8 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
70exchanges match
70on raw tape
3redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Why do you think we don't need PMs in this world? You dangled the carrot.

A Yeah. Yeah. It's, it's my fun joke. I think, well, first of all, I think it's incredibly hard to define what a PM is, what a product manager is. I kind of think of the role as like, Actually, explicitly undefined, and your goal is just to adapt to whatever the team or business needs. And, you know, often if you have a bunch of people, like, say here, like, trying to build as quickly as possible, then what a, what a product manager can do is spend time, like, taking a few steps back and trying to look around corners and figure out what to do, you know, collaborate with the folks and go to market, and maybe be the, the team's, like, greatest cheerleader and quality raiser. But, like, all of those things I just described, which are maybe my current role, could be done by a really strong edge lead, Or a designer who thinks a lot about product. And so I think it's like often useful to have product managers, but you probably don't want many of them until the team is really large.

AI assessment note: “all of those things I just described... could be done by a really strong edge lead”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q for busy humans. But I spoke to Anish Akaya, who's a GP at Andreessen, and he came out the other day, and he's like, no, no, no, this was created by Sam and Elon, and it works for very efficient people, but most of the planet want browser-based discovery, interactions, UIs. Do you think that chat will be the enduring UI in the next wave of AI interaction with humanity?

A The simple answer is yes, but actually I think there's two components here. Like, if we, if we just imagine the future, like, just like, let's think of some sci-fi movie, right? Like, what does AI look like? I, I, I believe that sci-fi is a really good predictor of what the simple, the future should look like, and usually it's pretty simple because it's a story, and I think simple is usually right. It's gonna be some just, like, entity that I can talk to however I want about whatever I want, right? I thought, like, I shouldn't have to navigate to a place where I work with, like, my coding AI, and then I have this, like, different place for my, like, sales AI, and I have to, like, be like, hey, I'm now talking to sales thing and, like, do that. It's just, like, I'm just gonna talk to a thing, and it's just gonna help. So I think what we're going to have is that we'll have chat or voice, basically conversational interface will be sort of the, the pillar of everything that you can talk to about anything. Um, and then you can add into any group chat or whatever so it can, like, discover how to help you. But then if you're, like, a power user and you're very good at a specific thing, you probably don't want to be disintermediated by having to talk to another person. It'd be like if you had an executive assistant, but you can only work by talking to them. That's, like, super annoying…

AI assessment note: “The simple answer is yes, but actually I think there's two components here.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q said about kind of FDs and implementation within enterprise is data security, sensitivity, permissioning, access provisions is really freaking hard, and people are much less intelligent and confident than we give them credit for, I think, especially in large enterprise, sorry. Um, and I think you actually need an FDE to go in and custom fit a lot of the different horizontal solutions to make it work. Am I wrong?

A I think you're right if you're trying to go, like, all the way from zero to one, and you have this, like, and I said, I don't mean grand negatively here, but if you have, like, a grand vision for some, like, ultimate workflow automation system, then yeah, you're going to have to clear through all of these security hurdles, all these, like, compliance hurdles that are really real, right? Build connections to all these data systems and, like, systems of record and action. Um, yeah, so you're going to need NFT to do that. What I've seen is that When we do these things top down, we end up, like, massively under leveraging the potential of AI in, like, helping that company. Whereas, If you can maybe do that in parallel, right? But if you can just give AI to the people, like, actually doing the work, um, they can start to, like, get a mental model for how AI can help, and then they can start pulling AI into their workflows at the same time. Here's just, like, an analogy or something here is, like, imagine if, um, you know, you work in, like, a customer support role, and AI is being brought into your role and starting to automate, like, meaningful chunks of your work. But you've never heard of Chachapiti, nor are you allowed to use it, right? So in this, in that scenario, you have, like, no intuition for what this thing is. Whereas in a world where actually you've been using Chachapit…

AI assessment note: “I think you're right if you're trying to go, like, all the way from zero”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Is weekly active a frequent enough metric, do you think? Sounds nice, but if this is actually replacing the IDE, is daily active not better?

A I think daily active will be better soon. Yeah. We just happened to use weekly active. It's like a standard here. And I think as we were getting started, it made sense. But I, I actually agree with the, the, the criticism there. It's like, we should probably just be a daily. Like, I think we, we need to be getting to a world where for any given task that you have, your first instinct is to ask an agent to help, right? It's kind of like, you know, how, like with Google search, it's just like, okay, anything I need to do, I just like go into this text box and I can get navigated to the right location. Then you had ChatGPT. It's like for any information I need, I can go into this text box, type it out and get information that helps me. And I think the next phase that we'll see this year is like for any task I need to do, as opposed to just get information, I go to this text box or this input and something happens that helps me, even if it's not the full task, even if it's only a small part of it.

AI assessment note: “I think daily active will be better soon. Yeah... we should probably just be a daily.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q how to engage in terms of developer experience, developer relations. Do you compete with a lovable and a rapid on a, like, low-end consumer basis in a year or two's time? Is that a business where you're like, you know what, Codex is not, For every person to create an about me or a small business to create their own site. How do you think about consumer in that way?

A Yeah, I would say that right now it doesn't feel like we're competing super directly. Um, but you know, I don't know if you saw our, our Superbowl ad, uh, the tagline of which is just, you can just build things. Um, with the app, we noticed that like many, many people who are less technical are starting to build things. And so the kinds of things they're building are much more hello worldy. And so I think that we will see some overlap in use cases, um, where you have, you know, people just pulling up codecs because they have it as part of their ChatGPT. Actually, like a big announcement last week was that we're now offering some codecs to people even on free ChatGPT plans or on the Go ChatGPT plan. So this is, this is massive just in terms of like bringing availability to everyone. Um, and so I think we're definitely going to see people with like a free chat GPT plan coming in and just like building simple things where they otherwise might have gone to a specialized tool.

AI assessment note: “I think that we will see some overlap in use cases”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q That's really interesting. You said it's been a very good few weeks for us and I feel that. Does the team feel changing winds of momentum, both in positive and negative cycles?

A Absolutely. We, we are very attuned to it, right? Like, if you look at the history of Codex, the first thing we launched last year was, like, this amazing idea that people were super excited about. It's like, hey, we're gonna give the agent its own computer in the cloud. You're gonna have as many of them as you want work for you in parallel on tasks. Super great idea. To be honest, it didn't work as well as what we shipped later. It was not the best. Um, and then since August with GPT five, we started pushing really hard on interactive coding, which is where most of the competition in the market is. And, you know, we went on an absolute tear. I feel like the public metric we have was like, since August, we grew by like, 20 X. And then like, even like late in the year, we like doubled from December to now. I forget the exact number there, but like, You know, that was competing neck and neck, but the, the shift that we feel last week is, you know, we, we felt like we had the most intelligent model that was cemented with five free codecs. We had feedback around our model being slower and like maybe less fun to work with and like being less good at communicating with you while it was working. We addressed that feedback. Um, and that's true even compared to, like, the, the other competitor model that launched, like, 20 minutes before us and was like, maybe this is spicy. It was like…

AI assessment note: “Absolutely. We, we are very attuned to it, right?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q terms of like market composition, as an investor, I have to think through how do I think about the eventual state of this given market, kind of a terminal state. How do you think about that? Is it like Uber and Lyft? And like the majority of the market will be on Codex or Claude Code, or is it like a AWS, Azure, Google Cloud, and a 33, 33, 33?

A I think this might end up with fewer providers that are capturing a lot of value in the long run. And here's why, like, and maybe this is a bit spicy, but I think that we are kind of in this temporary phase where we have agents that are really good at coding. Right, and, and if you look back last year, like, maybe more people thought we would have agents that are good at other domains, too, but that didn't happen last year. So we're, so we only have PMF for coding agents, like, in the industry overall, I would say, right? And then there's some, like, very narrow, narrow other use cases like customer support, et cetera. Um, but I think that's probably temporary, and then over time, I think we're gonna end up with agents that kind of can do anything for you. This is kind of what I was saying earlier. Like, there's just, like, a super assistant. You talk to it about anything. And then there is, like, specific, like, UI that you can go look at if you happen to be deep in a specific function. So in that world, I don't think you want, like, 12 agents at the company, and you have to, like, go, your employees have to go figure out the right one to talk to, because then they won't achieve fluency. And if they don't want to achieve fluency, then they will also won't, like, pull automation into their roles. But if you have this one thing that you can talk to about anything, right, so your…

AI assessment note: “I think this might end up with fewer providers that are capturing a lot of value”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q How on earth do you reinvigorate growth at Dropbox today?

A At least from when I was at Dropbox, the thing we were uniquely good at was desktop software. And Desktop software is, it's funny, it was never not back, but anyways, it's so back. Um, basically because if you're solving for productivity and knowledge work, um, yes, there are systems of record everywhere that you need to connect with, but everything at the end of the day happens on the user's computer, either in their browser or, you know, just like locally in apps on their computer. And so I do think that the, the, the fastest way we're going to see productivity gains from agents at work Is going to be at first meeting users on their computer, working with the stuff that they have available to them, you know, without having deployed FTEs to set anything up. And then over time, you'll connect in these various systems. And so if I was Dropbox, I'd be thinking about how do we leverage our unique domain expertise in like building really good like desktop software and this sort of collaborative layer on top of your computer. How do we leverage that to enable productivity agents? It's a bit broad, but I think that's the angle you go for.

AI assessment note: “How do we leverage that to enable productivity agents? It's a bit broad, but”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I'm a CS student, ok? I'm at Stanford, I'm at Imperial, I'm at Cambridge, I'm wherever, ETH. Great institution. What would you advise me, knowing all that you know now, that would help me navigate the next five years of my career? I want To be valuable to the AI ecosystem environment as an engineer entering the workforce in the next year.

A Basically there's actually never been a better time to be an engineer because you have incredible tooling available to you to get incredible amount done and your ability to like ramp into like a complex code base that you might be hired into has never been faster because you can go ask AI like a ton of questions about the code base. And you can ask it to plan out changes that would otherwise take you, like, days to research maybe. So, I think first off, I would say, like, you should be, like, very optimistic. But then of course, like about you, what your abilities once you're at the job, then all the question is, how do you get the job? I think that because it's never been like easier to build things, the thing that becomes scarcer is like agency taste and like quality. And so I would urge you to like, like just build things and, you know, demonstrate your agency and your taste around what you build and like build things that are of high quality and then share those things. Like, you know, we get a lot of inbound for, from folks, um, you know, both applying for jobs through the careers page or also on social. This is just me, but when someone writes to me with, like, some interesting thoughts and, like, a link to an interesting project, that gets my attention much more than, like, a normal resume does.

AI assessment note: “I would urge you to like, like just build things and, you know, demonstrate your agency”

Answered raw tape D 5 · C 5 · P 4 · Cm 3 4.45

Q I'm diving right in. Elon said that coding is one of the first professions to be largely automated. Do you agree, given your position and what you see day to day?

A I think for sure I would agree that coding is one of the first domains where LLMs are really good. You know, what does it mean for coding to be automated? It's like kind of a heavy statement, right? Like, for example, now that we no longer write assembly, like when that change happened and we moved to higher level languages, did we say coding is automated? Not really, right? We were just able to write much more code, and then as a result, actually, there was much more demand for code and there were many more software engineers required. But yeah, part of what they used to do is automate it in the same way that like, do you know the origin of the word computer? No. Um, I might pronounce the location wrong, but I think it was at Bletchley Park. There were all these machines for, like, decoding German Enigma, and, like, there were humans who would, like, punch out punch cards and, like, put them into the machine and do a bunch of, like, tabulated math. I'm probably butchering this, but basically there was an intensely manual part of work, and even, like, the first spreadsheet software was kind of loosely based off this idea that you would have an office full of desks arranged in a grid and people doing tabulations and then passing their sheets to the next person. And so all these things, like, Those specific tasks have become automated, but every time that's happened, there's been…

AI assessment note: “I think for sure I would agree that coding is one of the first domains”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Yeah. But does that not put the onus or the effort back on the user, back to the point of your bottleneck of human action and lack of activity on them? If you don't define the task, you put the responsibility on them for the defining the task, which humans lack the ability or inclination to do.

A Yeah. I think, so that's why I think it's the bottleneck. So, basically here are the three phases in my mind. First, Let's have agents work really well for software engineering and coding, because LLMs happen to be good at that. Next, let's realize that for an agent to be useful more generally, it using a computer is super valuable, and also we'll realize that all agents are actually coding agents, because coding is just the best way for an agent to use a computer. So let's take that same super flexible idea, but make it available to anyone who's excited to explore and tinker. And we're already seeing people start to do this with like the Codex app, like people, like Codex app is built for software builders. But we're seeing builders use it for all sorts of non-coding tasks. Then finally, once we see what's working, let's build that productization that you were talking about, where you have highly specific features that just work immediately out of the box for people. And I think we're going to speed run this entire, like, one, two, three journey, um, in the next months.

AI assessment note: “that's why I think it's the bottleneck. So, basically here are the three phases”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q it, how do you think about retention with this category? I remember Tom Blomfield, who's a YC partner, tweeted months and months ago, but it stuck with me, the weird brain, um, about the ease of transition between different Providers, whether it was Cursor or ClawCode or Codex. I can't remember which one it was, to be honest. But how sticky are users, and how do you think about retention?

A We've taken this, like, kind of counterintuitive approach with Codex to just build it super openly. So, like, the Codex core harness is open source, and we're always trying to make it easier for people to switch. So, for instance, um, when we first launched Codex last year, uh, we created, like, I mean, it's, Created is even a heavy word. It was just, we just established a convention, which is called agents.md. This is basically a file that you can put instructions for the agent in. And instead, we didn't call it codex.md. We just wanted it to be something that all agents can use. And pretty much every agent except Claude uses agents.md, which is awesome. And then just last week, actually, uh, we helped push for putting skills, which are a standard for like giving the agent instructions and scripts. We pushed for those to be sorted in sort of a neutral named folder called agents, um, instead of in like codex or something. And again, everyone has jumped on it except the usual suspect. Uh, so I think it's really great for the developers to have a lot of choice. Um, and we're trying to make it even easier for people to try different things. Now that said, I think These coding tasks, right, where you're asking an agent to write some code, they're quite hermetic. And what I mean by this is, it's like, or maybe an analogy in TV would be, like, episodic, right? Like, you can come in, …

AI assessment note: “we're always trying to make it easier for people to switch.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q is it like GTM, which is like the biggest enterprise in the world, do want to work with open AI. I have many friends in your sales team. The inbound that you get from the largest brands is incredible. So, GTM, because of the incredible brand, product execution and just codex being a freaking awesome product, or compute, inference, speed, actual, like, Compute advantage. Which one is the defining winner?

A Okay. So I think if we're going to talk about it more from an open AI perspective, obviously this is way above my pay grade, but I would say it's compute advantage and having the best models, right? And in order to achieve that, we then need to build businesses that generate revenue. And also that something we've, that's really interesting. We noticed with having the codex team, which is sort of Combined team of research and product is also by building these, these successful products, we create a lot of pressure to improve the model in sort of a faster way. So that's maybe the company perspective, right? If we come to the product perspective, I think the single most important thing we can do is build a, a really good product that people want to use. And like I was saying earlier, I think we really want to build products for individuals. And then allow, like, people to become fluent in those products and then, like, pull in automation. And I, I think that may be counterintuitive, but will result in way more impact than anyone purely approaching it from, like, the enterprise workflow perspective. Um, so, you know, I think that's mostly a question of product execution, and then that works for, say, like, prosumer. When it comes to enterprise, the go-to-market side is really important. Like, something that I've learned the hard way is if we go to an enterprise and we're just like,…

AI assessment note: “I would say it's compute advantage and having the best models”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q we spoke about, for example, going into large enterprises and how you can be helpful. I'm just using the most boring thing ever. Expense approval. You could have agent submission of expenses on my behalf for my trip to San Francisco, and then the agent on the flip side doing approvals for that from OpenAI's compliance department. Agent to agent. How do you think about that and that paradigm shift?

A My, like, quickest answer to this is that, like, we've noticed as we build Codex that the best, like, the best interfaces for Codex to do work are also tend to be the best interfaces for humans. So, like, when people ask, like, oh, like, how can I make my code base, like, more efficient for the agent to work with? The answer is often, like, well, have you looked at it yourself, and is it Is it easy for a human to work with? So, like, a very specific example would be, like, running tests in a code base. Naively, if you just, like, set up most test runners, they just, like, emit all the outputs of all the tests, and so, like, as a human, it's really annoying, because you have to go in and, like, find the one that failed, and it's like, you've got to read hundreds of thousands of lines. Turns out that's terrible for AI as well, but if you filter it down to just only emit the failed test, better for humans, also better for agents. So, Probably the agent to agent interaction points will be very similar to like, if there was a human in the loop. And that's nice because it means you can kind of atomically replace individual systems.

AI assessment note: “agent to agent interaction points will be very similar to like, if there was”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q it, how do you think about retention with this category? I remember Tom Blomfield, who's a YC partner, tweeted months and months ago, but it stuck with me, the weird brain, um, about the ease of transition between different Providers, whether it was Cursor or ClawCode or Codex. I can't remember which one it was, to be honest. But how sticky are users, and how do you think about retention?

A We've taken this, like, kind of counterintuitive approach with Codex to just build it super openly. So, like, the Codex core harness is open source, and we're always trying to make it easier for people to switch. So, for instance, um, when we first launched Codex last year, uh, we created, like, I mean, it's, Created is even a heavy word. It was just, we just established a convention, which is called agents.md. This is basically a file that you can put instructions for the agent in. And instead, we didn't call it codex.md. We just wanted it to be something that all agents can use. And pretty much every agent except Claude uses agents.md, which is awesome. And then just last week, actually, uh, we helped push for putting skills, which are a standard for like giving the agent instructions and scripts. We pushed for those to be sorted in sort of a neutral named folder called agents, um, instead of in like codex or something. And again, everyone has jumped on it except the usual suspect. Uh, so I think it's really great for the developers to have a lot of choice. Um, and we're trying to make it even easier for people to try different things. Now that said, I think These coding tasks, right, where you're asking an agent to write some code, they're quite hermetic. And what I mean by this is, it's like, or maybe an analogy in TV would be, like, episodic, right? Like, you can come in, …

AI assessment note: “we're always trying to make it easier for people to switch.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Dude, it's like pricing, grandfathering pricing is just, it's such a hard thing. What do we do today in engineering or product that in five years time you'll look back on and go, oh my God, can you believe that we did that?

A Well, one is just editing code by hand. Um, I think probably another one, this is maybe spicier, but another one might even be, uh, like actually managing the deployment and monitoring of, um, systems by hand. Like, I basically think that probably big companies will take a long time to, like, deploy this, but many startups might actually kind of start building on a completely new stack that's, like, fully AI managed. To be clear, the stack doesn't exist yet, but a fully managed AI stack where Because, like, basically it's been built to give you really strong deterministic guardrails over what the agent can do and, like, control over to, like, whirl back deploys and everything like that, and so we'll get to a world where the way you start a company is you start by getting an agent and just asking it to build things, and then you get more agents in that, and then maybe eventually you add, you add your co-founders to this service that you use to work with agents, and so you end up, like, maybe your main communication tool is actually your agent communication tool, And then maybe, ah, you're not actually, like, hand-holding this, like, very painful CI and deploy process, but you're just, like, having agents do things.

AI assessment note: “one is just editing code by hand. Um, I think probably another one”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q So you think we'll have more engineers in five years, not less?

A Yeah. And I, you know, sometimes we change what terms mean, right? Like the term computer now refers to something else, but now we have the term software engineer. And so I definitely think we'll have many more builders. You know, something interesting that I'm observing now is like, there's this compression of the talent stack. Like, you know, you still need software engineers today. You still need designers. I'm a PM. Do you need PMs? You know, you can have a fun, fun, some fun jokes about that. I don't think you need them. Um, but maybe, you know, maybe when you say engineer, you might be thinking of someone who's like much more full stack. Right. Then that has been true before. Like, even if you go back a few years, it was much, you had many more places where there was like the back end engineer and the front end engineer. Right. Whereas like now, at least if I think about the Codex team, like there's very few, like that's much less the case and things are much more full stack. Right. And so I think this, this talent stack will compress, but we'll still have people building.

AI assessment note: “Yeah... I definitely think we'll have many more builders.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q That's where I think, like, Claude has done well in terms of the packaging they've done, like Claude for legal, Claude for Excel, where you can implement it and have a DCF model. I'm not into models, but, like, better than one could do before. Do you think it is your job then to productize the prompts and the human actions to remove that bottleneck?

A Yeah, totally. So I, I think that it is our job to make sure that we have the models with the, with amazing capabilities. And then eventually to get to a world where this is, like, highly productized, and so you just have this, like, magic text box or audio input or whatever, or you can just add AI to your, like, group chat, and it just starts to help. But I think there's quite an interesting in-between stage, and I think that that is actually where the most value lies right now. So here's what I mean. You could try to productize, like, a specific feature of AI for a specific market, and, you know, many companies are doing this, But I think it's a little bit hard to know what exactly will work. What is the right form factor? You know, I, I, someone was on your podcast earlier and I, they said something that I thought was quite interesting about how you cannot adopt AI at enterprise without FDs.

AI assessment note: “Yeah, totally. So I, I think that it is our job”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q said about kind of FDs and implementation within enterprise is data security, sensitivity, permissioning, access provisions is really freaking hard, and people are much less intelligent and confident than we give them credit for, I think, especially in large enterprise, sorry. Um, and I think you actually need an FDE to go in and custom fit a lot of the different horizontal solutions to make it work. Am I wrong?

A I think you're right if you're trying to go, like, all the way from zero to one, and you have this, like, and I said, I don't mean grand negatively here, but if you have, like, a grand vision for some, like, ultimate workflow automation system, then yeah, you're going to have to clear through all of these security hurdles, all these, like, compliance hurdles that are really real, right? Build connections to all these data systems and, like, systems of record and action. Um, yeah, so you're going to need NFT to do that. What I've seen is that When we do these things top down, we end up, like, massively under leveraging the potential of AI in, like, helping that company. Whereas, If you can maybe do that in parallel, right? But if you can just give AI to the people, like, actually doing the work, um, they can start to, like, get a mental model for how AI can help, and then they can start pulling AI into their workflows at the same time. Here's just, like, an analogy or something here is, like, imagine if, um, you know, you work in, like, a customer support role, and AI is being brought into your role and starting to automate, like, meaningful chunks of your work. But you've never heard of Chachapiti, nor are you allowed to use it, right? So in this, in that scenario, you have, like, no intuition for what this thing is. Whereas in a world where actually you've been using Chachapit…

AI assessment note: “I think you're right if you're trying to go, like, all the way from zero to one”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q And so, is it like an inference monopoly? Like, you have it now, and competitors don't?

A This is just my opinion, but I don't think we're going to end up in, like, this kind of monopolistic world. I think there is so much competitive pressure that there'll be, like, multiple answers to this. But I will say that we have, like, news coming about, coming out about that partnership soon, and I'm very excited for these kinds of things to ship. It's going to be awesome. But even so, like, you know, with, uh, GPT-Codex, That model is, like, significantly, uh, more efficient than prior models, and so we've, in the feedback we've heard is that people actually feel like now this is, like, a very competitively fast model than before. So there's a lot of things you can do just in terms of the model. There are also things you can do, like, improving, uh, how you do inference. So we recently rolled out a change where in the API, like, those models are served, like, 40% faster, and in Codex they're served, like, a quarter fast, 25% faster. So I think, like, speed matters a lot, and we're kind of approaching it from all angles, like both the hardware, how you do inference, and the model level.

AI assessment note: “I don't think we're going to end up in, like, this kind of monopolistic world.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q That's where I think, like, Claude has done well in terms of the packaging they've done, like Claude for legal, Claude for Excel, where you can implement it and have a DCF model. I'm not into models, but, like, better than one could do before. Do you think it is your job then to productize the prompts and the human actions to remove that bottleneck?

A Yeah, totally. So I, I think that it is our job to make sure that we have the models with the, with amazing capabilities. And then eventually to get to a world where this is, like, highly productized, and so you just have this, like, magic text box or audio input or whatever, or you can just add AI to your, like, group chat, and it just starts to help. But I think there's quite an interesting in-between stage, and I think that that is actually where the most value lies right now. So here's what I mean. You could try to productize, like, a specific feature of AI for a specific market, and, you know, many companies are doing this, But I think it's a little bit hard to know what exactly will work. What is the right form factor? You know, I, I, someone was on your podcast earlier and I, they said something that I thought was quite interesting about how you cannot adopt AI at enterprise without FDs.

AI assessment note: “Yeah, totally. So I, I think that it is our job”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q is it like GTM, which is like the biggest enterprise in the world, do want to work with open AI. I have many friends in your sales team. The inbound that you get from the largest brands is incredible. So, GTM, because of the incredible brand, product execution and just codex being a freaking awesome product, or compute, inference, speed, actual, like, Compute advantage. Which one is the defining winner?

A Okay. So I think if we're going to talk about it more from an open AI perspective, obviously this is way above my pay grade, but I would say it's compute advantage and having the best models, right? And in order to achieve that, we then need to build businesses that generate revenue. And also that something we've, that's really interesting. We noticed with having the codex team, which is sort of Combined team of research and product is also by building these, these successful products, we create a lot of pressure to improve the model in sort of a faster way. So that's maybe the company perspective, right? If we come to the product perspective, I think the single most important thing we can do is build a, a really good product that people want to use. And like I was saying earlier, I think we really want to build products for individuals. And then allow, like, people to become fluent in those products and then, like, pull in automation. And I, I think that may be counterintuitive, but will result in way more impact than anyone purely approaching it from, like, the enterprise workflow perspective. Um, so, you know, I think that's mostly a question of product execution, and then that works for, say, like, prosumer. When it comes to enterprise, the go-to-market side is really important. Like, something that I've learned the hard way is if we go to an enterprise and we're just like,…

AI assessment note: “I would say it's compute advantage and having the best models, right?”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q we spoke about, for example, going into large enterprises and how you can be helpful. I'm just using the most boring thing ever. Expense approval. You could have agent submission of expenses on my behalf for my trip to San Francisco, and then the agent on the flip side doing approvals for that from OpenAI's compliance department. Agent to agent. How do you think about that and that paradigm shift?

A My, like, quickest answer to this is that, like, we've noticed as we build Codex that the best, like, the best interfaces for Codex to do work are also tend to be the best interfaces for humans. So, like, when people ask, like, oh, like, how can I make my code base, like, more efficient for the agent to work with? The answer is often, like, well, have you looked at it yourself, and is it Is it easy for a human to work with? So, like, a very specific example would be, like, running tests in a code base. Naively, if you just, like, set up most test runners, they just, like, emit all the outputs of all the tests, and so, like, as a human, it's really annoying, because you have to go in and, like, find the one that failed, and it's like, you've got to read hundreds of thousands of lines. Turns out that's terrible for AI as well, but if you filter it down to just only emit the failed test, better for humans, also better for agents. So, Probably the agent to agent interaction points will be very similar to like, if there was a human in the loop. And that's nice because it means you can kind of atomically replace individual systems.

AI assessment note: “Probably the agent to agent interaction points will be very similar to like, if there”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Final question, so we do a quick fire. You mentioned Dropbox earlier. It's, the, the alumni from Dropbox is incredible and really, like, amazing to see the talent that's come out of Dropbox. What's your single biggest lesson from Dropbox? That has shaped some of your thinking now with OpenAI.

A Oh, I don't need to think about that one. That's kind of the thing I was telling you about earlier, right? Like, I think when you're building tooling for people, like for end users, you have to think about, like, that tooling as a system of engagement, right? If people don't want to use your tool, if it doesn't like naturally feel like the easiest way to get something done, then people just won't use it, right? And so, like, again, I learned that from watching how Slack just absolutely took off. Um, and so I think about that a lot now when we're building these agents. I'm like, if we build our agent purely as like, you know, workflow automation, then it's always going to be like pulling teeth to get that thing started, right? You're going to need to hire Accenture or someone to come in. They're going to deploy FDEs. It's going to be tough. But if you can build a system that like people just love using, even if they only use it for partial tasks, over time, they'll get better and better at using it. And then that you'll get connected to the tools you want over time. And then you can start laddering in automation. Obviously these aren't mutually exclusive.

AI assessment note: “when you're building tooling for people... you have to think about that tooling as a system of engagement”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q I'm diving right in. Elon said that coding is one of the first professions to be largely automated. Do you agree, given your position and what you see day to day?

A I think for sure I would agree that coding is one of the first domains where LLMs are really good. You know, what does it mean for coding to be automated? It's like kind of a heavy statement, right? Like, for example, now that we no longer write assembly, like when that change happened and we moved to higher level languages, did we say coding is automated? Not really, right? We were just able to write much more code, and then as a result, actually, there was much more demand for code and there were many more software engineers required. But yeah, part of what they used to do is automate it in the same way that like, do you know the origin of the word computer? No. Um, I might pronounce the location wrong, but I think it was at Bletchley Park. There were all these machines for, like, decoding German Enigma, and, like, there were humans who would, like, punch out punch cards and, like, put them into the machine and do a bunch of, like, tabulated math. I'm probably butchering this, but basically there was an intensely manual part of work, and even, like, the first spreadsheet software was kind of loosely based off this idea that you would have an office full of desks arranged in a grid and people doing tabulations and then passing their sheets to the next person. And so all these things, like, Those specific tasks have become automated, but every time that's happened, there's been…

AI assessment note: “I think for sure I would agree that coding is one of the first domains”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q So you think we'll have more engineers in five years, not less?

A Yeah. And I, you know, sometimes we change what terms mean, right? Like the term computer now refers to something else, but now we have the term software engineer. And so I definitely think we'll have many more builders. You know, something interesting that I'm observing now is like, there's this compression of the talent stack. Like, you know, you still need software engineers today. You still need designers. I'm a PM. Do you need PMs? You know, you can have a fun, fun, some fun jokes about that. I don't think you need them. Um, but maybe, you know, maybe when you say engineer, you might be thinking of someone who's like much more full stack. Right. Then that has been true before. Like, even if you go back a few years, it was much, you had many more places where there was like the back end engineer and the front end engineer. Right. Whereas like now, at least if I think about the Codex team, like there's very few, like that's much less the case and things are much more full stack. Right. And so I think this, this talent stack will compress, but we'll still have people building.

AI assessment note: “Yeah. And I, you know, sometimes we change what terms mean, right?”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Dude, it's like pricing, grandfathering pricing is just, it's such a hard thing. What do we do today in engineering or product that in five years time you'll look back on and go, oh my God, can you believe that we did that?

A Well, one is just editing code by hand. Um, I think probably another one, this is maybe spicier, but another one might even be, uh, like actually managing the deployment and monitoring of, um, systems by hand. Like, I basically think that probably big companies will take a long time to, like, deploy this, but many startups might actually kind of start building on a completely new stack that's, like, fully AI managed. To be clear, the stack doesn't exist yet, but a fully managed AI stack where Because, like, basically it's been built to give you really strong deterministic guardrails over what the agent can do and, like, control over to, like, whirl back deploys and everything like that, and so we'll get to a world where the way you start a company is you start by getting an agent and just asking it to build things, and then you get more agents in that, and then maybe eventually you add, you add your co-founders to this service that you use to work with agents, and so you end up, like, maybe your main communication tool is actually your agent communication tool, And then maybe, ah, you're not actually, like, hand-holding this, like, very painful CI and deploy process, but you're just, like, having agents do things.

AI assessment note: “one is just editing code by hand.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Dude, it's like pricing, grandfathering pricing is just, it's such a hard thing. What do we do today in engineering or product that in five years time you'll look back on and go, oh my God, can you believe that we did that?

A Well, one is just editing code by hand. Um, I think probably another one, this is maybe spicier, but another one might even be, uh, like actually managing the deployment and monitoring of, um, systems by hand. Like, I basically think that probably big companies will take a long time to, like, deploy this, but many startups might actually kind of start building on a completely new stack that's, like, fully AI managed. To be clear, the stack doesn't exist yet, but a fully managed AI stack where Because, like, basically it's been built to give you really strong deterministic guardrails over what the agent can do and, like, control over to, like, whirl back deploys and everything like that, and so we'll get to a world where the way you start a company is you start by getting an agent and just asking it to build things, and then you get more agents in that, and then maybe eventually you add, you add your co-founders to this service that you use to work with agents, and so you end up, like, maybe your main communication tool is actually your agent communication tool, And then maybe, ah, you're not actually, like, hand-holding this, like, very painful CI and deploy process, but you're just, like, having agents do things.

AI assessment note: “Well, one is just editing code by hand.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q So you think we'll have more engineers in five years, not less?

A Yeah. And I, you know, sometimes we change what terms mean, right? Like the term computer now refers to something else, but now we have the term software engineer. And so I definitely think we'll have many more builders. You know, something interesting that I'm observing now is like, there's this compression of the talent stack. Like, you know, you still need software engineers today. You still need designers. I'm a PM. Do you need PMs? You know, you can have a fun, fun, some fun jokes about that. I don't think you need them. Um, but maybe, you know, maybe when you say engineer, you might be thinking of someone who's like much more full stack. Right. Then that has been true before. Like, even if you go back a few years, it was much, you had many more places where there was like the back end engineer and the front end engineer. Right. Whereas like now, at least if I think about the Codex team, like there's very few, like that's much less the case and things are much more full stack. Right. And so I think this, this talent stack will compress, but we'll still have people building.

AI assessment note: “Yeah... I definitely think we'll have many more builders.”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q If we assume a large percentage is done by Codex in terms of the code produced, how do you do coding reviews and is AI responsible for internal coding reviews?

A So, the, there are a few things here. Um, first off, the spec for what you want to do or the plan becomes more important than ever. Right? So, like, think, like, architecturally, like, how should this code work? Um, so, you know, we recently shipped, like, a very prominent plan mode that works a little differently than others where you have the agent go off and, like, propose how it's going to do something. It's, like, quite a long plan, and then asks you questions about if you agree on how it wants to do it or if you want to have input. And this is very similar to, like, if you had a new hire who was new to your code base and, um, You know, they had to present a sort of request for comments to the rest of the team before they started doing the work. So even though that's not formally code review, I would say review of the plan is actually something that's becoming more important because we're entering more of this like delegation phase of working with agents. So that's an underrated thing. Um, then, okay, there's actual code review. I think a problem that I hear a lot of people talking about, especially in the open source world, is like a lot of AI slop. Like people will just be submitting PRs to these open source repos. And they're trash, and like, maybe the user hasn't even, the person submitting the PR hasn't even tested them, or definitely hasn't reviewed the code. I think…

AI assessment note: “a common practice with Codex is to have Codex, like, review its own PR”

← previous page 2 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.