The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Alexander Embiricos argument clarity score 4.3/5 from 24 exchanges on raw tape · average scores: directness 4.7 · coherence 4.6 · precision 4 · compression 3.8 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
70exchanges match
70on raw tape
3redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q That's really interesting. You said it's been a very good few weeks for us and I feel that. Does the team feel changing winds of momentum, both in positive and negative cycles?

A Absolutely. We, we are very attuned to it, right? Like, if you look at the history of Codex, the first thing we launched last year was, like, this amazing idea that people were super excited about. It's like, hey, we're gonna give the agent its own computer in the cloud. You're gonna have as many of them as you want work for you in parallel on tasks. Super great idea. To be honest, it didn't work as well as what we shipped later. It was not the best. Um, and then since August with GPT five, we started pushing really hard on interactive coding, which is where most of the competition in the market is. And, you know, we went on an absolute tear. I feel like the public metric we have was like, since August, we grew by like, 20 X. And then like, even like late in the year, we like doubled from December to now. I forget the exact number there, but like, You know, that was competing neck and neck, but the, the shift that we feel last week is, you know, we, we felt like we had the most intelligent model that was cemented with five free codecs. We had feedback around our model being slower and like maybe less fun to work with and like being less good at communicating with you while it was working. We addressed that feedback. Um, and that's true even compared to, like, the, the other competitor model that launched, like, 20 minutes before us and was like, maybe this is spicy. It was like…

AI assessment note: “Absolutely. We, we are very attuned to it, right?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Now, this is a weird first start, but roll with it. You'll, you'll understand my British intricacies. I'm fascinated by people's motivations. Are you motivated more by the fear of losing, or like the thrill and excitement of winning?

A I, I'm a maximalist. I'm definitely much more motivated by the idea of winning than the fear of losing. But I'll admit, I'll admit to you something. One, I was running a startup before generating OpenAI, and one of my darkest moments, and there were many dark moments while I was running the startup, was recognizing that I had spent the past few months trying to avoid losing. And all of a sudden, I was like, oh my god, that is why I'm so unhappy, and that's probably why the startup isn't going well. And so when we flipped You know, I basically every now and then I have to catch myself and like flip back into this idea of winning. But really what motivates me even more than that is I think I just love building things and building things for people. And man, I am so excited for this year because many amazing things that don't exist yet are going to be built and given to a lot of people.

AI assessment note: “I'm definitely much more motivated by the idea of winning than the fear of losing.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q how to engage in terms of developer experience, developer relations. Do you compete with a lovable and a rapid on a, like, low-end consumer basis in a year or two's time? Is that a business where you're like, you know what, Codex is not, For every person to create an about me or a small business to create their own site. How do you think about consumer in that way?

A Yeah, I would say that right now it doesn't feel like we're competing super directly. Um, but you know, I don't know if you saw our, our Superbowl ad, uh, the tagline of which is just, you can just build things. Um, with the app, we noticed that like many, many people who are less technical are starting to build things. And so the kinds of things they're building are much more hello worldy. And so I think that we will see some overlap in use cases, um, where you have, you know, people just pulling up codecs because they have it as part of their ChatGPT. Actually, like a big announcement last week was that we're now offering some codecs to people even on free ChatGPT plans or on the Go ChatGPT plan. So this is, this is massive just in terms of like bringing availability to everyone. Um, and so I think we're definitely going to see people with like a free chat GPT plan coming in and just like building simple things where they otherwise might have gone to a specialized tool.

AI assessment note: “right now it doesn't feel like we're competing super directly”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Final one for you, my friend. What are you most excited about when you look forward 10 years?

A This is probably going to happen in much less than 10 years, but, like, my mission, sort of personally, when I joined the company, was I just felt like Even with the models we had a year and a half ago, there was so much just capability overhang or just ability for these things to be useful, but we hadn't built the right products around that. And so people like me were getting more benefit than like people like my grandma. And so, what I'm most excited for is to get to, like, a form factor for AI that means that they're just helping everyone, regardless of whether they're in tech, and especially if they're not in tech, or especially if they're older. Um, and so, you know, the concrete vision I have is, like, at some point, we'll, like, add an agent to, um, like, our family WhatsApp or something, and it'll just start, like, being useful, um, to the family without anyone having to think harder about it than that. Um, there, there are many other ways that that could happen, but I think concretely, that's the most obvious thing we can do with like my grandma.

AI assessment note: “what I'm most excited for is to get to, like, a form factor for AI”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I'm a CS student, ok? I'm at Stanford, I'm at Imperial, I'm at Cambridge, I'm wherever, ETH. Great institution. What would you advise me, knowing all that you know now, that would help me navigate the next five years of my career? I want To be valuable to the AI ecosystem environment as an engineer entering the workforce in the next year.

A Basically there's actually never been a better time to be an engineer because you have incredible tooling available to you to get incredible amount done and your ability to like ramp into like a complex code base that you might be hired into has never been faster because you can go ask AI like a ton of questions about the code base. And you can ask it to plan out changes that would otherwise take you, like, days to research maybe. So, I think first off, I would say, like, you should be, like, very optimistic. But then of course, like about you, what your abilities once you're at the job, then all the question is, how do you get the job? I think that because it's never been like easier to build things, the thing that becomes scarcer is like agency taste and like quality. And so I would urge you to like, like just build things and, you know, demonstrate your agency and your taste around what you build and like build things that are of high quality and then share those things. Like, you know, we get a lot of inbound for, from folks, um, you know, both applying for jobs through the careers page or also on social. This is just me, but when someone writes to me with, like, some interesting thoughts and, like, a link to an interesting project, that gets my attention much more than, like, a normal resume does.

AI assessment note: “build things that are of high quality and then share those things”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Is weekly active a frequent enough metric, do you think? Sounds nice, but if this is actually replacing the IDE, is daily active not better?

A I think daily active will be better soon. Yeah. We just happened to use weekly active. It's like a standard here. And I think as we were getting started, it made sense. But I, I actually agree with the, the, the criticism there. It's like, we should probably just be a daily. Like, I think we, we need to be getting to a world where for any given task that you have, your first instinct is to ask an agent to help, right? It's kind of like, you know, how, like with Google search, it's just like, okay, anything I need to do, I just like go into this text box and I can get navigated to the right location. Then you had ChatGPT. It's like for any information I need, I can go into this text box, type it out and get information that helps me. And I think the next phase that we'll see this year is like for any task I need to do, as opposed to just get information, I go to this text box or this input and something happens that helps me, even if it's not the full task, even if it's only a small part of it.

AI assessment note: “I actually agree with the, the, the criticism there.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q How much weight should we place on benchmarks and evals?

A I think probably, uh, this is an annoying answer for you. It's like some, right? Like they do tell you, they kind of, in my mind, they give you a good measure of intelligence. Right. Um, and so you can put weight on those for intelligence. Um, and especially before evals are saturated, I think you, when you see meaningful progress in those benchmarks, it's like very, um, very helpful. Um, and then I think you have to pair that though with like what it feels like to use the model. And that's, that's a vibes thing. Like whenever I talk to any, like even internally or even talking to like customers of our models, I'm always surprised by how vibes based the evaluation of how it feels to work with the model is.

AI assessment note: “It's like some, right? Like they do tell you... a good measure of intelligence.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q mentioned our show on LinkedIn and a wonderful investor from a different company. It's like Harry Potter, you know, Voldemort. And it's like, you know, he who shall not be named. Um, I don't want Sam to kill me, but from another company, I was like, you've got, you ask him, ask him, how do you think about a coding data moat? And does Anthropic have all the data now?

A I think that from what we've seen, and you know, I would defer to my research team on this, but I feel like we, we feel like we have plenty enough data to build really good coding models. I actually think the, the place that's more interesting for getting data now is, like, as we get into, like, knowledge work tasks, that's kind of data that's, like, not really, like, available most places on the internet, and so you start to have, like, really interesting brainstorms for, like, how to help a model be good at it. Like, maybe you have to, like, Pay people, um, to, like, simulate doing tasks so that you can, like, learn these trajectories for the model. Maybe you should acquire startups, you know, that are no longer a business or that, and, uh, but have a lot of, like, data, like, say, they're Slack or something. Um, yeah, I think that, that kind of knowledge work task distribution is, like, much harder than coding.

AI assessment note: “we feel like we have plenty enough data to build really good coding models”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q terms of like market composition, as an investor, I have to think through how do I think about the eventual state of this given market, kind of a terminal state. How do you think about that? Is it like Uber and Lyft? And like the majority of the market will be on Codex or Claude Code, or is it like a AWS, Azure, Google Cloud, and a 33, 33, 33?

A I think this might end up with fewer providers that are capturing a lot of value in the long run. And here's why, like, and maybe this is a bit spicy, but I think that we are kind of in this temporary phase where we have agents that are really good at coding. Right, and, and if you look back last year, like, maybe more people thought we would have agents that are good at other domains, too, but that didn't happen last year. So we're, so we only have PMF for coding agents, like, in the industry overall, I would say, right? And then there's some, like, very narrow, narrow other use cases like customer support, et cetera. Um, but I think that's probably temporary, and then over time, I think we're gonna end up with agents that kind of can do anything for you. This is kind of what I was saying earlier. Like, there's just, like, a super assistant. You talk to it about anything. And then there is, like, specific, like, UI that you can go look at if you happen to be deep in a specific function. So in that world, I don't think you want, like, 12 agents at the company, and you have to, like, go, your employees have to go figure out the right one to talk to, because then they won't achieve fluency. And if they don't want to achieve fluency, then they will also won't, like, pull automation into their roles. But if you have this one thing that you can talk to about anything, right, so your…

AI assessment note: “this might end up with fewer providers that are capturing a lot of value”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What have you changed your mind on most in the last 12 months?

A When I joined OpenAI, I thought that, and this was a little longer than 12 months ago, but when I joined OpenAI, I thought that we would all just be hanging out with our computer screen sharing within a year from there. You know, we'd have this agent that we're just talking to. Um, that was completely wrong. Um, I think the rate of, like, progress in, like, multimodal models was, like, slower than I expected. Uh, multimodal means, you know, like, models that work with, like, video and audio. So, instead, what happened was that we saw that, like, agents that work with your computer through code are the way. And so, for me, that's been a complete rethink in terms of, like, how we bring the benefits of AI to, like, just people generally. It's not, not through video and audio, primarily.

AI assessment note: “When I joined OpenAI, I thought that... that was completely wrong.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I'm diving right in. Elon said that coding is one of the first professions to be largely automated. Do you agree, given your position and what you see day to day?

A I think for sure I would agree that coding is one of the first domains where LLMs are really good. You know, what does it mean for coding to be automated? It's like kind of a heavy statement, right? Like, for example, now that we no longer write assembly, like when that change happened and we moved to higher level languages, did we say coding is automated? Not really, right? We were just able to write much more code, and then as a result, actually, there was much more demand for code and there were many more software engineers required. But yeah, part of what they used to do is automate it in the same way that like, do you know the origin of the word computer? No. Um, I might pronounce the location wrong, but I think it was at Bletchley Park. There were all these machines for, like, decoding German Enigma, and, like, there were humans who would, like, punch out punch cards and, like, put them into the machine and do a bunch of, like, tabulated math. I'm probably butchering this, but basically there was an intensely manual part of work, and even, like, the first spreadsheet software was kind of loosely based off this idea that you would have an office full of desks arranged in a grid and people doing tabulations and then passing their sheets to the next person. And so all these things, like, Those specific tasks have become automated, but every time that's happened, there's been…

AI assessment note: “I think for sure I would agree that coding is one of the first domains”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. But does that not put the onus or the effort back on the user, back to the point of your bottleneck of human action and lack of activity on them? If you don't define the task, you put the responsibility on them for the defining the task, which humans lack the ability or inclination to do.

A Yeah. I think, so that's why I think it's the bottleneck. So, basically here are the three phases in my mind. First, Let's have agents work really well for software engineering and coding, because LLMs happen to be good at that. Next, let's realize that for an agent to be useful more generally, it using a computer is super valuable, and also we'll realize that all agents are actually coding agents, because coding is just the best way for an agent to use a computer. So let's take that same super flexible idea, but make it available to anyone who's excited to explore and tinker. And we're already seeing people start to do this with like the Codex app, like people, like Codex app is built for software builders. But we're seeing builders use it for all sorts of non-coding tasks. Then finally, once we see what's working, let's build that productization that you were talking about, where you have highly specific features that just work immediately out of the box for people. And I think we're going to speed run this entire, like, one, two, three journey, um, in the next months.

AI assessment note: “Then finally, once we see what's working, let's build that productization”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q mentioned our show on LinkedIn and a wonderful investor from a different company. It's like Harry Potter, you know, Voldemort. And it's like, you know, he who shall not be named. Um, I don't want Sam to kill me, but from another company, I was like, you've got, you ask him, ask him, how do you think about a coding data moat? And does Anthropic have all the data now?

A I think that from what we've seen, and you know, I would defer to my research team on this, but I feel like we, we feel like we have plenty enough data to build really good coding models. I actually think the, the place that's more interesting for getting data now is, like, as we get into, like, knowledge work tasks, that's kind of data that's, like, not really, like, available most places on the internet, and so you start to have, like, really interesting brainstorms for, like, how to help a model be good at it. Like, maybe you have to, like, Pay people, um, to, like, simulate doing tasks so that you can, like, learn these trajectories for the model. Maybe you should acquire startups, you know, that are no longer a business or that, and, uh, but have a lot of, like, data, like, say, they're Slack or something. Um, yeah, I think that, that kind of knowledge work task distribution is, like, much harder than coding.

AI assessment note: “we feel like we have plenty enough data to build really good coding models.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q the examples. My question to you is, I'm an investor today. I'm looking for companies which will accrue value over time and provide incredible products to customers. There is a belief that the durability of revenue of large SaaS companies today is zero, and that SaaS is dead, because the model providers, you, Anthropic, others, are going to come for our lunch, so to speak. What would you advise me?

A Like, things are built for humans, like, otherwise, what's the point? Right? Even, even SaaS tools are built for humans. And so, for me, I think my question is, like, does this SaaS company own a relationship with a human? On the other end of things. And if it does, then I suspect it's, it's not going away. Um, you know, or does the SaaS company own some like really important system of record? It's probably not going away. Maybe those, both of those two things, the interaction with the human and the system of record are like more important than ever, actually. On the other hand, is the SaaS company like a kind of a glue layer? Uh, but it doesn't own either of those two things. Well, I'm not the expert here, but I'm more nervous about that kind of company.

AI assessment note: “does this SaaS company own a relationship with a human... or... system of record”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Why do you think we don't need PMs in this world? You dangled the carrot.

A Yeah. Yeah. It's, it's my fun joke. I think, well, first of all, I think it's incredibly hard to define what a PM is, what a product manager is. I kind of think of the role as like, Actually, explicitly undefined, and your goal is just to adapt to whatever the team or business needs. And, you know, often if you have a bunch of people, like, say here, like, trying to build as quickly as possible, then what a, what a product manager can do is spend time, like, taking a few steps back and trying to look around corners and figure out what to do, you know, collaborate with the folks and go to market, and maybe be the, the team's, like, greatest cheerleader and quality raiser. But, like, all of those things I just described, which are maybe my current role, could be done by a really strong edge lead, Or a designer who thinks a lot about product. And so I think it's like often useful to have product managers, but you probably don't want many of them until the team is really large.

AI assessment note: “all of those things I just described... could be done by a really strong”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q If we assume a large percentage is done by Codex in terms of the code produced, how do you do coding reviews and is AI responsible for internal coding reviews?

A So, the, there are a few things here. Um, first off, the spec for what you want to do or the plan becomes more important than ever. Right? So, like, think, like, architecturally, like, how should this code work? Um, so, you know, we recently shipped, like, a very prominent plan mode that works a little differently than others where you have the agent go off and, like, propose how it's going to do something. It's, like, quite a long plan, and then asks you questions about if you agree on how it wants to do it or if you want to have input. And this is very similar to, like, if you had a new hire who was new to your code base and, um, You know, they had to present a sort of request for comments to the rest of the team before they started doing the work. So even though that's not formally code review, I would say review of the plan is actually something that's becoming more important because we're entering more of this like delegation phase of working with agents. So that's an underrated thing. Um, then, okay, there's actual code review. I think a problem that I hear a lot of people talking about, especially in the open source world, is like a lot of AI slop. Like people will just be submitting PRs to these open source repos. And they're trash, and like, maybe the user hasn't even, the person submitting the PR hasn't even tested them, or definitely hasn't reviewed the code. I think…

AI assessment note: “a common practice with Codex is to have Codex, like, review its own PR”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q is it like GTM, which is like the biggest enterprise in the world, do want to work with open AI. I have many friends in your sales team. The inbound that you get from the largest brands is incredible. So, GTM, because of the incredible brand, product execution and just codex being a freaking awesome product, or compute, inference, speed, actual, like, Compute advantage. Which one is the defining winner?

A Okay. So I think if we're going to talk about it more from an open AI perspective, obviously this is way above my pay grade, but I would say it's compute advantage and having the best models, right? And in order to achieve that, we then need to build businesses that generate revenue. And also that something we've, that's really interesting. We noticed with having the codex team, which is sort of Combined team of research and product is also by building these, these successful products, we create a lot of pressure to improve the model in sort of a faster way. So that's maybe the company perspective, right? If we come to the product perspective, I think the single most important thing we can do is build a, a really good product that people want to use. And like I was saying earlier, I think we really want to build products for individuals. And then allow, like, people to become fluent in those products and then, like, pull in automation. And I, I think that may be counterintuitive, but will result in way more impact than anyone purely approaching it from, like, the enterprise workflow perspective. Um, so, you know, I think that's mostly a question of product execution, and then that works for, say, like, prosumer. When it comes to enterprise, the go-to-market side is really important. Like, something that I've learned the hard way is if we go to an enterprise and we're just like,…

AI assessment note: “I would say it's compute advantage and having the best models, right?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Is weekly active a frequent enough metric, do you think? Sounds nice, but if this is actually replacing the IDE, is daily active not better?

A I think daily active will be better soon. Yeah. We just happened to use weekly active. It's like a standard here. And I think as we were getting started, it made sense. But I, I actually agree with the, the, the criticism there. It's like, we should probably just be a daily. Like, I think we, we need to be getting to a world where for any given task that you have, your first instinct is to ask an agent to help, right? It's kind of like, you know, how, like with Google search, it's just like, okay, anything I need to do, I just like go into this text box and I can get navigated to the right location. Then you had ChatGPT. It's like for any information I need, I can go into this text box, type it out and get information that helps me. And I think the next phase that we'll see this year is like for any task I need to do, as opposed to just get information, I go to this text box or this input and something happens that helps me, even if it's not the full task, even if it's only a small part of it.

AI assessment note: “I think daily active will be better soon. Yeah. We just happened to use weekly active.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q for busy humans. But I spoke to Anish Akaya, who's a GP at Andreessen, and he came out the other day, and he's like, no, no, no, this was created by Sam and Elon, and it works for very efficient people, but most of the planet want browser-based discovery, interactions, UIs. Do you think that chat will be the enduring UI in the next wave of AI interaction with humanity?

A The simple answer is yes, but actually I think there's two components here. Like, if we, if we just imagine the future, like, just like, let's think of some sci-fi movie, right? Like, what does AI look like? I, I, I believe that sci-fi is a really good predictor of what the simple, the future should look like, and usually it's pretty simple because it's a story, and I think simple is usually right. It's gonna be some just, like, entity that I can talk to however I want about whatever I want, right? I thought, like, I shouldn't have to navigate to a place where I work with, like, my coding AI, and then I have this, like, different place for my, like, sales AI, and I have to, like, be like, hey, I'm now talking to sales thing and, like, do that. It's just, like, I'm just gonna talk to a thing, and it's just gonna help. So I think what we're going to have is that we'll have chat or voice, basically conversational interface will be sort of the, the pillar of everything that you can talk to about anything. Um, and then you can add into any group chat or whatever so it can, like, discover how to help you. But then if you're, like, a power user and you're very good at a specific thing, you probably don't want to be disintermediated by having to talk to another person. It'd be like if you had an executive assistant, but you can only work by talking to them. That's, like, super annoying…

AI assessment note: “The simple answer is yes, but actually I think there's two components here.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q we spoke about, for example, going into large enterprises and how you can be helpful. I'm just using the most boring thing ever. Expense approval. You could have agent submission of expenses on my behalf for my trip to San Francisco, and then the agent on the flip side doing approvals for that from OpenAI's compliance department. Agent to agent. How do you think about that and that paradigm shift?

A My, like, quickest answer to this is that, like, we've noticed as we build Codex that the best, like, the best interfaces for Codex to do work are also tend to be the best interfaces for humans. So, like, when people ask, like, oh, like, how can I make my code base, like, more efficient for the agent to work with? The answer is often, like, well, have you looked at it yourself, and is it Is it easy for a human to work with? So, like, a very specific example would be, like, running tests in a code base. Naively, if you just, like, set up most test runners, they just, like, emit all the outputs of all the tests, and so, like, as a human, it's really annoying, because you have to go in and, like, find the one that failed, and it's like, you've got to read hundreds of thousands of lines. Turns out that's terrible for AI as well, but if you filter it down to just only emit the failed test, better for humans, also better for agents. So, Probably the agent to agent interaction points will be very similar to like, if there was a human in the loop. And that's nice because it means you can kind of atomically replace individual systems.

AI assessment note: “agent to agent interaction points will be very similar to like, if there was a human”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q how to engage in terms of developer experience, developer relations. Do you compete with a lovable and a rapid on a, like, low-end consumer basis in a year or two's time? Is that a business where you're like, you know what, Codex is not, For every person to create an about me or a small business to create their own site. How do you think about consumer in that way?

A Yeah, I would say that right now it doesn't feel like we're competing super directly. Um, but you know, I don't know if you saw our, our Superbowl ad, uh, the tagline of which is just, you can just build things. Um, with the app, we noticed that like many, many people who are less technical are starting to build things. And so the kinds of things they're building are much more hello worldy. And so I think that we will see some overlap in use cases, um, where you have, you know, people just pulling up codecs because they have it as part of their ChatGPT. Actually, like a big announcement last week was that we're now offering some codecs to people even on free ChatGPT plans or on the Go ChatGPT plan. So this is, this is massive just in terms of like bringing availability to everyone. Um, and so I think we're definitely going to see people with like a free chat GPT plan coming in and just like building simple things where they otherwise might have gone to a specialized tool.

AI assessment note: “right now it doesn't feel like we're competing super directly. Um, but you know”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q That's really interesting. You said it's been a very good few weeks for us and I feel that. Does the team feel changing winds of momentum, both in positive and negative cycles?

A Absolutely. We, we are very attuned to it, right? Like, if you look at the history of Codex, the first thing we launched last year was, like, this amazing idea that people were super excited about. It's like, hey, we're gonna give the agent its own computer in the cloud. You're gonna have as many of them as you want work for you in parallel on tasks. Super great idea. To be honest, it didn't work as well as what we shipped later. It was not the best. Um, and then since August with GPT five, we started pushing really hard on interactive coding, which is where most of the competition in the market is. And, you know, we went on an absolute tear. I feel like the public metric we have was like, since August, we grew by like, 20 X. And then like, even like late in the year, we like doubled from December to now. I forget the exact number there, but like, You know, that was competing neck and neck, but the, the shift that we feel last week is, you know, we, we felt like we had the most intelligent model that was cemented with five free codecs. We had feedback around our model being slower and like maybe less fun to work with and like being less good at communicating with you while it was working. We addressed that feedback. Um, and that's true even compared to, like, the, the other competitor model that launched, like, 20 minutes before us and was like, maybe this is spicy. It was like…

AI assessment note: “Absolutely. We, we are very attuned to it, right?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q How much weight should we place on benchmarks and evals?

A I think probably, uh, this is an annoying answer for you. It's like some, right? Like they do tell you, they kind of, in my mind, they give you a good measure of intelligence. Right. Um, and so you can put weight on those for intelligence. Um, and especially before evals are saturated, I think you, when you see meaningful progress in those benchmarks, it's like very, um, very helpful. Um, and then I think you have to pair that though with like what it feels like to use the model. And that's, that's a vibes thing. Like whenever I talk to any, like even internally or even talking to like customers of our models, I'm always surprised by how vibes based the evaluation of how it feels to work with the model is.

AI assessment note: “It's like some, right? Like they do tell you, they kind of”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q terms of like market composition, as an investor, I have to think through how do I think about the eventual state of this given market, kind of a terminal state. How do you think about that? Is it like Uber and Lyft? And like the majority of the market will be on Codex or Claude Code, or is it like a AWS, Azure, Google Cloud, and a 33, 33, 33?

A I think this might end up with fewer providers that are capturing a lot of value in the long run. And here's why, like, and maybe this is a bit spicy, but I think that we are kind of in this temporary phase where we have agents that are really good at coding. Right, and, and if you look back last year, like, maybe more people thought we would have agents that are good at other domains, too, but that didn't happen last year. So we're, so we only have PMF for coding agents, like, in the industry overall, I would say, right? And then there's some, like, very narrow, narrow other use cases like customer support, et cetera. Um, but I think that's probably temporary, and then over time, I think we're gonna end up with agents that kind of can do anything for you. This is kind of what I was saying earlier. Like, there's just, like, a super assistant. You talk to it about anything. And then there is, like, specific, like, UI that you can go look at if you happen to be deep in a specific function. So in that world, I don't think you want, like, 12 agents at the company, and you have to, like, go, your employees have to go figure out the right one to talk to, because then they won't achieve fluency. And if they don't want to achieve fluency, then they will also won't, like, pull automation into their roles. But if you have this one thing that you can talk to about anything, right, so your…

AI assessment note: “I think this might end up with fewer providers that are capturing a lot of value”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I'm a CS student, ok? I'm at Stanford, I'm at Imperial, I'm at Cambridge, I'm wherever, ETH. Great institution. What would you advise me, knowing all that you know now, that would help me navigate the next five years of my career? I want To be valuable to the AI ecosystem environment as an engineer entering the workforce in the next year.

A Basically there's actually never been a better time to be an engineer because you have incredible tooling available to you to get incredible amount done and your ability to like ramp into like a complex code base that you might be hired into has never been faster because you can go ask AI like a ton of questions about the code base. And you can ask it to plan out changes that would otherwise take you, like, days to research maybe. So, I think first off, I would say, like, you should be, like, very optimistic. But then of course, like about you, what your abilities once you're at the job, then all the question is, how do you get the job? I think that because it's never been like easier to build things, the thing that becomes scarcer is like agency taste and like quality. And so I would urge you to like, like just build things and, you know, demonstrate your agency and your taste around what you build and like build things that are of high quality and then share those things. Like, you know, we get a lot of inbound for, from folks, um, you know, both applying for jobs through the careers page or also on social. This is just me, but when someone writes to me with, like, some interesting thoughts and, like, a link to an interesting project, that gets my attention much more than, like, a normal resume does.

AI assessment note: “I would urge you to like, like just build things and, you know, demonstrate”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Final question, so we do a quick fire. You mentioned Dropbox earlier. It's, the, the alumni from Dropbox is incredible and really, like, amazing to see the talent that's come out of Dropbox. What's your single biggest lesson from Dropbox? That has shaped some of your thinking now with OpenAI.

A Oh, I don't need to think about that one. That's kind of the thing I was telling you about earlier, right? Like, I think when you're building tooling for people, like for end users, you have to think about, like, that tooling as a system of engagement, right? If people don't want to use your tool, if it doesn't like naturally feel like the easiest way to get something done, then people just won't use it, right? And so, like, again, I learned that from watching how Slack just absolutely took off. Um, and so I think about that a lot now when we're building these agents. I'm like, if we build our agent purely as like, you know, workflow automation, then it's always going to be like pulling teeth to get that thing started, right? You're going to need to hire Accenture or someone to come in. They're going to deploy FDEs. It's going to be tough. But if you can build a system that like people just love using, even if they only use it for partial tasks, over time, they'll get better and better at using it. And then that you'll get connected to the tools you want over time. And then you can start laddering in automation. Obviously these aren't mutually exclusive.

AI assessment note: “when you're building tooling for people... you have to think about... a system of engagement”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q How on earth do you reinvigorate growth at Dropbox today?

A At least from when I was at Dropbox, the thing we were uniquely good at was desktop software. And Desktop software is, it's funny, it was never not back, but anyways, it's so back. Um, basically because if you're solving for productivity and knowledge work, um, yes, there are systems of record everywhere that you need to connect with, but everything at the end of the day happens on the user's computer, either in their browser or, you know, just like locally in apps on their computer. And so I do think that the, the, the fastest way we're going to see productivity gains from agents at work Is going to be at first meeting users on their computer, working with the stuff that they have available to them, you know, without having deployed FTEs to set anything up. And then over time, you'll connect in these various systems. And so if I was Dropbox, I'd be thinking about how do we leverage our unique domain expertise in like building really good like desktop software and this sort of collaborative layer on top of your computer. How do we leverage that to enable productivity agents? It's a bit broad, but I think that's the angle you go for.

AI assessment note: “if I was Dropbox, I'd be thinking about how do we leverage our unique domain expertise”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. But does that not put the onus or the effort back on the user, back to the point of your bottleneck of human action and lack of activity on them? If you don't define the task, you put the responsibility on them for the defining the task, which humans lack the ability or inclination to do.

A Yeah. I think, so that's why I think it's the bottleneck. So, basically here are the three phases in my mind. First, Let's have agents work really well for software engineering and coding, because LLMs happen to be good at that. Next, let's realize that for an agent to be useful more generally, it using a computer is super valuable, and also we'll realize that all agents are actually coding agents, because coding is just the best way for an agent to use a computer. So let's take that same super flexible idea, but make it available to anyone who's excited to explore and tinker. And we're already seeing people start to do this with like the Codex app, like people, like Codex app is built for software builders. But we're seeing builders use it for all sorts of non-coding tasks. Then finally, once we see what's working, let's build that productization that you were talking about, where you have highly specific features that just work immediately out of the box for people. And I think we're going to speed run this entire, like, one, two, three journey, um, in the next months.

AI assessment note: “Yeah. I think, so that's why I think it's the bottleneck.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q the examples. My question to you is, I'm an investor today. I'm looking for companies which will accrue value over time and provide incredible products to customers. There is a belief that the durability of revenue of large SaaS companies today is zero, and that SaaS is dead, because the model providers, you, Anthropic, others, are going to come for our lunch, so to speak. What would you advise me?

A Like, things are built for humans, like, otherwise, what's the point? Right? Even, even SaaS tools are built for humans. And so, for me, I think my question is, like, does this SaaS company own a relationship with a human? On the other end of things. And if it does, then I suspect it's, it's not going away. Um, you know, or does the SaaS company own some like really important system of record? It's probably not going away. Maybe those, both of those two things, the interaction with the human and the system of record are like more important than ever, actually. On the other hand, is the SaaS company like a kind of a glue layer? Uh, but it doesn't own either of those two things. Well, I'm not the expert here, but I'm more nervous about that kind of company.

AI assessment note: “does this SaaS company own a relationship with a human? ... not going away”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q mentioned our show on LinkedIn and a wonderful investor from a different company. It's like Harry Potter, you know, Voldemort. And it's like, you know, he who shall not be named. Um, I don't want Sam to kill me, but from another company, I was like, you've got, you ask him, ask him, how do you think about a coding data moat? And does Anthropic have all the data now?

A I think that from what we've seen, and you know, I would defer to my research team on this, but I feel like we, we feel like we have plenty enough data to build really good coding models. I actually think the, the place that's more interesting for getting data now is, like, as we get into, like, knowledge work tasks, that's kind of data that's, like, not really, like, available most places on the internet, and so you start to have, like, really interesting brainstorms for, like, how to help a model be good at it. Like, maybe you have to, like, Pay people, um, to, like, simulate doing tasks so that you can, like, learn these trajectories for the model. Maybe you should acquire startups, you know, that are no longer a business or that, and, uh, but have a lot of, like, data, like, say, they're Slack or something. Um, yeah, I think that, that kind of knowledge work task distribution is, like, much harder than coding.

AI assessment note: “we feel like we have plenty enough data to build really good coding models.”

page 1 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.