Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Is there a list published anywhere where founders who are listening in can go and see what they should be building or areas where there's more demand than either people understand or is anticipated or where the DOW is willing to help provide capital or other things to accelerate it?
A So the DIU for defensive issue has like four or five categories of things that they look at. So that's, that's sort of a published and public thing. The Office of Strategic Capital is a little more proactive. They're going to say, we have this whole, I'm going to go find companies that already exist that do this. So it's not for startups because it's a lending facility, if you will. So they're looking for companies that already exist and how do I expand their, their production? Um, but it's a good point. You know, at some point, I was talking to the secretary about this the other day. It's like, we should have some forum for, for talking about the things that we need more proactively so investors and entrepreneurs can think about them.
AI assessment note: “DIU for defensive issue has like four or five categories of things that they look at”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q like, the research and the product effort? Does that make sense? Or like thinking about new markets? And maybe wrapped up in that question too is just like, Well, where are we in quality on, on voice as well? Because if, if I, I would sort of claim like if the models are not good enough for certain use cases at all, like kind of doesn't make sense. Do product?
A And I think that's right. It's, it's almost exactly like when we, when we started originally, what we, what we did was try to actually use existing models that were in the market and kind of optimize them for our first use case was actually starting with combination of, of narration and dubbing, and then on that creative side. And, um, We realized pretty quickly that the models that existed just produced such a robotic and, and not, not good speech that people didn't want to listen to it. And that's where my co-founder's genius came in, where he was able to assemble the team and, and do a lot of the research himself to actually create new version of, of creating that work. But like to your question, I think that the way we are kind of organized internally and how we think about sequencing a lot of that was looking at the first problem. And then creating effectively a lab around that problem, which is like a combination of mighty researchers, engineers, operators to go after that problem. And the first problem was the problem of voice. So how can we recreate the, the voice? And like you say, it needs to have that research expertise to be able to do that well. So we started with effectively a voice lab, which was that mission of, can we narrate the work in, in, in a better way? There was a combination of roughly five people that were That we're doing that work, and then sequence …
AI assessment note: “sequence the research first, and then build a simple layer on top of that”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q imagine for a lot of your customers, it's not like they, like, know how to choose good voice. So how do you, how do you deal with that problem? Like, is it like, hey, I make a clone, and like, that sounds like me, and I believe it, and I'm gonna try all of these different options, or, or, or, You know, actually, are you teaching people to do eval?
A It's a great question, because I think there are, like, two big problems. One is, like, how do you benchmark the general space in audio, where, like you say, it's, like, so dependent on the specific voice, let alone, like, if you are training into interactive, then it's, like, even more tricky. Um, and then the second piece, which is, as you are working on a specific use case, how you select a voice. So I'll take the second one first, which is, uh, we have, like, a voice sommelier, effectively, with us. We work with, with enterprises. We, we, we deploy that person to Work with them and help them navigate. That person is like a voice coach, has an incredible voice themselves, and, uh, and now we have, like, a team under that person that, like, will partner to help you find what's the right branding.
AI assessment note: “we have, like, a voice sommelier, effectively... We deploy that person to Work with them”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I asked the internet through X, uh, what questions we should ask you. Um, and a popular one was like, why aren't you guys building a law firm? Are you going to build a law firm and compete with all your customers?
A Yeah, no, we, we get this question. And I mean, I think when we first started Harvey and we were doing research, we actually talked to 30 people from Atrium. And I think also, interestingly, Sam and Uh, Jason, who was the GC of OpenAI at the time, and was the GC at Y Combinator when they did the Atrium investment. What struck us was the people who worked there said it was a really good idea, and they were super excited about the prospects, and then there were some challenges around the legal and the execution. But when we dug more into it, the big challenge that they ran into was, you're essentially just building two different companies, right? You're building a law firm, And you're building a tech company, and it's already really hard to, like, build product engineering, do AI, scale sales, and I think the big issue you run into if you try to do both of these is, I think you can only do one thing well, and doing a law firm well is very different than building a software company well. I think that's one point. The bigger point is, for us, it feels like the best outcome is if we can figure out How do we make every law firm? How do we help every law firm become an AI first law firm? Not how do we build one ourselves? And I think the real problem we're trying to solve is can we make every law firm more profitable? And a part of that is how they work with their clients. And can you…
AI assessment note: “you're essentially just building two different companies, right? You're building a law firm, And you're building a tech company”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And how do you think about the different parts of the coding market? There's vibe coding, there's professional code. Like they're all, there's all just one thing. Are these separable things?
A I think it's, Pretty separable. So for Warp, our target is pro developers building, like, software that's economically meaningful. So we, we really want to, like, focus on actually people who are using agents to build kind of hard apps, uh, apps that might, like, Go into your Mac doc or be pinned as a Chrome tab as opposed to vibe coded apps where I think it's more of a long tail play. And so I do think by the way, it's, it's awesome that anyone can code at this point, but I think if you look at where most of the value is in the software market, it's, it's not in those long tail apps. It's in like a relatively small number of apps that are super heavily used. And that's my background. Like I, you know, I've worked on one of those apps, Google sheets and just like, I have a lot of passion in terms of helping people build real apps. It's much harder, by the way. Like, I think it's relatively straightforward at this point for a good agent to like, you know, with relatively few prompts, build like a web app. It's much harder to apply these agents successfully to pro code bases. Um, so that's where we're focused.
AI assessment note: “I think it's, Pretty separable. So for Warp, our target is pro developers”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And as you think forward and forward in terms of all these different tools and all these different use cases and vibe coding versus professional coding and that the role of a software developer, Where do you think we are in two, three years?
A Yeah. So the, the way that I'm thinking of it is there's sort of three phases here. For most of my career, we were in like the world of develop by hand is how I talk about it. So my workflow then was like, I would open up a code editor. I would find files I want to change. I would type some code. Uh, I'd have some assistive features and then I would go back to the terminal and I would type commands to build that code. And I think we're switching away from that to something like develop by prompt where I start most of the coding tasks that I do right now by prompting an agent, and that agent does some work. And I think there's a third phase, which is like, automated development. Honestly, I think that's like, Kind of the bigger market here and why people are so excited about this space is you can actually use these agents to automate some parts of the software development process. And so that's like, you know, cognition does that. We're moving into this space cursor as background agents. Um, the rate at which this stuff will happen is like not super clear to me, actually. Like the, the most recent iterations of the models, in my opinion, were not as big of a step change. As like, for instance, when like Sonnet four came out, that was a really big step change in coding capability. I think there's going to be a mix of interactive and automated pieces of development for a while. I,…
AI assessment note: “within a couple of years, you'll have everyone working by prompt”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And why do you say that? Is it because you need to correct errors that the agents are making? Is it because things may be architected in a way that isn't scalable?
A Is it something totally, so it's like, it's like the agents, you can think of them kind of as junior engineers. So If you didn't have someone who was senior watching them, you end up in a situation where these agents will make code that creates bugs. It could create security issues. It can cause your code base to become really unmaintainable. And so there's actually like a premium right now on these senior engineer skills where you can architect, where you can review code, you can make sure the system doesn't degrade. And so again, I would be, if I were like early in my CS career, I would be racing towards building that expert Where you don't want to be, I think is like someone who is just like perpetually in the junior engineer state. Cause I do think that's at risk.
AI assessment note: “these agents will make code that creates bugs. It could create security issues.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Like what's lacking today, or what technology is most promising, or what do you think are the most interesting shifts happening? Yeah, I mean, talk in my book a little bit.
A I think batteries are the key. Batteries are the unlock to the energy transition, and will continue to drive out cost, uh, both the hard cost and the soft cost, right? I mean, you have the cells, and the modules, and the power electronics, and, uh, the bus bars, and the current collectors, and all the things that go into a battery, but then you have the EPC, you know, getting the battery in the ground, the logistics, the transportation, Uh, all the software that goes on top of it, um, monetizing these assets, and there's costs to attack in all parts of the stack. Um, and so I think that, um, battery storage will be, uh, will define the next chapter of the energy transition. And I'm, I'm super hopeful that new kinds of generation will break through as super economic. And I hope that breakthroughs happen in nuclear and, you know, there's some interesting companies working on geothermal and hydroelectric. And I think that has a lot of promise, but, you know, as we discussed, energy is a geographically defined problem. You know, parts of the Pacific Northwest and in Canada where there's tons of hydro, it makes a ton of sense to, to use that as a generation type. And, you know, in Texas, you don't have a lot of that, but you do have a bunch of wind and a bunch of solar. And so I think it'll look a little different in different parts of the world, but really, ah, you know, another wa…
AI assessment note: “I think batteries are the key. Batteries are the unlock to the energy transition”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So what do models need to, um, know about people? Or, like, what capabilities are they, um, either missing or have not been elicited from them?
A The most fundamental thing. Is that the models kind of don't understand the long-term implications of the things that they do and say. When you treat every turn of a conversation as kind of its own game, and you, you know, you basically think of it as like, okay, you had this interaction. You're done. You need to make sure that this one response has all of the possible answers, has all of the possible content. You don't ever, like, ask questions. You don't ever, like, try to clarify things. You don't really tend to express uncertainty. Um, you don't tend to be proactive. You don't tend to think about the long-term, uh, like, like, you see a lot of, like, even single-term side effects of this kind of regime. Like, and most of them are treated as kind of their own problems to solve. You see issues around, like, that, that people highlight around, like, sycophancy. You see issues that, you know, that there was recent news around, like, you know, the psychosis stuff. There, there's a lot of these, like, uh, harmful effects that you, Get if you think about things in this very single task or like task centric way. Um, but if you have models that actually consider, you know, the long term implications of, oh, hey, if I tell this person to start like a, you know, a company that, you know, sells gloves for catching ice cream. If I like tell them that that sounds like a good business ide…
AI assessment note: “models kind of don't understand the long-term implications of the things that they do”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Ah, you try one percent, one percent a day, right? Tom and Cass. But, but you, you know, now, how concerned are you about like deep fakes and generated spear phishing and, you know, voice attacks and all that stuff?
A So to the extent they enable the act of social engineering, yes, those are concerning because I think most forms of two-factor authentication are going to be, you know, out of the window. I still won't say which bank it is going to go and call it, but I call them and they say, oh, can you please confirm your identity? And they ask me three arcane questions, which I'm pretty sure chat GPT or Gemini will be able to answer in some seconds because they're, they're only scouring the web to find public information about me and ask me questions. So I think all those forms of Authenticating who you are, are getting an easier and easier to compromise. So the question becomes, so let's, so the problem you have to figure out is, you can solve it their way or our way, as in the way they're looking at it, the way we're looking at it. At the end of the day, every one of these social engineering attacks, credential takeovers, eventually initiates some bad activity in enterprise. And the bad activity in the enterprise often takes the form, takes on the form of what I will call anomalous behavior, right? Suddenly Sarah decided to exfiltrate all the data in Ila's company, even though she used to do email with him every day, today suddenly she's logged in and she's downloading everything onto her laptop.
AI assessment note: “to the extent they enable the act of social engineering, yes, those are concerning”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So you're basically providing these customer service, AI agents slash workflows that help I guess, function 24 seven and multiple different languages out of the box, and do you basically do like a lot of integrations into what they are already providing, or how do you tend to work with folks?
A Yeah, I think the way you should think about agents here are that, uh, it's more of a substitute for the mundane human labor. So whatever systems they're already using, generally an AI agent, at least when you first deploy, is not going to disrupt the tooling you currently have. So whatever CRM they're using, whatever, you know, telephony stack, We will just integrate with that. And then it's, it's kind of doing all the tasks you would expect a human to do. And, uh, over time that's just continued scaling. And so one of the benefits of AI agents is that they're always on, you know, either awake, 24 seven there. You don't have to train them really. There's no churn. You can just like scale them out.
AI assessment note: “whatever CRM they're using, whatever, you know, telephony stack, We will just integrate with that”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What made you realize that this was the thing to do?
A The, the real answer is we just saw a lot of folks that were willing to pay us like, you know, six figure contracts, which at the time when you're at zero ARR, it's like, oh wow, that's huge. And a lot of folks that were willing to, you know, do the same thing. And it was the only idea we really explored that really had that property where people were like, Hey, yeah, like if you did this, I would literally pay you money because I can justify it. The sort of flip side of that at the time was more just, oh, well, this is such an obvious idea. Like why do this? Cause you know, people would have thought of this stuff before, but Uh, that, that's a whole nother thing. I think like once you start doing anything, once you get into it, you, you, you understand there's way more nuance than the overall narratives. The sheer fact that people are willing to talk to us, like, you know, two people and willing to pay us money was signal enough that it was worth doing.
AI assessment note: “we just saw a lot of folks that were willing to pay us like, you know, six figure contracts”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What makes somebody a good customer, like the right type of early customer for you?
A I would start with, uh, do they have a complex problem? Um, that we think there's a lot of efficiency to be created by applying AI. Uh, number two is, do they have the right data, um, that we can use in order to create the solution? Uh, number three is, do they have the right command and control structure in the company where people are willing to buy into it? Because you could have one or two, but if you have a lot of resistance, you know, in the ranks or in leadership and they're not willing to kind of force people to try, um, There can be a lot of resistance because people are afraid of change. You know, change is like heaven. Everyone wants to go there, but nobody wants to die. And what we found is that some of the people who were are most, um, are most resistant initially, once they see what this can do for them and how much better it's making them at their job, they're actually becoming our biggest champions. And so that's been very gratifying to see in some of the customers, um, so far.
AI assessment note: “I would start with, uh, do they have a complex problem?”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q maybe it's by the label or something else, and there's a bunch of stuff that's a bit more TBD in terms of those kind of trials that may connect each other a little bit, or maybe other information that may be a bit more sporadic. How do you, how do you deal with that third bucket of ambiguity, and how do you think and tell about capturing that broader knowledge?
A So the first way to deal with that third bucket of ambiguity is ensure that your users are physicians and not patients. And we've made that strategic decision. And we keep thinking we're going to change that decision. And, uh, we've been talking about changing that decision since the inception of the company and so far have not changed that decision for all the reasons implicit in your question. There's an enormous luxury that we have as builders, um, in having doctors as users because the MD is attached to their name. Right. So they need to protect that MD and they're going to use us as a tool in the same way as a Wall Street trader might use a Bloomberg terminal. If a Bloomberg terminal, for example, produced, you know, an inaccurate quote on a bond that was very obviously inaccurate, you know, it was off by an order of magnitude and the, and the trader, you know, in a hedge fund just sort of, well, I mean, that's odd.
AI assessment note: “ensure that your users are physicians and not patients”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you actively seek to find more motivation for yourself?
A No. Uh, I, I, I, and, and actually the opposite. One of the things I think is unhelpful about the contemporary cult of psychoanalysis and in psychology and psychiatry that sort of traces its origins to early 20th century and Freud and these guys is, um, it doesn't appreciate that in the, Analysis and description of something, you kill it. So I've actually resisted exploring trauma. I've resisted going back to the origins of my motivation. I've resisted going back to the origins of my aggression. I have kind of like a partially developed map from childhood and other experiences, but the second I feel myself going close to analyzing it, I, I, I resist the urge to analyze it because in, in the, in the analysis of something, um, Is, is the deletion of it, in a way.
AI assessment note: “No. Uh, I, I, I, and, and actually the opposite.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q One last question for you on just, like, what other parts of the economy you focus on. So there's energy, there's obviously intelligence, um, there's sort of inputs and, um, rare earth, uh, magnets and, uh, minerals and, like, what other, Are there other domains where you think competitiveness is like essential from a security perspective or from a strategic perspective?
A Yeah, so I tend to think of, of my work being very supply chain focused is because, um, it allows, it gives me a mental framework for thinking about these issues holistically, um, by looking at the supply chain as a layered pyramid that, uh, includes energy, minerals, uh, component manufacturing, semiconductor manufacturing, data centers, models, and apps. And, and I think, you know, you, you, As a country, we kind of need a strategy that is holistic at the different layers of the supply chain. Um, we're actually in a really good position at most of them. We have abundant energy, although we need to increase our supply. Um, we, we're, our, our biggest exposure points are component manufacturing, semiconductor manufacturing, and minerals, and, you know, there's a lot that we can do to move the needle there. Um, but, but transportation logistics is another Really interesting, um, really interesting area where, um, a policy can actually help play a role, and the Chinese have been masters at, you know, through their Belt and Road Initiative at actually having a supply chain plan that includes a global transportation logistic network to get minerals from Africa back to China, refined in China, exported back everywhere else, And, and I think we need, as you know, we used to do this with the Panama Canal with, you know, these big investments that we used to make in transportation logi…
AI assessment note: “transportation logistics is another Really interesting, um, really interesting area”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q There's certain things that were missing initially that are now in place in terms of Everything from certain forms of inference time compute on through to forms of memory and other things that allow you to maintain some sort of state against what you're doing. What do you view are the things that are still missing, or need to get built, or what sort of foment progress on that end?
A I think the technology component level, there's stuff that I hope will improve. For example, computer use, you know, kind of works, often doesn't work. Um, I think, so the god rails, evals is a huge problem. How do you quickly evaluate these things and drive evals? So I think the, the component is this room for improvement. But what I see is the single biggest Barrier to getting more, uh, agentic AI workflows implemented is, is actually talent. Uh, so when I look at the way many teams build agents, the single biggest differentiator that I see in the market is, does the team know how to drive a systematic error analysis process with evals? So you're building the agents by analyzing at any moment in time, what's working, what's not working, what do you improve? As opposed to, uh, less experienced teams kind of try things in a more random way, then it just takes a long time. And we're looking for a huge range of businesses, small and large. It feels like there's so much work that can be automated through agentic workflows, but, you know, the talent, the skills, and maybe the software tooling, I don't know, just isn't there to drive that disciplined engineering process to get this stuff built.
AI assessment note: “the single biggest Barrier to getting more, uh, agentic AI workflows implemented is, is actually talent.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What is the, um, if you just look at the spectrum of agentic AI, what's the strongest example of agency you've seen?
A I feel like Bleeding edge of agentic AI. I've been really impressed by some of the AI coding agents. Um, so I think in terms of economic value, I feel like there are two very clear, very apparent buckets. One is answering people's questions, uh, probably, you know, open AI, chat, GPT seems to mark the leader of that, with real takeoff, lift off velocity. The second massive bucket of economic value is, uh, coding agents, where coding agents, like my, my, my personal favorite, Claude Deva Tu, Right now it's cloud code. Maybe it'll change at some point, but I, I, I just use it. Love it. Uh, highly autonomous in terms of planning out, you know, what to do to build a software, building a checklist, going through it one at a time. So this ability to plan a multi-step thing, execute the multiple steps of a plan, uh, is one of the most highly autonomous agents out there being used that, that actually works. Uh, there's other stuff that I think doesn't work, like, Some of the computer use stuff, like, you know, go shop for something for me and browse online. Some of those things are really nice demos, but, but not yet production.
AI assessment note: “I've been really impressed by some of the AI coding agents.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What do you think investing firms Or incubation studios like yours will not do two years from now? Like, not do manually. Sorry.
A I think there's a lot could be automated, but the question is, what are the tasks we should be automating? So, for example, you know, we don't make follow-on decisions that often, right, because of portfolio of some dozens of companies, so do we need to fully automate that? Probably not, because we're very, very hard to automate. Um, I feel like doing Deep research on individual companies and competitive research, that seems right for automation. Uh, so I, I might, I don't know, I, I personally use whatever Open Eyes Deep Researcher and other Deep Researcher types of tools a lot, uh, to just do at least a cursory market research things. Um, LP reporting, that is a massive amount of paperwork that maybe you could simplify.
AI assessment note: “doing Deep research on individual companies and competitive research, that seems right for automation.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Amazing. And I, I don't think there's a way to ask this question without, like, somewhat trivializing the journey, but, like, how did you become so dominant?
A I don't know. I, I mean, I think we just focused on how did we do the right thing for our customers? How do we solve the problems that were there? And, you know, at some level the story of Cloudflare is that Um, our, we have been customer zero along the entire journey. So every, you know, thing that started from, could we take a firewall, put it in the cloud? How would we get the data to populate that firewall? We had to have a free service. Once we had a free service, all of a sudden we had to be able to figure out how to scale, um, you know, enormously across, you know, millions of customers, uh, in, in an efficient way. That meant that we had a whole bunch of, you know, weird stuff that was using us. We got attacked by Every which direction. We had to build a public policy team in order to deal with those issues. We had to build our own security. We, like, someone almost hacked or stole our domain at some point as a way of hacking into us. So the next thing you know, we built our own registrar. And so to some extent, Cloudflare has been about, you know, start with a relatively simple idea, um, make it as broadly available as possible, and then solve all the problems that become sort of inherent once you've done that.
AI assessment note: “we have been customer zero along the entire journey”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q at least expert content that you mentioned, um, if you look at some of the models like MedPalm II from Google, which is a couple years old now, it outperformed human physicians, the average human physician in terms of output. So if you rated its output against people, at what point do you think we've run out of good content from people? In other words, there, there is some limit
A I don't think that's true. I mean, I do think that there, there will be some, like there's always going to be people running new experiments and new tests and finding new things and, and new discoveries. And yeah, maybe we can imagine some distant future where it's all robots that are, that are doing this in, in, in the labs, but that's, that's a long ways off. And so in, in the meantime, you know, I think we can do that. My, my black mirror kind of version of the future though, Is, is actually one where we're not going to get rid of journalists. We're not going to get rid of scientists. We're not going to get rid of researchers. You're going to still need that work. What I worry about is we don't figure out how to compensate broadly content creators who are independent, that we actually go back to almost a time in the Medici's where the web had historically been this incredible sort of, um, distributor of value of creation and, and knowledge creation. You could imagine a world in which All of a sudden you have five big AI companies. You have the conservative one, and you have the liberal one, you have the European one, and the Chinese one, and, and they all actually hire and run their own team of journalists, researchers, academics, the experts that fill in the cheat, the sorts of holes in their cheese. And, and again, that's not too hard to imagine that in some not so distant…
AI assessment note: “I don't think that's true... there's always going to be people running new experiments”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Does that imply any particular belief around like open or closed models as, you know, people continue to develop capabilities?
A You know, we have closed models that run on us. Um, We have a lot more open models that run on us. Um, I, I tend to, um, we, we have historically been a, a company that believes very much in open source. And we, most of the things that we build internally, um, as long as, as long as we can, we try to open source all of that technology. And so I, I tend to be in the pro open models. Um, you know, we work very closely with the meta team and llama and, And everything that they're, they're doing. But again, I think there's going to be different, different flavors of this. And, and I, you know, again, I, I'm, we're happy to have customers in either end of, of that, that spectrum. I'm, I'm a little bit skeptical. I'm actually quite skeptical of the, if we allow open source models, the world is going to end arguments. That seems histrionic to me. There are things we should worry about. Like it is, you know, I, I think some of the, um, you know, uh, synthetic pathogens and other things that can be created, but It seems to me like the, the place to regulate that and control that is in the machine that can actually print the pathogens, not in the AI model that can come up with, with, uh, with, with what it is. That, that seems like a, a pretty flimsy argument for why we shouldn't have open source.
AI assessment note: “I tend to be in the pro open models.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And they can't, it's not easy to predict what is incremental to models. What advice would you have for them?
A So I think the first thing is you've got to, you've, you've got to get back to controlling your content. So you have to create scarcity from the beginning. So how do you make sure that you're not just giving your content away for free? And again, we've made that easy. There are other companies, uh, that are working to try and make that easy, um, as well. And so one way or another creates scarcity and then start to have conversations. You can see which AI companies are the most likely, uh, to, to deal with it. So we're, you know, just today there was news that Google is starting Uh, pilot project to start to pay news providers, something they swore they would never do. Um, but again, I think that they can see, and because they do believe in the ecosystem, they can see that this has to happen. If the incentives for creating content go away, if you can no longer sell something, if you can no longer sell ads against something, if you can no longer even get the ego hit, because if people aren't going to the original source, you don't even know. If you write some incredibly influential piece that ends up in, you know, millions of AI responses, You don't actually ever even know that happened. You're, you're yelling into the void. We've got to figure something out around that, uh, around that piece. So I think the first step for content creators is recognize that you have a, that the b…
AI assessment note: “you've got to get back to controlling your content. So you have to create scarcity”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What does it mean to win the AI race? Like, why, why do we need to win that? And what, like, what would losing mean? How do we know we've won?
A Well, I'm, I'm suspect you folks agree. AI might be the most transformational economic cultural force of our lifetime. And I believe that if the country or the ecosystem Which winds up getting ahead is going to have these cyclic effects, right? Like you're going to, you know, you're going to power productivity. You're going to have drug discovery. You're going to, you know, discover, uh, new material sciences, new technologies, which then feed back into your infrastructure, feed back into your economy. And you're going to get this flywheel effect where who winds up getting ahead could wind up really accelerating ahead in kind of a classic network effect ecosystem way that all of us in Silicon Valley will understand. Now that is purely on the civilian economic Context. You can also imagine a military context, right? Think about everything from drones to autonomous weapons. I'm pretty sure it's not in our best interest to have another country have that same economy of scale and flywheel and race ahead of us. So that's the race. Now, one interesting question that we have been pondering, which we can get it to, is how do you actually measure what it means? How are we doing in the race? And one measure I've been playing around with, and maybe I want to get your take on, is I think Google just announced this morning, uh, that they inference one quadrillion tokens a month or a quarter…
AI assessment note: “What share of those tokens are being inferenced on American hardware on American models”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q These, you know, course, uh, as you said, 5:02 reaction, academic benchmarks, or even non-academic industrial benchmarks are, uh, easily hacked or like not the right gauge of performance against any given task. They are very popular. What is the alternative, um, for somebody who's trying to, like, choose the right model or understand model capability?
A So the alternative that I think all the Frontier Labs view as a gold standard is basically human evaluation. So again, proper human evaluation where you're actually taking the time to look at the response. You're going to fact check it. You're going to see whether or not it followed all the instructions. You have good pace so you know whether or not the model has good writing quality. Like this concept of, like, doing all that and spending all the time to do that as opposed to just vibing for five seconds. I think actually is really, really important because if you don't do this, you basically, you're basically just training your models on the analog of clickbait. Um, so I, I think it actually is really, really important for model progress.
AI assessment note: “the alternative that I think all the Frontier Labs view as a gold standard is basically human evaluation”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Okay. Yeah. I mean, so Sarah brought up earlier, um, how everybody kind of wants high quality data. What does that mean? How do you think about that? How do you generate it? Can you tell us a bit more about your thoughts on that?
A So let's, let's say you wanted to train a model to write an eight line poem about the moon. And so the way most companies think about it is, well, let's just hire a bunch of people from Craigslist or through some recruiting agency and let's ask them to write poems. And then the way they think about quality is, well, is this a poem? Is it eight lines? Does it contain the word moon? If so, like, okay, yeah, I hit these three checkboxes. So yeah, sure. This is a great poem because it follows all these instructions. But if you think about it, like the reality is you'd get these terrible poems, like sure it's eight lines and it has the word moon. But they feel like they're written by kids from high school. And so other companies be like, okay, sure. These people on Craigslist don't have any poetry experience. So I'm going to do instead is hire a bunch of people with PhDs in English literature. But this is also terrible. Like a lot of PhDs, they are actually not good writers or poets. Like if you think, like think of Hemingway or Emily Dickinson, they definitely didn't have a PhD. I don't think they even completed college. And like, one of the things I will say is like, yeah, I, I, I went to MIT. I think Eli, you went, you went there too. And a lot of people I knew from MIT who graduated with a CS degree, they're terrible coders. And so we think about quality completely differently. …
AI assessment note: “we think about quality completely differently. Like what we want isn't poetry that checks”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q ways. So I guess it depends on what you mean by external validation. Like in my mind, again, like I often think about things from a perspective of, are you trying to build a startup that's actually going to change the world? Like, do you have this big thing that you're dreaming of? And if you have this big thing that you're dreaming of, you, Like, why do you care?
A Maybe the way to think about it is in Sarah's context, like if you haven't, say you're a YC founder, you haven't been at Google, you haven't been at Meta, you haven't been at Twitter, you don't have this network of engineers, you're a complete unknown, you haven't worked with very many people, you're straight out of school. How do you then attract that talent? And to your point, you can tell a story of how you're going to build things or what you're going to do, but it is a harder, um, obstacle to basically convince others to join you or for others to come on board or to have money to pay them if you haven't. If you don't have long work at history. So I think maybe that's the point Sarah's making.
AI assessment note: “it is a harder obstacle to basically convince others to join you”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What is an example as you guys open up the waitlist that you want users to try where it should just be like obvious that the answers are, are better than other coding agents?
A I think the kinds of, um, queries that it tends to be better at are I guess what we would call Semantic queries. So let's say like an example of a query where this is not the best system to use. It's like file level. If you're looking at a file, and there's like a specific thing in that file, and you're just trying to get a quick answer to it, you don't really need the hammer of like a deep research like experience. Um, you don't need to wait, you know, like tens of seconds or a minute or two, uh, to, to get that answer, because that should just be delivered snappily. But if you, um, Don't exactly know where you're looking for, and you, you know, you don't know the function name, or you don't, you know, something, and this is kind of the hard problems that engineers are usually in. Like, there's a flaky test. I mean, you know that this test is flaky, but that's where your knowledge stops, right? And that's when you usually go to Slack and ask some engineers, like, this test is flaky. What's going on? Does anyone know? Um, you know, we've had, uh, the way we've used it is when you're training these models, there's a lot of infrastructure work that goes into it. And, um, it fails in interesting ways all the time. Uh, and asking things like, you know, my jobs are running slowly, five times more slowly than usually. Why is that? That's kind of a vague query that would be very hard …
AI assessment note: “asking things like, you know, my jobs are running slowly... Why is that?”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q You have to have a concept of authority, right? People are going to say things that are wrong.
A The way it's worked with customers we've started working with is, uh, they typically have, they, they want to start off with kind of a group of trusted kind of senior, you know, staff level plus engineers who are kind of the gatekeepers, which is a very, I think, common notion. Um, you have permissions, right? And ownerships, uh, ownership structure and code bases. And they basically are the ones who kind of populate the memory first. Um, and then sort of expand the scope, but I think it works. It's actually a much more complex feature to build, uh, because it touches on, um, yeah, org wide permissions. Um, there's some parts of the code where a certain engineer should be able to edit the memory, but other engineers shouldn't. Um, and so it, it actually starts looking like the new way of, um, versioning code effectively, right? It's kind of a GitHub plus plus, uh, because you're not versioning the code, you're kind of versioning the meta knowledge around it. That helps language models understand it better. Uh, but definitely that is something that we built, but I think it's a thing to iterate a lot until you kind of get the right design here because you're effectively building and yeah, and you, and you get from scratch.
AI assessment note: “staff level plus engineers who are kind of the gatekeepers”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Okay, last question, Misha, where would you characterize us as, like, being on the path toward deployment of these capabilities in, in different fields?
A I think we're a lot earlier than most people think, uh, that this is going to be one of those areas where the technological building blocks, um, outpace their deployment. And so, yeah, within the next couple of years, uh, the blueprint roughly for, you know, how to build ASIs will have been set more or less. Like, uh, maybe there's still some, um, efficiency Breakthroughs that need to happen, um, but more or less there'll be a blueprint for how do you build a super intelligence in a particular category? Actually going in and deploying it and, and building it for, you know, specific categories of work. There are going to be a lot of product and kind of research innovation specific to those categories, um, that will probably make this a multi-decade thing. Um, so I don't think that it's a couple of years from now and, uh, GDP starts growing 10%, um, you know, year over year globally. I think We're actually going to get there, uh, but it's going to be a, kind of, multi-decade, uh, endeavor. I tend to, kind of, see a lot of patterns, uh, now in, kind of, real-world deployment with, uh, reinforcement learning, um, research as it worked, again, before large language models. Um, and before large language models, it used to be, kind of, you pick an environment, like, you pick Go, you pick, um, StarCraft, you pick something else, and You go and try to solve it with, you know, some combi…
AI assessment note: “I think we're a lot earlier than most people think”