Every argument clarity score on this site is built from rows on this page, here across
all 44 shows. Each
question and answer was assessed with names hidden, the hosts' own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Um, and... Can I ask you, we, we, we, we saw Benioff say that he spends three hundred million a year on Anthropic, which equates to about 3.8% of developer salaries on Anthropic. To make it justify the valuations that we're seeing for these companies, it needs to be 20%. Do you have any concern in that movement from 3.8% to 20%?
A No. I mean, I, I think if you look at I've never done this in any detail, but if you look at what we pay hardware engineers, and you look at what the tools, which we, the EDA tools they use, I bet you're much closer to 15 or 20% than two or three percent. What's happened is historically software engineers use very low-cost tools, and hardware engineers used extremely expensive EDA tools. And so, that's interesting, isn't it? I mean, I, I, I think we, the cost of bugs And hardware is so high that we became accustomed to using many expensive tools. And in software, we threw people at the problem rather than tools. And as AI becomes more productive, I certainly don't see a problem where software engineers using 50 or a 100,000 a year each in tokens. There are forty-seven million software engineers in the world. I mean, that's five trillion dollars just in software engineering token use.
AI assessment note: “No. I mean, I, I think if you look at what we pay hardware engineers”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Again, going back to my idea of what the future looks like, what should one expect from that? Does it ease? Does it ease over time? What happens to the cost?
A Well, I, I think the, the challenge here is that, that these are extremely, uh, lumpy. Items, right? You can't just add a little bit of manufacturing capacity at a fab. You have to build a fab for forty billion dollars, and it takes five years to build. So if you see demand explode, you cannot respond quickly. All you can do is fill your factory. Once your factory is filled, you've got to build another factory, right? It's a step function in your ability to meet that demand. And the step is huge and takes years, and so if demand stays high, uh, they are, we're going to continue to see memory shortages for at least the next several years.
AI assessment note: “we're going to continue to see memory shortages for at least the next several years”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q in that this market would be really important, and you more than others, right, since you actually started a company in it. But then it took some time for the market to really expand to the point where, uh, to your point now, it's, it's this massive use case. People really care about speed of inference and other things. Um, what gave you the conviction back then to do this?
A Combination of, of vision, um, the right co-founders, and a little bit of arrogance, a little bit of luck. You know, we, we saw AI on the horizon as a new workload. And as computer architects, new workloads are opportunity, right? It's very, very hard to, to, to enter in the x-eighty-six world, right? Where there's not, nothing new is happening there, and nothing has happened for generations. But You know, when graphics emerged, you got the discrete GPU and you, you, you got, uh, Nvidia and, and when, uh, when the mobile, uh, compute hit, you, you got arm. And it was interesting that, that not Intel, not AMD, not all sorts of people who you would have thought have been really well positioned to win in that business. They all got no share. And so we knew that, that this new workload would eat a lot of compute. It would require Uh, a new architecture, dedicated architecture, and that ought to be very different. The architecture could not be a derivative of what's existing. Those were our big bets, and they were a hundred percent contrarian, and they turned out to be dead right.
AI assessment note: “Combination of, of vision, um, the right co-founders, and a little bit of arrogance”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q terms of the different models there. Again, sorry to cite it, but it's, it's kind of handy having just done it. Jonathan said that you would definitely have OpenAI and Anthropic build out their own ships, because then they would have control of their own destiny. Do you think OpenAI and Anthropic build their own ships so they don't have self-reliance on NVIDIA in the way that they do today?
A I think that there is a long history of software companies failing to build chips. The list is, is very large. I think, uh, whether, uh, OpenAI can do it, uh, whether they can do it through partnership with other vendors, with Broadcom, with smaller, more innovative companies is an open question. Um, but I, I think that, uh, you know, companies at the size of Microsoft have been unable to deliver, uh, chips, right? I think, uh, uh, there are plenty of examples as you look across the Fang, uh, group where chips were tried. I mean, probably the most successful is Google and they're 10 years in, right? Maybe longer. You know, soft, modern software does not fit well In a chip making framework. I mean, weekly sprints don't work well on two year long projects. Um, you know, move fast, break things often is not the way you think in the chip world. The way you think in the chip world is measure twice before you cut once because your bugs cost you six months and tens of millions of dollars. And so it's a very different mentality. And where there's been success, it has frequently been acquired. Apple got into the chip business through buying PA Semi. Amazon got into the chip business through acquiring Annapurna. Um, uh, Google acquired the talent from a collection of companies, uh, and then set it in a BU that was a side and under somebody who, who had sort of enormous respect in the org…
AI assessment note: “there is a long history of software companies failing to build chips.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q are putting a lot of money into compute and hardware and compute and hardware startups. And I wanted to ask you, should I be following them? What should I be looking for in them? I'm seeing some incredibly young founders. I was with a nineteen-year-old founder trying to take on NVIDIA this morning, um, raising twenty million dollars for a pre-seed. How, how should I be thinking about this, Andrew?
A Well, I, if you don't know a lot about hardware, I wouldn't invest in hardware. I think, uh, Harry, it's probably the same in many things, is that, um, I think hardware is not an easy place to make money. It's a place that has historically rewarded experience, both from investors and from entrepreneurs. I think the number of different technologies involved in designing a chip, Uh, is extraordinary. Not just the, the logic, which is what most people think about. When, when you think about chip design, that's just the front end part that that that's writing in very low level software. Um, but the, the selection of tools, right? We pay millions of dollars a year in tools, selection of geometry, right? Which fab and having a relationship with a fab, you're going to pay 20 or thirty million in NRE. And if you have a bug in your, in your chip, you got paid again.
AI assessment note: “if you don't know a lot about hardware, I wouldn't invest in hardware.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q so I want to do a bit of a deep dive on that on itself in a, in a second. Um, uh, but to close on this, so three bottlenecks, um, you also mentioned CPUs a couple of times in this conversation, and there seems to be a theme around the emergence of like a CPU shortage as well. Is that, so what is that true? Two, what causes it?
A So agentic AI Is a world in which AI doesn't just provide answers. It initiates action. So that action might be go to a website. It might be learn, gather some data from a website, bring it back, take another action. Those actions are done by CPEs. And so as AI gets better and better at doing things, at making instructions, calls for things to get done, we're using more and more CPUs, right? And that is driving up the consumption of CPUs and therefore the demand for CPUs. And so this sort of huge push for more CPUs is being driven by AI On machines like ours and GPUs doing agentic work and asking the, the, the CPUs to take an action, to go to a website, to order a burrito, to find a piece of information, to pull it from storage to all that work is being done by the CPUs.
AI assessment note: “huge push for more CPUs is being driven by AI”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Great. Let's talk about that over in that deal since it's such like a major historical milestone record making. Um, so it's, it's providing up to with seven and 50 megawatts, which is, which is interesting by the way, as a, as a metric because we're in a chip provider, but this is power. So is that shorthand for?
A It's a shorthand. I mean, it turns out right now, and we didn't talk about this because there's Sort of in the adjacent supply chain. We, we went through the shortage of, of memories. We went through the shortage of a process called COOS, three nanometer capacity. The, the other limitation in our industry right now is data center availability. And, I mean, Uh, that is a limiting factor for everybody. And that's why Anthropic did a huge and sort of very expensive deal with Elon for data center capacity. Um, uh, our deal with them, with, uh, with OpenAI was because data center capacity is a limiting constraint, measured the way data centers are measured in, in megawatts. The deal is 760 megawatts, 250 megawatts in 26, uh, on a multi-year lease. An additional 250 megawatts in 27, on a multi-year lease, and an additional in 28, a multi-year lease.
AI assessment note: “It's a shorthand. I mean, it turns out right now”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Co-design has become an important part of how do you create that kind of platform for the next three years because you're, you're building for the next three years. So how do you think about that and how do you go about chip design?
A Well, I, I think historically you, you made chips and you ran software on them and, uh, There wasn't surprisingly a close interaction because there was a layer called an operating system that lived between the chip and the software. And so, uh, you know, Intel and AMD made chips and people wrote software for the operating system or for the chip. AI has gotten so large and, uh, speed is so important that, that what they're doing is they're thinking about sort of the design together. That what changes could we make in software that would advantage the hardware or as we're designing the hardware, what changes could we make that would make the software easier to run? And so they're being designed sort of at the same time. And like anything, when you sort of design things together, um, the, the advantages are enormous. And, and so this is something that's really taken shape right now. Um, you know, one of the advantages of our relationship with open AI is we, we get to see exactly where the frontier is going and we get a chance to, to roll that into our designs. One of the advantages Google has is that their TPU can be designed in collaboration with the team building Gemini or the team of DeepMind, and so they can, they can inform their choices back and forth, and that's an enormously powerful thing that is surprisingly relatively new in our space.
AI assessment note: “they can inform their choices back and forth, and that's an enormously powerful thing”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q We've spoken before and off record about, kind of, personal lives. I'm intrigued. When you are a public company CEO and you're going public, the world wants a piece of you. You're public, now you're public. Any advice on how to sustain an amazing marriage And an amazing relationship while also being a public company CEO and going through that process?
A I, I would say that pick a wife with patience. Pick a partner, a husband or a wife, partner, um, who understands w w what it is to, to be an entrepreneur. I, I don't think, and I look at my co-founders and, and our leaders, it, it is every day when you're a leader of a, of a startup, a pressure test on your soul. Every single day. And if you're a real leader, when you are 30 people, a little, a little company picnic, you look out, and what you see are mortgage payments and braces that need to be done that you're responsible for. And that doesn't change. And, uh, I, I think that if you really believe that, and you, you hold that in your heart every day, you carry real weight with you. And I, I think, uh, you have to share that with your partner so they understand. It's really hard if they don't. I think almost everybody, and maybe your, your, your, your partner has felt this, and I think every CEO I know has told the story of their partner telling them that They're more lonely when you're sitting next to them thinking about work and your mind is just ripping on work than they were when you weren't in the house. And I, I, I think that what we do is a family thing. There's a price to be paid and how often you see your wife. I mean, I'm on the road three weeks a month. I mean, put it this way, uh, Emirates airline sends me a Christmas basket. This is an Arab airline sending a Jewis…
AI assessment note: “pick a wife with patience. Pick a partner, a husband or a wife”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Do you think we should be selling chips to China as a result?
A No. No. I, I think, um, I think, uh, let, let's remove all of us that are self-interested, and even though I'm arguing against my self-interest, right, um, if you remove me, and you remove Jensen, you remove Lisa, and you remove everybody in the chip industry, and, and you say, if we sell to, to somebody in the security business, and you ask this question, if we sell leading edge technology to China, will their military use it? Everybody says yes. There is no debate on that point. Their military will use it. You ask a second question, which is, if you sell our leading edge technology, will they, will their government use it through their industry to compete with us, all right, in an advantaged way? The answer is also yes. In that, and so that's where I stop. There's complete agreement that those two things are true by everybody in the security business and outside of the chip business. And now, you can say that keeping them in our ecosystem is the best way to manage that problem. That's one argument, and there's some merit to that. There is keeping them from building their own Their own ecosystem is something that's in our interest. There's real merit in that. I don't agree with either of those arguments, but, but they're real arguments and they have real merit. They are at least today our industrial adversary. And, uh, as you travel the world and you see sort of the results of…
AI assessment note: “No. No. I, I think, um, I think, uh, let, let's remove”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q If I said that you have one policy change that you could usher through with no resistance, what would it be?
A Um, I, I would allow TSMC and, uh, Samsung, uh, both to, uh, a 20 year period free from all local and legal, local ordinances, all of them, to build fabs in their desired location in the US. If that's Arizona, that's great. If that's Texas, that's great. 20 years, no local rules, no, allow them to build fabs And I, I would say that we use the same rules we use in Taiwan. Don't, don't build garbage. Use exactly the same construction techniques and rules, et cetera, that you, that you've built fabs successfully elsewhere in the world. But local ordinances are disastrous. And not intended to cover pyramids, right? Fabs are modern pyramids, Harry. I mean, they are the greatest things humans make in manufacturing, in the manufacturing world, by far. Nothing's closed.
AI assessment note: “allow TSMC and, uh, Samsung, uh, both to, uh, a 20 year period free”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Can I ask you as an aside, actually, just because you, you have for more than a decade believe that this revolution is going to happen. Uh, how much is all of this, um, AI generated coding relevant for Cerebris internally?
A Hugely. I would say that, that, you know, eight months ago, we weren't spending a thousand dollars in engineer on tokens, and we're probably at 25 or 30,000 right now, and it's ripping. I, I think it's not useful for everybody. I, I think that's the truth. I, I think there are some, some people who have sort of the perfect mindset for it, right? And you, you, they are running eight or 10 agents, seven by 24. They've moved their coding, Style to being one in which they govern agents, whether they think about how to QA, so they've got a QA agent running, they think about how to sort of remedy some of the weaknesses in the coding models, right, they're often verbose, they often cut out comments, so they've really thought about, and it's a type of puzzle that's the perfect fit for their mind, and they've gone from being sort of 10 X guys to being hundred X guys. I think the rest of us, myself included, we're sort of Limping along. We, we, we're trying to figure out how, how we can make it work for, for our different jobs, for being the CEO, for being the CFO, for being accountants, for being in marketing. Um, but for a, a small number, it is such a tool. And then the rest, we try and, try and show them what, what, what others are doing, what best practices are.
AI assessment note: “Hugely. I would say that, that, you know, eight months ago, we weren't spending”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Dude, you are too kind. Uh, listen, I want to start with the billion dollar raise that you just announced yesterday. Um, can you just talk to me about the billion dollar raise, why it's important, why now, and what it means for the company?
A Well, look, it was the largest raise ever done in, in our category. Uh, it was done at the highest valuation and with the, the premier investors. So, uh, At, at late stage investing, you're looking for, uh, the likes of Fidelity. They are the, uh, What would the English call it? The sort of Oxford or Cambridge of investing, right? I mean, they, they are the, uh, the premier, uh, public market investors, and when they choose to lead a round, uh, it, it brings the Wall Street a great deal of confidence, and so, uh, we were really happy to partner with them and with the treaties to lead the round, and then we, uh, were able to get enormous participation from Tiger Global, From Valor, from 17, uh, 89. So that's .1. I think .2 is that, um, We, we now have sort of the dry powder to, uh, to really push and to take the opportunities in front of us, uh, to build out our manufacturing to the scale and scope we want, to add new data centers. We added five this year in the US to add more data centers, and we have more big ideas. I, I think incremental improvements, Uh, uh, make believe gains achieved by dropping from, from, you know, eight bit to four bit. Uh, those aren't gonna get us to, uh, to the promised land in AI. We, we need, we've got real work to do as a community. And I, I think this, this funding puts us in the catbird seat for that.
AI assessment note: “we now have sort of the dry powder to, uh, to really push”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q but SRAM is obviously memory, like, on chip versus off chip. Seemingly great, but he said it's completely unable to handle scale, and so although it may be quicker, for anyone who wants to do large scale, It is incapable at present of doing that, and that's a fundamental need and requirement of any of the large providers. Do you think that's fair, and how do you think about that?
A Well, not only is it fair, it's the reason we went to wafer scale. So let me explain. What your friend said is strictly true, in that SRAM is blazing fast and low capacity. HBM is a flavor of DRAM. It has high capacity, and it's very slow. Now, NVIDIA and all GPUs, including, uh, uh, AMDs chose a big capacity memory that is slow because it's perfect for graphics. You don't have to go to memory very often. You can hold a lot, you don't go very often. SRAM Is blazing fast, but it can't hold very much. So the problem on traditional chips is if you put memory on the chip, you are using space that could be otherwise used for compute. You have a fixed amount of real estate. And so if you put half memory, then you have half your real estate's available for compute. And so, our idea was that if we, we sort of, if we built a chip that was the size of a dinner plate, we could stuff it to the gills with fast SRAM, overcoming the limitation of SRAM, which is it doesn't store very much, by putting a huge amount down, by using more silicon area. Now, if you're an SRAM solution today in a normal size chip, And you're trying to do a trillion parameter model, use four or 5000 chips. What a mess. You know how many cables that is? Do you know the, the, the impact to the AI? It's a horrible mess, right? And it limits you from doing things you want to do with the AI, uh, like speculative decode. It…
AI assessment note: “Well, not only is it fair, it's the reason we went to wafer scale.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q but SRAM is obviously memory, like, on chip versus off chip. Seemingly great, but he said it's completely unable to handle scale, and so although it may be quicker, for anyone who wants to do large scale, It is incapable at present of doing that, and that's a fundamental need and requirement of any of the large providers. Do you think that's fair, and how do you think about that?
A Well, not only is it fair, it's the reason we went to wafer scale. So let me explain. What your friend said is strictly true, in that SRAM is blazing fast and low capacity. HBM is a flavor of DRAM. It has high capacity, and it's very slow. Now, NVIDIA and all GPUs, including, uh, uh, AMDs chose a big capacity memory that is slow because it's perfect for graphics. You don't have to go to memory very often. You can hold a lot, you don't go very often. SRAM Is blazing fast, but it can't hold very much. So the problem on traditional chips is if you put memory on the chip, you are using space that could be otherwise used for compute. You have a fixed amount of real estate. And so if you put half memory, then you have half your real estate's available for compute. And so, our idea was that if we, we sort of, if we built a chip that was the size of a dinner plate, we could stuff it to the gills with fast SRAM, overcoming the limitation of SRAM, which is it doesn't store very much, by putting a huge amount down, by using more silicon area. Now, if you're an SRAM solution today in a normal size chip, And you're trying to do a trillion parameter model, use four or 5000 chips. What a mess. You know how many cables that is? Do you know the, the, the impact to the AI? It's a horrible mess, right? And it limits you from doing things you want to do with the AI, uh, like speculative decode. It…
AI assessment note: “Well, not only is it fair, it's the reason we went to wafer scale.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q How do you think about that? Is that just pricing power, which they are taking advantage of?
A Absolutely. I mean, the, the, the short answer is, why does it make sense for, uh, AWS to, to, to build a training part? Well, because they want to get rid of the 78% gross margin that NVIDIA is charging. That's why it makes sense. Just like that. It might be, on the high-end chips, 85%. And I, I think people don't like that historically. Historically, people sort of put that in the back of the mind and they remember it. And, you know, when Intel stumbled, the, the number of people who came out of the, the woodwork to kick them when they were down was extraordinary. And it was sort of years of pent up Sort of frustration, uh, came out when, when the giant stumbles. And I, I think, uh, we've seen that again and again.
AI assessment note: “because they want to get rid of the 78% gross margin that NVIDIA is charging.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q are putting a lot of money into compute and hardware and compute and hardware startups. And I wanted to ask you, should I be following them? What should I be looking for in them? I'm seeing some incredibly young founders. I was with a nineteen-year-old founder trying to take on NVIDIA this morning, um, raising twenty million dollars for a pre-seed. How, how should I be thinking about this, Andrew?
A Well, I, if you don't know a lot about hardware, I wouldn't invest in hardware. I think, uh, Harry, it's probably the same in many things, is that, um, I think hardware is not an easy place to make money. It's a place that has historically rewarded experience, both from investors and from entrepreneurs. I think the number of different technologies involved in designing a chip, Uh, is extraordinary. Not just the, the logic, which is what most people think about. When, when you think about chip design, that's just the front end part that that that's writing in very low level software. Um, but the, the selection of tools, right? We pay millions of dollars a year in tools, selection of geometry, right? Which fab and having a relationship with a fab, you're going to pay 20 or thirty million in NRE. And if you have a bug in your, in your chip, you got paid again.
AI assessment note: “if you don't know a lot about hardware, I wouldn't invest in hardware.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Can I just interrupt and ask what is yield and why is it impossible to solve?
A Ok, ah, a wafer become, begins, it's a, ah, ah, a 12 inch, ah, diameter circle, slice of, of silicon. And your, your chip is punched out of this the way your mother might take a cookie cutter and cut out cookie dough. And, ah, during the process at some point, just like your mom might have done, she lifts up the edges, and all the little bits are removed, and what's left are just the cookies. Those are your chips. Um, now what happens is there are a set of naturally occurring flaws, and that's like your mother closing her eyes and throwing up a handful of, of M&Ms. Now, the bigger the cookie, right, the higher probability you hit an M&M. The bigger the chip, the higher the possibility that you have a flaw. And traditionally what you did when you had a flaw was you threw away the chip, or you sold it as a less valuable part. You shut down part of the chip and sold it as a less valuable part, something called binning. So, every wafer is going to have flaws. The bigger your chip, the higher probability you hit a flaw, and the more way, the more part of silicon is wasted when you throw it away. This is what everybody thought was known truth. And one of the, the things our, our team realized was that there are other ways to handle flaws. Like, what, what if instead You built your computer, you built your processor out of hundreds of thousands of identical tiles, and say there was a …
AI assessment note: “The bigger the chip, the higher the possibility that you have a flaw.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Are export controls being implemented properly? Do you think that is a good idea? You know, everyone was going with Deepsea, wow, how did this happen? They must have stolen chips. How could this be? What do you think about export control?
A It turns out that they probably did use chips in Singapore. Um, I think the following, I, I think managing Software and managing hardware compliance are, are extremely different things. Because their, their vector of diffusion is different. There's different weights. Right? If you sell a, a server that weighs five or 600 pounds, arrives on a pallet, you can go visit it. Right? You, you want to deploy it in Kazakhstan. You can put a data center, and you can have somebody from the embassy visit it. Take photos of it once a month. It's not going anywhere. Right? You, you can keep track of who uses it and provide logs, and that's much, much harder with software. And open source is a, a whole nother level. Right? And so, um, that's the first observation. The second is that We had, I got to know the, the, the leadership in commerce in the previous administration. I didn't always agree with their policies, but it is a world of unintended consequences. You, ah, sought to limit Chinese access to EDA tools to delay the growth of a Chinese chip market. And so U.S. venture capitalists backed tons of Chinese companies in Shenzhen to build EDA tools, right? I mean, right? Right. This is a unbelievably slippery, dynamic, challenging problem, and I don't know if it's a tractable problem to delay another nation's progress on a technical trajectory Is an enormously challenging thing. And, ah, I,…
AI assessment note: “This is a unbelievably slippery, dynamic, challenging problem, and I don't know if it's a tractable problem”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q It's been an extraordinary nine years. Can you just help me understand how does the movement into an age of AI change the requirements from a chip perspective of what is needed for a provider and how that then resulted in how you built Cerebris?
A The way to think about, uh, A chip is, it does two things. It does calculations, and it moves data, right? This is what, what, what a chip does. And, ah, sometimes along the way, it stores data. And so, ah, what AI presented was a very unusual combination of challenges. First, the underlying calculation is trivial. It's a matrix multiplication, and an FMAC can be developed by any second year electrical engineering student. So you say to yourself, holy cow, this has a huge number of very, very simple calculations. The hard part with AI work is results and intermediate results have to be moved a lot. And therein is the most complicated part. They have to be moved to memory and from memory. And they have to be broken up and moved among GPUs. And what we saw was that this was going to be the hard problem. And that if we could solve for that problem, we would build an AI computer that was faster and use less power.
AI assessment note: “The hard part with AI work is results and intermediate results have to be moved”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q When you speak about kind of being the fastest and across all benchmarks being the fastest, what matters the most? Is it being the fastest? Is it being the most efficient? Is it being the least costly? How do you think about the stack of prioritization for your customers?
A I think it varies. I, I think, uh, um, look, if, if, if you go to, to, to, to get a cancer diagnosis on, God forbid, your mother or, uh, your wife, I think, uh, uh, 93% accuracy is just plain not as good as 94% accuracy. And you pay a lot and wait another week to understand what the accuracy is, right? You pay a lot. Right. Now, on the other hand, uh, if you want Llama- four or five B to generate data to help you tune Llama- seven DB, uh, maybe you can wait a few days, three days, a week, more. You don't, there's no urgency there. On the other hand, if you want an answer from perplexity, Right? You don't want to wait 45 seconds for a search answer. Right? You don't want to wait in a chat. You don't want to wait three minutes for R-one on GPUs to give you an answer. What we know is that in interactive mode, milliseconds matter. In interactive mode, what Urs Holtz over at Google years ago showed was that you can destroy your user's attention With milliseconds of delay. So being the fastest matters everything in that domain. So I, I think what you have to do is sort of be thoughtful and say in some cases being the fastest doesn't matter. Uh, we'll call those batch. Uh, lots of maybe the cheapest matters there. In other domains, there is no search if you gotta wait eight minutes to get an answer. Right? That, that, that's not a product. That, when you go fast, a whole set of, of ne…
AI assessment note: “in some cases being the fastest doesn't matter... In interactive mode, milliseconds matter.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q How do you think about how the cost of inference goes down with the surge of demand that we mentioned, you know, over a hundred X, does the price reduce a hundred X? Does it follow Moore's law continuously? How do we think about the ever reducing price of inference?
A Look, there, there are, ah, the cost of inference is built up of, of several pieces, right? There's the power and space that is consumed to generate the response, right? That, that, that's a data center cost. That's an OpEx item, number one. Number two, there's the, ah, cost of, of the, the computer. We can drive down the cost of the computers with each generation by driving up their performance, et cetera. The other thing we can do is we can develop more efficient algorithms. Our AI algorithms today are not particularly efficient. There's a tremendous amount of room. Uh, in a GPU, most of the time it's doing inference, it's five or seven percent utilized. That means it's 95 or 93% wasted. So, over time, I think as an industry, we get better at things. We can drive the cost of compute down, we can build more efficient data centers with lower PUEs, and our algorithms will get more efficient so that our utilizations on our now cheaper computers are higher, so you get a higher percentage of the maximum number of flops. You get more tokens per unit time for the same power.
AI assessment note: “the cost of inference is built up of, of several pieces”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Um, is there a role for, uh, Local AI and local chips. NVIDIA had some announcement around just building chips for Windows computers. Is it for inference? Is that something that you guys, uh, think should be part of the multi-cylical ecosystem?
A Yeah, I think, um, if you look at the way the ecosystem for apps and cloud emerged, uh, everything you can do, you should do on your phone or your laptop. But the, the, the ability to get real processing power to a phone or to a laptop is constrained because they're generally working off a battery. And so you, you want to do as much as you can, as, as close to the data as you can. But the truth is, is in most situations, for real compute, you have to go to the cloud, to the data center. And that's exactly the way it's going to be with AI. We're going to do a lot of work Uh, uh, on the cell phones, uh, on a laptop, but for the big work, you're going to go to, to the data center. And that's where our focus is. Our focus is, is data center compute free.
AI assessment note: “for the big work, you're going to go to, to the data center”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah. And to the general chip versus pistolized chip, uh, it was actually interesting that you guys in your city celebrated when Grok was, uh, acquired. Was that, uh, what was that? Was it a recognition by NVIDIA that your vision was right all along?
A Yeah. I, I think, uh, One of those ideas sort of most durable modes was the perception that the GPU could do everything, and it was all you needed for AI. And the acquisition of Grok for twenty billion dollars, and the structure, and the speed with which they chose to do it, made clear to everyone that that wasn't true. That, uh, the GP architecture couldn't do, could not do fast inference, and that this market was large and growing quickly. And we were the fastest at it, and the largest, and, you know, our sales were more than 10 times the Grox, and they paid twenty billion dollars for the number two collector. So that was a good day. That was a great day.
AI assessment note: “made clear to everyone that that wasn't true”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q comes from the big labs, which, uh, you know, themselves are financed by, uh, venture capital, private equity, edge funds, where you call it, you know, sovereign investors. Um, is, is there any concern that, you know, demand for chips comes from the labs, which Are maybe artificially financed and that the, you know, if you had a true circle of deals here and there, then you sense any infrigility?
A Um, I, I, that's not exactly our experience. I mean, obviously we have, uh, enormous demand from pull open in. But we have huge, you know, dozens of other customers who, who are trying to place very big orders. And historically bubbles were when supply got out ahead of demand, right? When, uh, in the nineties, we built out a telco infrastructure, right? We built out fiber years before it was Going to be used and it took six or eight years and it all got used, but it was a sort of, if you build it, they will come mentality. Whereas what's different about AI right now is we're all trying to catch up. Um, we're trying to build data centers faster. We're trying to increase our, our demand, our, our, our supply chains for demand that's already here today. And only a very small portion of the world are using AI anywhere close to its potential. And we're already sort of overwhelmed with compute. We're, um, overwhelmed with, uh, the demand for, for memory, which was a real weakness in the GPUs. It's not a problem we face. And the ecosystem, there are constraints left and right, and that, that doesn't feel like a bubble.
AI assessment note: “that's not exactly our experience”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Okay. Wow. And so what lesson do you learn from that as a CEO? Like what is in your mind?
A I think a couple of things. I think to, to do the job that I love to build. I think, uh, making money is really great and making money for people you care about is really, really great. And when you get a chance to deliver for people who bet on you, who, who bet chunks of their career, right? Your investors, it's great to deliver for them. They bet on you, but they're diversified. They bet on you and 20 other companies. When someone bets five or seven years of their career and a career is 30 years, Right. They're making a sixth of their professional career. And when you get to deliver for them and, and, uh, they get to, to achieve the financial goals that they wanted. That's a great feeling and I'm proud of every day.
AI assessment note: “making money for people you care about is really, really great”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q runs on Brex, so I can spend time on building and not busy work. It's time to get Brex AF. Learn more at brex.com slash sorcery. That's B-R-E-X dot com slash S-O-U-R-C-E-R-Y. Bye. So looking forward, I'm sure this is also just a mark in your journey because you're a builder. So what are you most looking forward to over the next couple months as we said across this year?
A Look, getting to an IPO is, is not the, the, the end of a journey. It's sort of a plateau. It's sort of the arrival at corporate adulthood. It is the achieving what one plateau so that you can climb others. And our opportunity has gotten bigger. We have more resources. We have, uh, we're better recognized. Uh, we can reach more people. Um, and we can sort of prosecute our vision and our, our ambitions, uh, with more fuel. And so that, that, that's what we're excited about every day. You know, building more chips, building more data centers, inventing technology that, that, that moves the industry forward. Um, that's what, what drives us and gets us out of bed every day.
AI assessment note: “building more chips, building more data centers, inventing technology that moves the industry forward”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What do you think that the biggest misconception with that process is and the challenges are?
A Well, I, I think that the misconception is that it's easy and all you need to do is get in a room and, um, it's a, uh, it's a very hard problem. Um, you know, the, the, the software guys think one way. The hardware guys think a slightly different way. Um, you know, anything you do to make it easier to, to, To, to write the software makes it harder to do the hardware, right? And these are really hard trade-offs, and so bringing them together, um, and means these sort of compromises where it will be harder here to make it easier here. And that means somebody's schedule is going to be impacted. Somebody's got to add resources. Um, those discussions are enormously difficult.
AI assessment note: “the misconception is that it's easy and all you need to do is get in a room”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q As we look forward, I guess more on the macro lens, the proliferation of AI and everything that we're be able, we're able to now create and build and do. What are you excited about on the externalities that come with all of this?
A Look, I, I think, um, that we have a chance for our children or the next generation, not only to not die from cancer, but to not know anybody who died from cancer. I think that is a, a real, uh, achievable goal in 25 years. Wouldn't that be something? Um, you know, when we think about what AI can do, uh, writing better code is cool. And there's a huge market for that. But I, I think what, what it can do to better humanity is rid us of the number one killer of, of adults. And, um, I think, you know, that, that's when, when I think of what we're all doing this for, it's an outcome like that. You know, pancreatic cancer had a huge breakthrough recently. I, I think they're, uh, the opportunity for, for breakthroughs right now is, It has never been better, and AI is a, an extraordinary tool in, in pursuit of, of, of knocking down, um, major, major human killers, right? I mean, if you think, if you, if you took out cancer and you took out automobile accidents, you're taking out huge numbers of deaths a year. Um, and you say to yourself, well, that's a lot of good. That we did. And I, I think, you know, self-driving, humans are terrible drivers.
AI assessment note: “we have a chance for our children or the next generation, not only to not die from cancer”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you think we will see a peaking of demand? You've seen so many different...
A Not if AI continues to improve in usefulness. I mean, what's happened here is, and this is something that I haven't heard others sort of talk about, is that somewhere in twenty-twenty-five, the models got smart enough to be really useful. Before that, Harry, these were sort of, sort of a novelty. AI was like, cool, and then nobody used it. Remember, we, we make AI with training, and we use it with inference. And so once the, the AI we made, 20, twenty-five-ish, first half, got smart, we began using it. And this explosion in demand that Jensen described, alright, and that we very much agree with is happening, that's because people are using it every day. And they're using it on more and more problems. They're using it on harder problems, and it is sweeping through different demographic groups. It's not just twenty-eight-year-olds in Silicon Valley. It's my eighty-five-year-old father. It's, right, it's eleven-year-old, my eleven-year-old niece. It is, right, it is sweeping through demographic groups, and they're using it all the time. And that is what's driving this demand. And so, um, If we continue to find ways to make the AI, the frontier models, smarter and more useful, we'll keep using it. The demand will continue to, to con, on this sort of exponential curve.
AI assessment note: “Not if AI continues to improve in usefulness.”