Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q The spend has to materialize into actual tangible revenue back. And if it doesn't, whether you're in the mag seven or not doesn't, doesn't matter. Correct?
A That's correct. But, um, right now, AI is returning massive value already. It's very lumpy in the, uh, in the applications, but it's returning massive amounts of value. Let me talk about an example that actually happened for us. So I've tried a little bit of vibe coding. Um, I'm not the best in the world at it. We've got some, uh, interns who are amazing at it, and we, we had this customer visit us, or, um, and I had a meeting with them, and so they asked for a feature, and I spec'd it out, very high-level, vibey, Um, so I was prompt engineering the engineers, and four hours later, it was in production. Not a single line of code was written by a human being. There was no debugging done by a human being. It was all prompting. Um, I think we even have Slack integration now, where you push in, like, you commit things through Slack. So, all that was done, four hours later, it's in production. Think about the value there. But now, imagine, fast forward six months from now, when that could happen before the customer meeting's over. It's a qualitative difference. It's not even just a dollar amount difference. Yes, um, you know, when you're able to do it that fast, you spend less to get the feature into production. That's real ROI. However, qualitatively, when you can do that before the customer meeting is over, you're gonna be able to win deals that your competitors won't.
AI assessment note: “That's correct. But, um, right now, AI is returning massive value already.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And who produces HBM? I'm sorry for the dumb questions.
A There's three companies in the world that do this. Um, SK Hynix, Samsung, and Micron. And it's a specialty memory. It's only used in high-end servers, so there's a limited quantity that's built. It's very expensive to ramp up. It's a very technically challenging type of memory to build, more so than others. And so there's a very limited supply. And GPUs are so fast, computationally, that if you were using regular memory, it'd be like drinking out of a martini straw. It would just take forever. This is why you see people preferring to do, um, even inference, but especially training on, um, GPUs rather than CPUs, because the memory bandwidth is too limited. And CPUs rarely use HBM. They're mostly regular memory. Um, our architecture, the observation that we had when we started Grok, everyone knows Moore's law. Every 18 to 24 months, like clockwork, double the transistors. It means double the compute. But we noticed that AI was getting better faster, and it clearly wasn't the algorithms, because algorithms have sort of discontinuous jump. It also, uh, didn't seem to be, um, the data, because there wasn't that much more data. And the transistors were only doubling every 18 to 24 months, so where was all of this capability coming from? Turns out, the number of chips was also doubling every 18 to 24 months. So rather than two X, it was four X. So the question we asked was, if you're …
AI assessment note: “There's three companies in the world that do this. Um, SK Hynix, Samsung, and Micron.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q But there has to be a ceiling on efficiency. No?
A Does there? So there's a mathematical limit. So if you, if you study computer science, you've probably heard of something called big O complexity. Big O complexity is, um, you know, if I am solving a problem, And I look at how I solve it. I might need to take more steps if I solve it with one algorithm versus another. So for example, quick sort versus bubble sort. Quick sort, I need n log n steps. Bubble sort, I need n squared. What's the difference? If I'm sorting 1000 numbers, n log n, that's 10,000 steps. But with, um, n squared, that's a million steps because it's either 10 times a thousand or thousand times a thousand. One of the reasons that these LLMs struggle to multiply large numbers is because, um, multiplying is not linear. These LLMs could do anything linear without, you know, needing to think, but just like on a piece of paper how you need to write out all those intermediate steps, these LLMs need that intermediate space in, in order, in, in those steps in order to compute these things. It's a mathematical requirement. There's nothing, you cannot train a model enough so that it'll see any arbitrarily large number, just be able to multiply it. But you can choose bigger and bigger groupings of numbers for it to memorize, in which case it can do it in fewer steps. And effectively, as you are training the model on more and more data, it's seeing more and more examples.…
AI assessment note: “So there's a mathematical limit.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Why is 40% of their revenue inference then, and what Why have you not taken so much more of that?
A At the beginning of 2024, we only had 640 chips. At the end we had 40,000. We're not at that scale yet. So you have to, you have to provide quality, you have to provide low cost, you have to provide speed, but you also have to provide capacity. And so this is where that, that, um, most important part of not using HBM came in. It means that we effectively have no scale limits. So the GPU itself is actually Um, manufactured using the same process that, um, you use for your mobile phone, right? So the same silicon that's in your mobile phone is the same silicon for the GPU. In fact, they build the mobile phones first, uh, the mobile phone chips first because they're smaller, so they yield better. So the, the NVIDIA actually gets it after Apple. The difference is that memory. That's the only difference, but that memory is the hard part to manufacture that's, that's limited in scale. So by us avoiding that, we effectively have almost no limit on how much we can scale up, and that's important for inference.
AI assessment note: “At the beginning of 2024, we only had 640 chips... We're not at that scale yet.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q But there has to be a ceiling on efficiency. No?
A Does there? So there's a mathematical limit. So if you, if you study computer science, you've probably heard of something called big O complexity. Big O complexity is, um, you know, if I am solving a problem, And I look at how I solve it. I might need to take more steps if I solve it with one algorithm versus another. So for example, quick sort versus bubble sort. Quick sort, I need n log n steps. Bubble sort, I need n squared. What's the difference? If I'm sorting 1000 numbers, n log n, that's 10,000 steps. But with, um, n squared, that's a million steps because it's either 10 times a thousand or thousand times a thousand. One of the reasons that these LLMs struggle to multiply large numbers is because, um, multiplying is not linear. These LLMs could do anything linear without, you know, needing to think, but just like on a piece of paper how you need to write out all those intermediate steps, these LLMs need that intermediate space in, in order, in, in those steps in order to compute these things. It's a mathematical requirement. There's nothing, you cannot train a model enough so that it'll see any arbitrarily large number, just be able to multiply it. But you can choose bigger and bigger groupings of numbers for it to memorize, in which case it can do it in fewer steps. And effectively, as you are training the model on more and more data, it's seeing more and more examples.…
AI assessment note: “So there's a mathematical limit. So if you, if you study computer science”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And who produces HBM? I'm sorry for the dumb questions.
A There's three companies in the world that do this. Um, SK Hynix, Samsung, and Micron. And it's a specialty memory. It's only used in high-end servers, so there's a limited quantity that's built. It's very expensive to ramp up. It's a very technically challenging type of memory to build, more so than others. And so there's a very limited supply. And GPUs are so fast, computationally, that if you were using regular memory, it'd be like drinking out of a martini straw. It would just take forever. This is why you see people preferring to do, um, even inference, but especially training on, um, GPUs rather than CPUs, because the memory bandwidth is too limited. And CPUs rarely use HBM. They're mostly regular memory. Um, our architecture, the observation that we had when we started Grok, everyone knows Moore's law. Every 18 to 24 months, like clockwork, double the transistors. It means double the compute. But we noticed that AI was getting better faster, and it clearly wasn't the algorithms, because algorithms have sort of discontinuous jump. It also, uh, didn't seem to be, um, the data, because there wasn't that much more data. And the transistors were only doubling every 18 to 24 months, so where was all of this capability coming from? Turns out, the number of chips was also doubling every 18 to 24 months. So rather than two X, it was four X. So the question we asked was, if you're …
AI assessment note: “There's three companies in the world that do this. Um, SK Hynix, Samsung, and Micron.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Why is 40% of their revenue inference then, and what Why have you not taken so much more of that?
A At the beginning of 2024, we only had 640 chips. At the end we had 40,000. We're not at that scale yet. So you have to, you have to provide quality, you have to provide low cost, you have to provide speed, but you also have to provide capacity. And so this is where that, that, um, most important part of not using HBM came in. It means that we effectively have no scale limits. So the GPU itself is actually Um, manufactured using the same process that, um, you use for your mobile phone, right? So the same silicon that's in your mobile phone is the same silicon for the GPU. In fact, they build the mobile phones first, uh, the mobile phone chips first because they're smaller, so they yield better. So the, the NVIDIA actually gets it after Apple. The difference is that memory. That's the only difference, but that memory is the hard part to manufacture that's, that's limited in scale. So by us avoiding that, we effectively have almost no limit on how much we can scale up, and that's important for inference.
AI assessment note: “At the beginning of 2024, we only had 640 chips... We're not at that scale yet.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Will you and NVIDIA move into the model there? Everyone talks about model providers becoming application providers or infrastructure providers become model providers.
A We have decided that we're not going to train our own models. We'll do a little fine tuning for specific cases or whatnot, but we don't want to compete. And that's really important because people are putting their models with their weights on us. Right? And they don't want us to learn from and take that stuff for our own benefit. This is the problem you have when you work with a hyperscaler because you know they're also doing everything that you are doing. So we've decided model providers, you make the model, we don't do that. I think there's also the data side of the users and the queries. So the other thing that we could do that we do not do is log the queries, and then we've got data if we wanted to train. We don't train. We have no reason to hold the data. So we, we only temporarily store things in the DRAM, so there's no persistent storage. If the power went out, everything's gone. And DRAM is limited, so we can't hold things for a long time. So you know that we don't have your data. Now, people who are building businesses on top of us, you can obviously keep the data from your customers if you want. We have no control over that. That's fine. But we don't take any data.
AI assessment note: “We have decided that we're not going to train our own models.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I'm sorry, does this news not ridicule the five hundred billion dollar announcement? At a time when we've seen increasing efficiency to a scale like never before with DeepSeat today, the five hundred billion dollar seems ridiculed.
A Actually, I don't think it's enough spending. And, and the reason is, so we saw this happen at Google over and over again, right? So we, we do the TPU and the TPU. So why did we do the TPU? The speech team trained a model. It outperformed human beings at speech recognition. This was like back in 2011, 20 12, right? It was the first time. And so Jeff Dean, most famous engineer at Google, Um, gives a presentation to the leadership team. It's two slides. Slide number one. Good news. Machine learning finally works. Slide number two. Bad news. We can't afford it. And we're Google. We're gonna need to double or triple our global data center footprint at probably a cost of 20 to forty billion dollars, and that'll get a speech recognition. Do you also want to do search and ads? So it turns out there's always this giant mission accomplished banner every time someone trains a model, and then they start putting it into production. And then they realize, oh, this is gonna be expensive. This is why we've always focused on inference. And so, now think about it this way. At Google, we always ended up spending 10 to 20 times as much on the inference as training back when I was there. Now the models are being given away for free. How much are we gonna spend on inference? And I, I, Garrett, and now with the test time compute, Right. And like, I I've, I've asked questions of deep seek where it to…
AI assessment note: “Actually, I don't think it's enough spending.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I'm sorry, does this news not ridicule the five hundred billion dollar announcement? At a time when we've seen increasing efficiency to a scale like never before with DeepSeat today, the five hundred billion dollar seems ridiculed.
A Actually, I don't think it's enough spending. And, and the reason is, so we saw this happen at Google over and over again, right? So we, we do the TPU and the TPU. So why did we do the TPU? The speech team trained a model. It outperformed human beings at speech recognition. This was like back in 2011, 20 12, right? It was the first time. And so Jeff Dean, most famous engineer at Google, Um, gives a presentation to the leadership team. It's two slides. Slide number one. Good news. Machine learning finally works. Slide number two. Bad news. We can't afford it. And we're Google. We're gonna need to double or triple our global data center footprint at probably a cost of 20 to forty billion dollars, and that'll get a speech recognition. Do you also want to do search and ads? So it turns out there's always this giant mission accomplished banner every time someone trains a model, and then they start putting it into production. And then they realize, oh, this is gonna be expensive. This is why we've always focused on inference. And so, now think about it this way. At Google, we always ended up spending 10 to 20 times as much on the inference as training back when I was there. Now the models are being given away for free. How much are we gonna spend on inference? And I, I, Garrett, and now with the test time compute, Right. And like, I I've, I've asked questions of deep seek where it to…
AI assessment note: “Actually, I don't think it's enough spending.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q One would think that with commoditization of models and with cheaper inference that actually big tech wins, right? Have you seen the stock market today? They've been hit hard. How do you think about that?
A What you see is a bunch of people who are concerned about training and the need for it, and everyone's still thinking that most of compute is training and that there's going to be less of it because someone trained a model on Um, 2000 GPUs and the nerfed, you know, a 800 version with slower memory or whatever it is. And they're like, oh, people aren't going to need as many chips. But again, like Jeven's paradox, right? Which is the more you bring the cost down, the more people consume. So for the last five to six decades, like clockwork, once a decade, the cost of compute has gone down a thousand X. People buy 100,000 Uh, X as much compute spending a hundred times as much. So every decade they spend a hundred times as much. So you make it cheaper and they want more. And so what's really happening is every time one of these models gets cheaper, we see our developer count just skyrocket. It just like goes up and then it comes back down a little bit, but the, the slope is higher than when it started. So better models create more demand for inference, more demand for inference. Then has people going, I should train a better model, and the cycle continues.
AI assessment note: “Which is the more you bring the cost down, the more people consume.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you worry that if we have a speed bump in the short term, it will derail significant parts of the economy given the concentration of value? Everyone rips today, but if NVIDIA, MATA, Google, Microsoft suddenly hit speed bumps and the AI speed train It's just slowed down. The consequent multiplier effect is mega. Do you worry about that?
A Yeah, and this is, this is independent of the value of AI. This is the sort of control system, um, theory of what's going on, right? So, a stock market could inherently be on an upward trajectory. It can overheat, and that overheating causes it to run away. People bid things up, they realize they've made a mistake, and then it has to come back down, and then it dips below where it should be, spending, um, retreats, And then people don't have the funds they need to build their businesses. Uh, a lot of good businesses can die during one of these downward trends, but this is also where the best businesses are made. How many times do you see a downturn, um, and a ton of amazing businesses come out of it?
AI assessment note: “Yeah, and this is, this is independent of the value of AI.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you think OpenAI will be able to move into the chip player? At some point, Nvidia must be concerned that the OpenAI will want to verticalize and own the chip player as well. Do you think they will be able to make that successful transition?
A Um, I think one of the problems in building your own chip is it's really, first of all, everyone thinks that building the chip is the hard part. Uh, and then as you do it, you start to realize building the software is the hard part. And then as you do it, you realize keeping up with where everything is going starts to become the hard part. I have no doubt that, uh, OpenAI will be able to build its own chips. I have no doubt that, Um, eventually, Anthropic will be building their own chips, that every hyperscaler will build their own chip. One of the things that, um, I, I had this experience when I was at Google, where, um, I, I got a lab tour. And this was before AMD was doing a great job, right? AMD was struggling for a little while, and now they're doing great. But, um, they had built 10,000 servers, and those 10,000 servers of AMD chips, I was walking through the lab, and they were pulling the servers out of the racks, Taking the AMD chip, popping it off, and throwing it in a trash can. And the funny thing was, it was almost preordained, because everyone knew that in that generation, Intel was gonna win. So, why did Google build 10,000 servers? Because they wanted to, um, get a discount on the Intel chips they bought. And when you're at that scale, the cost to design your own server, because they had to design their own motherboard in order to fit the AMD chip, uh, and, and t…
AI assessment note: “I have no doubt that, uh, OpenAI will be able to build its own chips.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Will, will China not just subsidies the inference and the running there? I understand. Yes. So does it matter if their cost of running is higher, but their Chinese, the CCP will just subsidize it. Does it matter?
A There's a home game and there's an away game. The home game is we want to, um, build enough, uh, compute for the United States. The away game is we want to build it for our allies, right? Europe, South Korea, Japan, um, India, and so on. And the advantage that the, so China can, can win their own home game. They're going to build a 150 nuclear reactors. So they're going to have enough energy, even though their chips aren't as energy efficient. Uh, and they can subsidize as you mentioned, but the away game is different. If a country only has a hundred megawatts of power, What are they going to do? Build another nuclear power plant? Like, that's just not a realistic thing. You can do that in China. You can't do that elsewhere. So having a better chip gives you an advantage in the away game. So my expectation is that right now for the next two to three years, the United States has a clear advantage in that away game over China. And if we move very quickly, then we're going to be able to bring a bunch of allies into the AI race.
AI assessment note: “having a better chip gives you an advantage in the away game”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q do just have to ask about it. Do you think this is an enduring and sustainable market? When you look at, um, a lot of the use cases today, they're quite transient. Do you, how do you analyze the future of the Vibe coding market, having played with it a little bit, and having seen also interns, as you said, who are very good at it internally, use it well?
A Um, vibe coding is going to be, so let's take reading. Reading used to be, reading and writing used to be a career. If you were a scribe, you were one of the small percentage of people who knew how to read and write, and people would hire you just to record things, and you, you did much better than the average person in the economy because of that, because it was a specialized skill. Coding has been the same thing. Very small percentage of the population did it, Took, you know, a couple years to learn how to do it well. Some people were really good at it. Now everyone reads. Everyone writes. It's not a special skill. It's expected in every job. And coding is gonna become the same thing. For you to be in marketing, you're gonna have to be able to code. For you to be, um, in customer service, you're gonna have to be code, uh, be able to code. Uh, I was having dinner with someone who runs a chain of 25 coffee shops, has never coded in their life, and they vibe coded, uh, a supply chain Tool that allowed them to check inventory. They didn't write a single line of code. They got it to work. And it was funny because they discovered all the problems that we software engineers discover over time. They started getting feedback from their employees like, this feature doesn't work. It doesn't, this thing doesn't work when I do this. All the little edge cases. And then you just started fix…
AI assessment note: “coding is gonna become the same thing. It's expected in every job.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q The spend has to materialize into actual tangible revenue back. And if it doesn't, whether you're in the mag seven or not doesn't, doesn't matter. Correct?
A That's correct. But, um, right now, AI is returning massive value already. It's very lumpy in the, uh, in the applications, but it's returning massive amounts of value. Let me talk about an example that actually happened for us. So I've tried a little bit of vibe coding. Um, I'm not the best in the world at it. We've got some, uh, interns who are amazing at it, and we, we had this customer visit us, or, um, and I had a meeting with them, and so they asked for a feature, and I spec'd it out, very high-level, vibey, Um, so I was prompt engineering the engineers, and four hours later, it was in production. Not a single line of code was written by a human being. There was no debugging done by a human being. It was all prompting. Um, I think we even have Slack integration now, where you push in, like, you commit things through Slack. So, all that was done, four hours later, it's in production. Think about the value there. But now, imagine, fast forward six months from now, when that could happen before the customer meeting's over. It's a qualitative difference. It's not even just a dollar amount difference. Yes, um, you know, when you're able to do it that fast, you spend less to get the feature into production. That's real ROI. However, qualitatively, when you can do that before the customer meeting is over, you're gonna be able to win deals that your competitors won't.
AI assessment note: “That's correct. But, um, right now, AI is returning massive value already.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q financial question, but we see the S&P about to hit 7000. We see this ripping of the mag seven, like we haven't seen a concentration of value in many, many years. And people suddenly start to feel like, wow, it's getting toppy. I listen to you, and I hear all of this, and I think, it's just the start. How should I think about the duality of those two thoughts?
A There's two components to the value. Um, one is the weighing machine and one is the popularity contest. And there are some products that are pure popularity contest like crypto. I have never bought a Bitcoin, you know, I missed out. Why? Because I can't play in the popularity contest. I'm not good at it. I don't know what's going to be popular and what isn't. All I can do is I can see value. When I look at AI, I see real value being delivered. Best example, PE firms are all over us. They want access to cheap AI compute, because every time they get more cheap AI compute, they can bring the, the, they can change the bottom line of their businesses. It has real value. When PE firms go after something and see value in it, it's not a popularity contest. It's pure value. And so what happens is the reason companies get a large multiple is people see that the valuation is, the actual value is going to accrue, or they get hype cycled on it, and it's pure popularity contest. And there are different participants in the market. Some of them are just playing the popularity contest. Others are looking at the value, and they may come to the same conclusion for different reasons. Coming at it from the value point of view, the weighing machine point of view, the most valuable thing in the economy is labor. And now we're going to be able to add more labor to the economy by producing more compute…
AI assessment note: “There's two components to the value. Um, one is the weighing machine and one is the popularity contest.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you think OpenAI will be able to move into the chip player? At some point, Nvidia must be concerned that the OpenAI will want to verticalize and own the chip player as well. Do you think they will be able to make that successful transition?
A Um, I think one of the problems in building your own chip is it's really, first of all, everyone thinks that building the chip is the hard part. Uh, and then as you do it, you start to realize building the software is the hard part. And then as you do it, you realize keeping up with where everything is going starts to become the hard part. I have no doubt that, uh, OpenAI will be able to build its own chips. I have no doubt that, Um, eventually, Anthropic will be building their own chips, that every hyperscaler will build their own chip. One of the things that, um, I, I had this experience when I was at Google, where, um, I, I got a lab tour. And this was before AMD was doing a great job, right? AMD was struggling for a little while, and now they're doing great. But, um, they had built 10,000 servers, and those 10,000 servers of AMD chips, I was walking through the lab, and they were pulling the servers out of the racks, Taking the AMD chip, popping it off, and throwing it in a trash can. And the funny thing was, it was almost preordained, because everyone knew that in that generation, Intel was gonna win. So, why did Google build 10,000 servers? Because they wanted to, um, get a discount on the Intel chips they bought. And when you're at that scale, the cost to design your own server, because they had to design their own motherboard in order to fit the AMD chip, uh, and, and t…
AI assessment note: “I have no doubt that, uh, OpenAI will be able to build its own chips.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And so what you, the risk that you could take is to what, sorry, just specifically?
A We could just double, um, the rate at which we're building out supply. I mean, with this fundraise, we, we ended up raising, um, you know, more than twice what we were, you know, expecting to raise, and then we were forex oversubscribed over, um, over what we did raise, and so we could have raised a lot more money, it would have been more dilutive, um, and I'm trying to be dilution sensitive for investors and everyone else, Um, but on the other hand, we could have just raised more money, and we could have just built a ton of compute. Um, the other advantage that we have is, versus anyone else, our cost per token, especially given a, given a speed, um, is very advantageous. So, we know that we can charge less than the rest of the market, um, which matters when you're trying to build these businesses, not because people are, are spend conscious. If we lower, um, what we charge 50%, people are gonna buy twice as much. They're spending as much as they're making because whatever they spend increases the quality of the output.
AI assessment note: “We could just double, um, the rate at which we're building out supply.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So should founders building today, should they build with the assumption that scaling laws will continue? Should they build with what we have today? How do you advise them on that?
A I would advise you to build based on things getting better, but I would also focus a little more on the sort of big quantum steps. So, the analogy that I like is, if you look at the information age, we went through, we had the printing press, we had the telephone, we had the telegram, we had the, um, internet, and we had smartphones, right? And if you had built Uber back when we had, um, uh, internet, it wouldn't have worked because you'd book a ride, you'd go somewhere, how do you get home? Exactly, right? And, uh, we're in the same sort of space now, so we don't, the models hallucinate. So it would be hard to build a medical diagnosis company. It would be hard to build a legal company, right? However, if you are doing that, And the algorithmic enhancements happen that get the hallucination rate down, you are perfectly positioned, just like Grok. We were around for seven years before we had product market fit, right? We were around. Our bet was scaled inference. That inference was going to be the bottleneck that we were going to need to run really big, heavy models. Like, everyone was assuming you would have a single PCIe card running inference because training was the complicated part, right? The reality was we made the right bet ahead of time, and then we were perfectly positioned. Your job is not to follow the wave. Your job is to get positioned for the wave, and that's the…
AI assessment note: “I would advise you to build based on things getting better”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q We mentioned the different huge amounts of money that's being spent here. Is this a good bubble that bluntly lays the foundations for an incredible next 10 to 20 years where, bluntly, the capital actually turns out to be productive but not seemingly so on paper? Or is it where actually just a huge amount of money is incinerated on depreciating assets?
A I can guarantee you that a huge amount of money will be incinerated. But I also bet that in total more money will be made than will be put in. And so this is the problem. You have to look at it either in aggregate or individual bets, right? When everyone is making investments in the market, some people are going to lose money because not every company is going to be successful. So what you always see is when there is some real tech improvements or things coming, you've got the things that were early, that people are investing in heavily, that are super successful, and then everyone else wants to get in on it. And, you know, it goes from you have AI chips and AI models to now you've got AI, you know, t-shirts, and next thing you know, you've got AI thermal grease, right? It, it just, like, people just start applying AI to everything. Next thing you know, you'll have an AI condo.
AI assessment note: “I can guarantee you that a huge amount of money will be incinerated. But I also bet”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q China is quite opaque in everything. What do we not know about China that we would like to know?
A I think the most important thing to understand is where they're going to end up on the, the censorship and privacy of these models. We come from democratic countries. We have an expectation that companies can build something that says anything. Are they going to be permissive and allow models to make mistakes and hallucinate? Or are they going to shut it down? Because I think if you know that, you know whether or not China has a shot. One of the biggest nightmares that they have is free speech. It's the exact opposite of that vulnerability we talked about earlier. Can you imagine Xi Jinping going out and saying, country, we've lost our advantage in AI. I need your help. Never. Ever. It's always going to be, we're the greatest, we're the best. Everyone's going to know differently, but they're all going to have to toe the party line. Right now, because of that, I think it's really hard for them to just allow these models to, to say anything. Say, you know, the US is great and better at AI. That's a bad thing for them. And so that's going to really tell you a lot about the AI story in China.
AI assessment note: “where they're going to end up on the, the censorship and privacy of these models”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So what does a world where inference is so crucial and 20 times more important than training, what does that world look like?
A I think the simplest way to understand it is equate an LPU or a GPU to an employee, right? If, if you have Enough of them, the LPUs or GPUs, you can do work just like with an employee, but it's, it's a little different in the sense that, ah, they can't quit and take another job. Ah, you don't have to retrain. Once you get a model to a certain capability, it'll always be at least that capability, right? It's not going to regress. You, you, you know, um, so you get the consistency out of it. But now imagine that you're a startup And rather than having to go out and hire a hundred people, you hire 10, and you buy the amount of compute equivalent to 90 employees worth. That's a very different way of thinking about the world, because now CapEx, or in some cases different types of OpEx, can, can be used instead of just employees. And so, and in terms of inference, just to give you a sense of our scaling, we started 2024 with about 640 chips in production. We ended with over 40,000. This year we want to be at over two million, and next year the number is much, much, much larger.
AI assessment note: “rather than having to go out and hire a hundred people, you hire 10”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q We mentioned the different huge amounts of money that's being spent here. Is this a good bubble that bluntly lays the foundations for an incredible next 10 to 20 years where, bluntly, the capital actually turns out to be productive but not seemingly so on paper? Or is it where actually just a huge amount of money is incinerated on depreciating assets?
A I can guarantee you that a huge amount of money will be incinerated. But I also bet that in total more money will be made than will be put in. And so this is the problem. You have to look at it either in aggregate or individual bets, right? When everyone is making investments in the market, some people are going to lose money because not every company is going to be successful. So what you always see is when there is some real tech improvements or things coming, you've got the things that were early, that people are investing in heavily, that are super successful, and then everyone else wants to get in on it. And, you know, it goes from you have AI chips and AI models to now you've got AI, you know, t-shirts, and next thing you know, you've got AI thermal grease, right? It, it just, like, people just start applying AI to everything. Next thing you know, you'll have an AI condo.
AI assessment note: “I can guarantee you that a huge amount of money will be incinerated.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q China is quite opaque in everything. What do we not know about China that we would like to know?
A I think the most important thing to understand is where they're going to end up on the, the censorship and privacy of these models. We come from democratic countries. We have an expectation that companies can build something that says anything. Are they going to be permissive and allow models to make mistakes and hallucinate? Or are they going to shut it down? Because I think if you know that, you know whether or not China has a shot. One of the biggest nightmares that they have is free speech. It's the exact opposite of that vulnerability we talked about earlier. Can you imagine Xi Jinping going out and saying, country, we've lost our advantage in AI. I need your help. Never. Ever. It's always going to be, we're the greatest, we're the best. Everyone's going to know differently, but they're all going to have to toe the party line. Right now, because of that, I think it's really hard for them to just allow these models to, to say anything. Say, you know, the US is great and better at AI. That's a bad thing for them. And so that's going to really tell you a lot about the AI story in China.
AI assessment note: “most important thing to understand is where they're going to end up on the, the censorship”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Okay. Why is it such a huge deal? Let's unpack that.
A So up until recently, the Chinese models have been behind, um, sort of Western models. And I say Western, including like Mistral as well, and some other companies. And it was, um, largely focused on how much compute you could get. Most people actually, most don't realize this. Most companies have access to roughly the same amount of data. They buy them from the same data providers. And then they just churn through that data with a GPU, and they produce a model, and then they deploy it. And they'll have some of their own data, and that'll make them subtly better at one thing or another, but they're largely all the same. And the more GPUs, the better the model, because you can train on more tokens. It's the scaling law. This model was supposedly trained on a smaller number of GPUs and a much, much tighter budget I think the way that it's been put is less than the salary of many of the executives at Meta, and that's not true. It's, it's actually, there's an element of marketing, uh, involved in the deep seek release.
AI assessment note: “This model was supposedly trained on a smaller number of GPUs”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q So how is it gonna cost us to have a cup of coffee because of AI?
A Because you're gonna have, uh, robots that are gonna be farming the coffee more efficiently. You're gonna have better supply chain management. You're gonna, um, it's just gonna be across the entire supply chain. Um, you're gonna be able to genetically engineer the coffee so that you get more of it per, um, watt of sunlight. Right? Just across the entire spectrum. So, you're gonna have massive deflationary pressure. That's number one. And what that means is people will need to work less. And that's gonna lead you to number two, which is people are gonna opt out of the economy more. They're gonna work fewer hours, they're gonna work fewer days a week, and they're gonna work fewer years. They're gonna retire earlier because they're gonna be able to support their lifestyle working less. And then number three is we're gonna create new jobs and new, uh, company, uh, new industries that don't exist today. Um, think about a hundred years ago. 98% of the workforce in the United States was in agriculture. Two percent did other things. When we were able to reduce that to two percent of the population working in agriculture, we found things for those other 98% of the population to do. The jobs that are going to exist a hundred years from now, we can't even contemplate. A hundred years ago, the idea of a software developer made no sense. A hundred years from now, it's going to make no sense…
AI assessment note: “So, you're gonna have massive deflationary pressure.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q So what does a world where inference is so crucial and 20 times more important than training, what does that world look like?
A I think the simplest way to understand it is equate an LPU or a GPU to an employee, right? If, if you have Enough of them, the LPUs or GPUs, you can do work just like with an employee, but it's, it's a little different in the sense that, ah, they can't quit and take another job. Ah, you don't have to retrain. Once you get a model to a certain capability, it'll always be at least that capability, right? It's not going to regress. You, you, you know, um, so you get the consistency out of it. But now imagine that you're a startup And rather than having to go out and hire a hundred people, you hire 10, and you buy the amount of compute equivalent to 90 employees worth. That's a very different way of thinking about the world, because now CapEx, or in some cases different types of OpEx, can, can be used instead of just employees. And so, and in terms of inference, just to give you a sense of our scaling, we started 2024 with about 640 chips in production. We ended with over 40,000. This year we want to be at over two million, and next year the number is much, much, much larger.
AI assessment note: “rather than having to go out and hire a hundred people, you hire 10”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q Can you just walk me through how that deal is structured?
A Yeah. So we started off last year, right? And, and we got to, um, 19,000 of our chips deployed. We did that in about 51 days. And The question was, what can we do this year? So they've gone off, they've collected up a bunch of power in the country, and the deal is structured so that they will put up the capex for us to deploy our chips in that data center, or those data centers, and we pay back based on the money that we make. So it's, it's sort of a, it's a little bit different than debt, In that they participate in the upside, um, but it, it's similar in nature, but it is revenue, because we actually make profit up front.
AI assessment note: “the deal is structured so that they will put up the capex for us”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q So who wins and who loses? Like is massive and incinerate the largest amount of cash ever?
A I think the Keynesian beauty contest No longer applies here, because there's so much money available being spread out, and, and I think you're, you're gonna see that the people who have the best products are actually gonna be the winners, because everyone can be capitalized. But there will be problems for the winners because of this. The problems are gonna be of the sort, you, you had this employee that you were gonna hire, and someone offered them a ridiculous amount of money. Yeah, you see this all the time now. And they could have gone and contributed to the winner, but now they're contributing to a competitor that shouldn't even exist, right? Or is equally likely to win, and now you're splitting the talent.
AI assessment note: “the people who have the best products are actually gonna be the winners”