The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Andrew Feldman argument clarity score 4.2/5 from 12 exchanges on raw tape · average scores: directness 4.3 · coherence 4.4 · precision 4 · compression 3.8 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
12exchanges match
12on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q so I want to do a bit of a deep dive on that on itself in a, in a second. Um, uh, but to close on this, so three bottlenecks, um, you also mentioned CPUs a couple of times in this conversation, and there seems to be a theme around the emergence of like a CPU shortage as well. Is that, so what is that true? Two, what causes it?

A So agentic AI Is a world in which AI doesn't just provide answers. It initiates action. So that action might be go to a website. It might be learn, gather some data from a website, bring it back, take another action. Those actions are done by CPEs. And so as AI gets better and better at doing things, at making instructions, calls for things to get done, we're using more and more CPUs, right? And that is driving up the consumption of CPUs and therefore the demand for CPUs. And so this sort of huge push for more CPUs is being driven by AI On machines like ours and GPUs doing agentic work and asking the, the, the CPUs to take an action, to go to a website, to order a burrito, to find a piece of information, to pull it from storage to all that work is being done by the CPUs.

AI assessment note: “huge push for more CPUs is being driven by AI”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Great. Let's talk about that over in that deal since it's such like a major historical milestone record making. Um, so it's, it's providing up to with seven and 50 megawatts, which is, which is interesting by the way, as a, as a metric because we're in a chip provider, but this is power. So is that shorthand for?

A It's a shorthand. I mean, it turns out right now, and we didn't talk about this because there's Sort of in the adjacent supply chain. We, we went through the shortage of, of memories. We went through the shortage of a process called COOS, three nanometer capacity. The, the other limitation in our industry right now is data center availability. And, I mean, Uh, that is a limiting factor for everybody. And that's why Anthropic did a huge and sort of very expensive deal with Elon for data center capacity. Um, uh, our deal with them, with, uh, with OpenAI was because data center capacity is a limiting constraint, measured the way data centers are measured in, in megawatts. The deal is 760 megawatts, 250 megawatts in 26, uh, on a multi-year lease. An additional 250 megawatts in 27, on a multi-year lease, and an additional in 28, a multi-year lease.

AI assessment note: “It's a shorthand. I mean, it turns out right now”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Um, is there a role for, uh, Local AI and local chips. NVIDIA had some announcement around just building chips for Windows computers. Is it for inference? Is that something that you guys, uh, think should be part of the multi-cylical ecosystem?

A Yeah, I think, um, if you look at the way the ecosystem for apps and cloud emerged, uh, everything you can do, you should do on your phone or your laptop. But the, the, the ability to get real processing power to a phone or to a laptop is constrained because they're generally working off a battery. And so you, you want to do as much as you can, as, as close to the data as you can. But the truth is, is in most situations, for real compute, you have to go to the cloud, to the data center. And that's exactly the way it's going to be with AI. We're going to do a lot of work Uh, uh, on the cell phones, uh, on a laptop, but for the big work, you're going to go to, to the data center. And that's where our focus is. Our focus is, is data center compute free.

AI assessment note: “for the big work, you're going to go to, to the data center”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. And to the general chip versus pistolized chip, uh, it was actually interesting that you guys in your city celebrated when Grok was, uh, acquired. Was that, uh, what was that? Was it a recognition by NVIDIA that your vision was right all along?

A Yeah. I, I think, uh, One of those ideas sort of most durable modes was the perception that the GPU could do everything, and it was all you needed for AI. And the acquisition of Grok for twenty billion dollars, and the structure, and the speed with which they chose to do it, made clear to everyone that that wasn't true. That, uh, the GP architecture couldn't do, could not do fast inference, and that this market was large and growing quickly. And we were the fastest at it, and the largest, and, you know, our sales were more than 10 times the Grox, and they paid twenty billion dollars for the number two collector. So that was a good day. That was a great day.

AI assessment note: “made clear to everyone that that wasn't true”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q comes from the big labs, which, uh, you know, themselves are financed by, uh, venture capital, private equity, edge funds, where you call it, you know, sovereign investors. Um, is, is there any concern that, you know, demand for chips comes from the labs, which Are maybe artificially financed and that the, you know, if you had a true circle of deals here and there, then you sense any infrigility?

A Um, I, I, that's not exactly our experience. I mean, obviously we have, uh, enormous demand from pull open in. But we have huge, you know, dozens of other customers who, who are trying to place very big orders. And historically bubbles were when supply got out ahead of demand, right? When, uh, in the nineties, we built out a telco infrastructure, right? We built out fiber years before it was Going to be used and it took six or eight years and it all got used, but it was a sort of, if you build it, they will come mentality. Whereas what's different about AI right now is we're all trying to catch up. Um, we're trying to build data centers faster. We're trying to increase our, our demand, our, our, our supply chains for demand that's already here today. And only a very small portion of the world are using AI anywhere close to its potential. And we're already sort of overwhelmed with compute. We're, um, overwhelmed with, uh, the demand for, for memory, which was a real weakness in the GPUs. It's not a problem we face. And the ecosystem, there are constraints left and right, and that, that doesn't feel like a bubble.

AI assessment note: “that's not exactly our experience”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Yeah. There's even more specialized ones, uh, right? Like, uh, developed specifically for transformers. Those are called ASICs as well. Can you maybe define what that term means?

A An ASIC is an application-specific integrated circuit, and it's, um, uh, it's a word that now has a wide range of meaning. It, it means that you have made a series of choices away from general towards a narrower class of, uh, problem solving. And that you've made some decisions in your architecture that make it much better at some things and much, much worse than others. Right. And that's, uh, choices that are made across the spectrum. So the TPU has made some choices like that. It can't do graphics. It is very good at matrix multiply. It's not very good at a collection of other things that we do in, in mathematics. Um, uh, same for the GPU. We, we've all made different choices. I was right now in production. There are, uh, in RIDI and AMD, there are a collection, there's a TPU for Google, Traniums. Uh, just coming up is a Maya part. He's a part from, uh, uh, Microsoft, uh, Cerebrus, and there was one other, uh, Grok that got acquired by, by NVIDIA.

AI assessment note: “An ASIC is an application-specific integrated circuit, and it's”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Yeah. Now why, why so much? What was the cost just to understand how much businesses?

A Because what everybody thought was hard, we solved quickly. And what nobody else knew about, because they'd never actually done it, turned out to be really hard. You know, imagine, I tell people, imagine the first group that was going to climb Everest. And they're at base camp, and they're having tea with a group that just failed, and the group that just failed says halfway up, there's this part, it's unbelievably hard, we couldn't do it. Ok. Your team climbs up, makes it all the way to the top, comes back, They're having tea again, and the team that made it leans over the team that hadn't made it, said that part in the middle, that wasn't the hard part, because nobody had gotten past it. Nobody had gotten past certain things, so they didn't even know what to be afraid of. We now know. And it was something, a step called packaging, and that's how you affix a wafer to a motherboard, how you deliver power to it, and how you call it. And, uh, nobody had done it before. And over that 18 month period, we became the best in the world at it from approximately zero. And we did that by sailing again and again and using good engineering methodology and doing a failure analysis, every single failure. So we failed differently again and again and again and again. And we told our board, you know, we met with our board over six weeks and yeah, this is the strategy. Nobody's ever done this bef…

AI assessment note: “Because what everybody thought was hard, we solved quickly.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Um, and, um, so going back to supply chain, uh, do you need to think about, uh, on-shoring, diversification?

A So, it's very hard to diversify away from TSX. Chips are so hard, and you actually, when you design a chip, part of the design is for the rules of that factory. Right? So you can't take your design from TSMC and go to somebody else because a huge amount of the work has been to be sure your design is within their rules. And so even in, I think, only with one or two exceptions in history, uh, each chip generation goes to one fat. So that, that's, we're going to be with TSMC for our next generation as well. I think somehow Uh, we have, uh, a supply chain that is built in many parts, but we bring the chips back from TSMC to the U.S., we package in the U.S., and we assemble in the U.S. We do our manufacturing in the U.S., and then we ship from the U.S. I think when you're growing as fast as we are, right, the overall range of garden variety supply chain challenges, You know, a vendor screws up a batch. It gets stuck in customs. The number of, of ways that, that things can go wrong in the supply chain is unbelievable. But we manage these every day, and we're increasing our manufacturing through quite exponentially, and so it's really, uh, that part of the business in Cromer.

AI assessment note: “it's very hard to diversify away from TSMC”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Where does a multimodality, uh, for the larger models, uh, falling in your world?

A We just announced, uh, uh, sort of that we were fastest in the world on, on one of Google's multimodal models. I, I think the truth is, is that there's very little text that doesn't have charts and graphs, right? You, you, you must to understand text, uh, be able to Understand, uh, illustrations, graphs, uh, charts, and so that, that's sort of the, the first and easiest part, and then you ought to be able to create both things, and then you ought to understand images, and, uh, I, I think the, the new models are very, very good at that. Obviously, what follows that Is video, because a video is just a collection of images. Um, but that takes an enormous amount of compute right now. And that's one of the reasons it's been sort of set aside by the leading labs. So unbelievably computation intensive.

AI assessment note: “we were fastest in the world on, on one of Google's multimodal models.”

Redirected raw tape D 2 · C 5 · P 4 · Cm 4 3.70

Q So in Chengdu, uh, pre-training, I mean, post-training was RL, some pre-training for the other labs. Um, GPUs are still better than you, for what job?

A So GPUs in training, uh, have Some challenges that have been solved by a very narrow selections of the community. GPU is a very small chip, and the calculations that we need to do in training are very large. And one of the most complicated parts of training is the breaking up the calculations and spreading them apart on multiple GPUs. And that's called distributed compute. That has historically been the domain of the super compute world and is very difficult, not just because cracking a problem and having lots of others work on it is hard, but they have to constantly share information. And that sharing information is why they needed to buy Mellanox. Right. Is they needed to control a fabric over which all is sharing in order to get an answer would happen. That breaking up the, a big matrix multiply, a big calculation is called running tensor model parallel. You are breaking up the tensor and spreading it apart. And, and the best labs in the world are good at that. But nobody else is. When we run training, we don't have to run that way. We run what's called model data parallel. And data parallel is very simple. And so it allows teams who are good and very good to quickly test ideas in training. And so we are easier to use and faster, um, because we allow them to use a technique, which is much simple.

AI assessment note: “GPUs in training, uh, have Some challenges that have been solved”

Answered raw tape D 4 · C 3 · P 3 · Cm 3 3.30

Q Was that a, they had a prepared mind or they weren't just exceptionally fast on their feet?

A Um, first the salesperson had gathered the decision rankers. Second, uh, we, our, our proposal was sort of really good at allowing them to use what they were good at. And it didn't require them to change a huge amount, but it did require them to make real changes. And I think they saw this as sufficiently bold that they would learn as they did it. And they also knew that AI was better on big chips. And so the combination of fair mind, the willingness to take some risk, bold thinking from a very large company. Fascinating. It is fascinating. I mean, that's how big companies went. Is that right? Um, and how rare is that? It was extraordinary.

AI assessment note: “combination of fair mind, the willingness to take some risk, bold thinking”

Answered raw tape D 3 · C 3 · P 3 · Cm 3 3.00

Q Yeah. So you started in 2016. Um, you had a fire company that you sold to your AMD. What were some of the lessons you learned there that you took into Cerebris?

A I think the lessons are large and many. I think what, what are the things, you know, I guess experiences is another name for having made mistakes and learning from them, right? I, I think around You know, we, we have as a team built dozens of chips over the past 20 o'clock years, and the returns to experience in chip making is in Argus. And we built a, a different type of computer at, at C-Micro, a type of computer optimized for low power and optimized for a workload that was very different than AI for something like web browsing. And, uh, but the, the fundamental underpinnings, the questions you ask as a computer architect are always to say, well, what can I do to, to, to make this work faster? Right. And is there enough of it to make it worthwhile? These are the two questions we ask. Um, should we build a part for it? Should we, what could we do to build a chip optimized for AI? And will there be enough AI so that you can build business around it? Those are the questions we asked in 2016. And, you know, the, the flip side of that was only Wouldn't it be a surprise if the GPU, which had been optimized for graphics for 20 years, had been pushing pixels to a monitor, was suddenly good at a new world? Wouldn't that be serendipitous? And we came to believe that it wasn't the right architecture for it, it was just better than the CPU. Right. And that we could build an architecture …

AI assessment note: “the questions you ask as a computer architect are always to say”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.