The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Dylan Patel no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q uh, just a general question about, you know, I'm a fellow writer on, on Substack. You are obviously managing your, your consulting business while you're also publishing these amazing posts. Uh, how do you, what's your writing process? How do you Source info. Like, when you sit down and go, like, here's the theme for the week. Do you, do you have a pipeline coming out? Just anything you describe.

A So, um, I'm thankful for my, uh, you know, my teammates, because they are actually awesome, like, uh, and they're much more, um, you know, directed, focused to working on one thing, you know, or not one thing, but a number of things, right, like, you know, someone who's this expert on X and Y and Z and the semiconductor supply chain, so that really helps with the, the, that side of the business. I, most of the times, only write when I'm very excited, or, you know, it's like, hey, like, we should work on this, and we should write about this, so, like, You know, one of the most recent posts we did was we explained the manufacturing process for three D NAND, uh, you know, flash storage, uh, gate all around transistors and, um, three D DRAM and all this sort of stuff. Cause there's a company in Japan that's going public, uh, Kokosai Electric, right? And it's like, okay, well, we should do a post about this and we should explain this. But like, it's like, okay, we, we, you know, and so Myron, um, he, he did all that work, Myron, she, and, and most of the work and, and awesome. But like, usually it's like, there's a few like very long in-depth back burner type things, right? Like that, Took a long time. Took, you know, over a month of, of research, and Myron knows this stuff already really well, right? Like, but also furthermore, it's like, you know, so there's stuff like that that w…

AI assessment note: “I, most of the times, only write when I'm very excited”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q Google, and he did mention that The goal of Google is, like, make TPUs go fast with TensorFlow, but then you also had a post about PyTorch kind of stealing the, the thunder, so to speak. Uh, how do you see that changing? If, like, now that a lot of the compute will be TPU-based, um, and Google wants to offer some of that to the, to the public, too?

A I mean, Google internally, and I think, you know, is, is obviously on JAX and XLA and all that kind of stuff, right? But, uh, externally, like, they've done a lot, a really good job. Like, I mean, I wouldn't, you know, I wouldn't say, like, TPUs through PyTorch XLA is amazing, but it's, it's not bad, right? Like some of the numbers they've, they've shown, some of the, you know, code they've shown for TPUv-V-E, which is not the TPUv-V-V that I was referring to, which is, uh, in, in the sort of the post, the GPU poor post I was referring to, but TPUv-V-V-E is, like, uh, the new one, but it's mostly, mostly an inference chip. It's a small chip. It's a little bit, it's about half the size of a TPUv-V-V. That chip, you know, you can get very good performance on, like, Of, of Lama, 70 B inference, right? Like, you know, very, very good performance. So, like, when you're using PyTorch and XLA. Now, of course, you're gonna get better if you go Jaxx, XLA, but, uh, I think Google is doing a really good job after the restructuring, um, of focusing on external customers, too, right? Like, hey, like, TPV-V, well, probably won't focus too much on TPV-V for everyone externally, but V-V, well, we're, we're also building a million of those, right? Hey, a lot of companies are using them, right? Or will be using them because it's gonna be, An incredibly cheap form of compute. The world of, like, …

AI assessment note: “I wouldn't say, like, TPUs through PyTorch XLA is amazing, but it's not bad”

Redirected raw tape D 2 · C 4 · P 4 · Cm 4 3.40

Q Yeah, excellent. Um, by the way, I like to quantify things when you say make things over, like, uh, is there a target range of, like, MFU that you typically talk about?

A Yeah, there's, there's sort of two metrics that I like to think about a lot, right? So in training, everyone just talks about MFU, right? But then on inference, right, which I think is, you know, one, LLM inference will be bigger than training, or multimodal, whatever, blah, blah, blah inference will be bigger than training, you know, probably next year, in fact, um, at least in terms of GPUs deployed, The other thing is, like, you know, what, what's the bottleneck when you're running these models? So, like, the simple, stupid way to look at it is training is, you know, there's six flops, uh, floating point operations you have to do for every byte you read in, right? Every, every parameter you read in. So if it's fp-eight, then it's a byte. If it's fp-sixteen, it's two bytes, whatever, right? Um, on training. But on inference side, the ratio is completely different. It's two to one, right? There's two flops per parameter that you read in, and parameters, maybe one byte, right? Because it's fp-eight or int-eight. Right, eight bits, bytes. But then when you look at the GPUs, right, the GPUs are a very, very different ratio. The H-One-Hundred has 3.35 terabytes a second of memory bandwidth, and it has a thousand teraflops of FP-sixteen, B-foot-sixteen. Right, so that ratio is like, I'm sorry, I'm going to butcher the math here and people are going to think I'm dumb, but two 56 to …

AI assessment note: “in training, everyone just talks about MFU, right? But then on inference”

Answered raw tape D 4 · C 3 · P 3 · Cm 3 3.30

Q One, one point that you, Didn't touch on it, and Taiwan Company is famously very chatty about the fruit company. Um, should we take Apple seriously at all in this game, or they're just in a difference? Well, no, it's good there.

A I think, I think, you know, just from my view of Apple, I, I, I don't personally use Apple products, but every, I mean, like, my, my mom, I buy her a new iPhone every year, right? Just to be clear, right? Like, yeah, no, mom, you, you know, new Apple watch every couple years, right? Like, of course, right? Um, so I, I respect their products, but like, I don't think Apple will ever release a model that you can get to say, you know, really bad things, right? Or, you know, racist things or whatever, right? I don't think they can ever do that, but like, you know, frankly, like, I'm sure, I'm sure OpenAI releasing, you know, 3.5 and four has had people like, you know, break jail breaks for the, you know, kind of, uh, old terminology from iPhone jailbreak the model and get it to do bad things, right? Teach me how to make anthrax, right? Like, or like say these like hateful things, like race rank, rank the races of the world, right? Like, you know, like crazy. I mean, I've seen it on Twitter. I've seen all these three of these things, right?

AI assessment note: “I don't think Apple will ever release a model that you can get to say”

Answered raw tape D 4 · C 3 · P 3 · Cm 3 3.30

Q know you had some safety hot takes, and I think it's like an interesting dynamic because, you know, Anthropic came out of OpenAI, and then you can kind of make the case that like by having more labs, if you're really worried about safety, you're like accelerating the unsafe because you have more pull-ups and more compute. Uh, yeah, what, what's your thought on like this whole, this whole space?

A So obviously I think safety is probably important, but like, I mean, it is important, right? Like, I mean, I've read sci-fi normally, right? It's clearly important, right? Like, um, I could easily see how an LLM could, I, I wrote about this the other day, but it was like, it's like, hey, like, if you just look at the demographics across the world, there are like, there's like, 30 to fifty million more men than there are women, and they will never get married, ostensibly, obviously, on population-level dynamics, You know, LGBTQ, all that stuff happens, and it's great, but like, you know, like, there are 30 to fifty million more men across the world. They'll always be single. Why can't an LLM, like, like, radicalize them, right? By being its AI girlfriend, and then all of a sudden, like, inciting, like, you know, and also, yeah, I don't know, there's like all sorts of stuff like that can happen, of course, right? Or like, you know, teach some person to create, manufacture what they thought was a good thing, and it ends up wiping out humanity. Like, all these sorts of stuff can happen. But, you know, at the end of the day, I think, Security through obscurity doesn't work, right? So that's, that's the approach that the labs take. I truly, I truly do believe it, right? Like, you know, they, they're very open internally, at least, at least anthropic and open AI are. I know Google's a…

AI assessment note: “Security through obscurity doesn't work, right? So that's, that's the approach”

Partly raw tape D 3 · C 3 · P 4 · Cm 3 3.25

Q your, yeah, your Google Gemini hits the world, um, blog post. One, did you know that this thing was gonna blow up so much? Sam Allman even tweeted about it. He said, incredible Google got the semi-analysis guide to publish their internal marketing recruiting chart. Uh, and yeah, it's helped people Who are the GPU boys? Who are the GPU rich? Like, what's this framework that they should think about?

A It's, um, you know, some of this work we've been doing for a while is just on infrastructure and like, hey, like when something happens, I think it's like a, you know, sort of competitive advantage of our firm, right? Me, myself, and my colleagues is like, we go from software all the way through to like low-level manufacturing, and it's like, who, you know, oh, Google's actually ramping up TPU production massively, right? Um, and like, I think people in AI would be like, well, duh. But like, okay, like who, who has the capability of figuring out the number? Well, one, you can just get Google to tell you, but they don't, they won't tell you, right? That's like a very closely guarded secret, and most people that work at Google DeepMind don't even know that number, right? Um, two, you go through the supply chain and see what they've placed in orders, right? Um, but then you, you know, three is sort of like, well, who's actually winning from this, right? Like, hey, oh, Celestica's building these boxes. Wow. Oh, interesting. Okay. Um, you know, uh, this company's involved in testing for them. Oh, okay. Oh, this company's providing design IP to them. Okay. Okay. Like, that's like, like, You know, very valuable in a monetary sense. Um, but, you know, you have to understand the whole technology stack, but on the flip side, right, is like, Well, why is Google building all these? Uh, wha…

AI assessment note: “people just brag about how many GPs they have. Like, it's happened”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.