The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Steeve Morin argument clarity score 4.1/5 from 45 exchanges on raw tape · average scores: directness 4.7 · coherence 4 · precision 3.8 · compression 3.5 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
45exchanges match
45on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Why do you think that's a bubble that will blow some time? Why is that not legitimate?

A Because, um, it was built on the A 100, uh, I would say financial model, which was At Generation Zero, we do training. Uh, but when it's last generation, we do inference and it worked beautifully, right? Uh, for a 100, then H 100 comes along and inference is it's worth five times the price. Uh, and it may be runs twice, uh, in terms of performance on inference. That is on training. It's a lot better, but on inference, it's like maybe twice as fast when it actually, when it came out, it ran at the same speed than the a 100. So there's a money gap. That's going to have to, you know, be bridge sometime. Right. And the, the part that worries me is that I see, you know, amortization plans on like, you know, six, seven years with the GPUs at the collateral. And I'm like, well, I'm not sure how it's going to work because, you know, there were, at least when they came out, there were five times the price and they're just two times, you know, faster. So something's going to, something has got to give.

AI assessment note: “So there's a money gap. That's going to have to, you know, be bridge sometime.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What is forcing the price of a Cerebris to be so high? And then you heard Jonathan at Grok on the show say that they're 80% cheaper than NVIDIA.

A Ah, it's, so there's this trick. Cause here's the thing, uh, there's no magic. This little trick is called SRAM. Uh, SRAM is memory on the chip directly. So that is very, very fast memory. But here's the problem with SRAM, is that SRAM consumes, you know, surface on the chip, right, which makes it a bigger chip. Which, you know, is very hard in terms of yield, right? Because there's the chances of like a problems are higher and so on. SRAM is, I would say, very, very, very fast memory, which gives you a lot of advantage when you do very, very high inference, very, very, sorry, very high speed inference. But it's terribly expensive. And if you look at, for instance, Grok, they have on their generation, this generation, they have 230 megabytes of SRAM per chip. A seven TB model is, you know, uh, what's called BF-sixteen is a 140 gigabytes. So you do the math, right? Uh, Cerebras has a 44 gigabytes of SRAM into what they're called their wafer scale engine, which is a chip the size of a wafer. I mean, most likely it's interconnected, but it's huge, right? And it's, it has to be watercooled. They have Uh, copper, you know, I would say needles that touch the chip. It's crazy stuff. Um, uh, very, very impressive technology, mind you, but very, very expensive. So my bet is, I think there will be, you know, chips on the market that do that at, at much lower price. Um, and there's two co…

AI assessment note: “SRAM consumes, you know, surface on the chip, right, which makes it a bigger chip.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Can I ask you then, if we think about sitting between any model and any provider there in terms of AMD and video, do you think then we will be existing in a world where people are using multiple models simultaneously and that is concurrently running?

A Yes. Um, You, you actually can see it. It's been happening for a while. Models now are not the right abstractions. At least if you look at close source model, they're not really models. They're more like backend, right? Uh, and there are a lot of tricks that you feel like you're talking to one model, but ultimately you're talking to a constellation, an assembly of backends that produces, you know, a response. Uh, probably the number one, you know, I would say obvious thing would be that if you ask a model to generate an image, Then it will, you know, switch to a diffusion model, right? Not an LLM. So, and there's many, many more tricks. The Turbo models and OpenAI do that. There's a lot of tricks. So definitely, uh, models in the sense of, you know, getting, you know, weights and running them is something that is ultimately going away because, uh, you know, in favor of like full blown backends, right? You feel like you're talking to a model, but ultimately you're talking to an API. The thing is, that API will be running locally, right? Locally, I mean, in your own, you know, cloud, you know, instances, and so on.

AI assessment note: “Yes. Um, You, you actually can see it. It's been happening for a while.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Totally get you there. Okay. So when we move to actually inference and training, I think everyone's focused so much on training. I'd love to understand what are the fundamental differences in infrastructure needs when we think about training versus inference?

A These two obey fundamentally different, I would say, tectonic forces, if you will. So, in training, more is better. You want more of everything, essentially. And you, and the recipe for success is the speed of iteration. You change stuff, you see how it works, and you do it again, hopefully it converges. And it's like, you know, changing the wheel of a moving car, so to speak. Some training runs are, that is. Uh, so that is training. On inference, this is a complete reverse. Less is better. You want less headaches. You don't want to be woken up at night because inference is production, right? You could say that training is research and inference is production and it's fundamentally different. In terms of infra, probably the number one thing that is the number one difference between these two is the need for interconnect. So if you do, you know, production, you, if you can avoid to have interconnect between, you know, let's say a cluster of GPUs, of course you will not go, you, you will, you know, avoid that, right? If you can. And this is why models have the sizes they have is so that people can run them without the need to connect multiple machines together. It's, it's very constraining in terms of the environment. So That is probably the fundamental difference, the need for interconnect. And number two is, ultimately, do you really care about what your model is running on as …

AI assessment note: “the number one difference between these two is the need for interconnect”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Totally get you there. Okay. So when we move to actually inference and training, I think everyone's focused so much on training. I'd love to understand what are the fundamental differences in infrastructure needs when we think about training versus inference?

A These two obey fundamentally different, I would say, tectonic forces, if you will. So, in training, more is better. You want more of everything, essentially. And you, and the recipe for success is the speed of iteration. You change stuff, you see how it works, and you do it again, hopefully it converges. And it's like, you know, changing the wheel of a moving car, so to speak. Some training runs are, that is. Uh, so that is training. On inference, this is a complete reverse. Less is better. You want less headaches. You don't want to be woken up at night because inference is production, right? You could say that training is research and inference is production and it's fundamentally different. In terms of infra, probably the number one thing that is the number one difference between these two is the need for interconnect. So if you do, you know, production, you, if you can avoid to have interconnect between, you know, let's say a cluster of GPUs, of course you will not go, you, you will, you know, avoid that, right? If you can. And this is why models have the sizes they have is so that people can run them without the need to connect multiple machines together. It's, it's very constraining in terms of the environment. So That is probably the fundamental difference, the need for interconnect. And number two is, ultimately, do you really care about what your model is running on as …

AI assessment note: “number one difference between these two is the need for interconnect.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What's one piece of advice you'd give to AI startups navigating the changing landscape of training, inference, and hardware?

A Probably the number one thing I would say is, um, do not resell compute if you can. A lot of, you know, AI startups, uh, That are building on top of AI are trying to make a margin, you know, on top of a very big cake. And ultimately what they sell is compute. If you look at the dollar of spend, um, you know, for one dollar of spend, maybe 98% of it goes to somebody else's margin. So if you, if you do AI, uh, as much as you can try to verticalize on, on the product, but not on the compute. So if you are, you know, if your business model implies, you know, buying a lot of tokens, it's a very hard circle to square to, you know, put that into 20 dollars, right? Um, a month. So, you know, I always say like, please, you know, look at it from that angle. And if you can try and avoid it.

AI assessment note: “Probably the number one thing I would say is, um, do not resell compute”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q How does that change the situation? It makes it much more efficient, but what does that actually mean in reality?

A It means you get maybe not SRAM level performance, but you get very, a lot faster performance, uh, in terms of compute. So, and if you translate that to LLMs, let's say you get much, much higher tokens per second in a single stream, which is exactly what you want, uh, when you go into reasoning. Right. You want your model to maybe think, let's say for like half a second and then boom. Right. You don't want to wait 50 seconds and, you know, context switch to some other thing, which is the problem everybody has today, mind you. So yeah, I think inference will be major, or at least this is my thesis is that it will be pushed. The compute landscape will be pushed to change because of these two constraints. I know I'm working on it.

AI assessment note: “you get much, much higher tokens per second in a single stream”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q At what point do we stop and say, Hey, there is a lot of wastage and we could do more better. How far away are we from that?

A I think until somebody does it. Deep Seek was a good, you know, wake up call, right? Uh, suddenly efficiency is in, right? Um, that's number one. And number two is until there's a new architecture that comes out and changes the game. So in the case of LLMs, for instance, you have these, what's called non transformer models that changes fundamentally the compute requirements. So that might be a frontier that, you know, completely obsoletes the transformers. Right. And if the trans, sorry, the transformers are the, um, I would say the building block by which current model work. Right. So the way they work is that for each, you know, token or syllable, if you will, the model will look at everything behind it. So you can see that as you add more text, you have more work to do. Right. Uh, so there are these new architectures, uh, um, that do not require this, uh, that might Change, you know, these things, uh, and probably shift the amount of compute needed to do training or to do inference. And then there's the new thing, which is, uh, Jan's, um, uh, thesis, which is the, um, uh, word model, right? As in LLMs are at that end, what we need is something that understands the world fundamentally, and this is the, it's, uh, uh, it's JEPA thesis, it's called. I'm very bullish on this, but it's, it's very frontier.

AI assessment note: “I think until somebody does it. Deep Seek was a good, you know, wake up call”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Why are you bullish on it? And why is it so frontier?

A Cause it's Sian Lecun. It's hard to. Um, he's no bullshit, right? So, um, he explained to me how it worked and I was blown away. Simple as this. Uh, but it makes a lot of sense. Um, you know, we are creeped out because the machine talks back to us. That's it. But it's not a new thing, right? It used to, you know, this is not new technology when it came out. Well, like when it, when it exploded, it wasn't new technology. But suddenly it was talking back, and that freaked us out. Uh, and we got crazy on it, right? But language is one form of communication, uh, but it, it is ultimately a very narrow, uh, window into, you know, the world. Uh, we use it to describe the world arguably with some loss, right? Um, and so there's this, um, the JEPA approach is, long story short, is that You have essentially two things you want to do and you try and minimize the energy to do them. But, and, and from this understanding emerges, physics emerges and et cetera, because you're trying to minimize the amount of energy to go from one state to the other. And that actually makes sense. Like if you try and, you know, pick this air pod, you know, case, I'm not going to go round trip around the block to get it right. I just get it. And in my brain, it's wired to Just do the, the thing. If I If I go and, you know, talk to myself out loud, you know, put the, the, the hand down, move to the left and what…

AI assessment note: “the JEPA approach is, long story short, is that You have essentially two things”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q What's one piece of advice you'd give to AI startups navigating the changing landscape of training, inference, and hardware?

A Probably the number one thing I would say is, um, do not resell compute if you can. A lot of, you know, AI startups, uh, That are building on top of AI are trying to make a margin, you know, on top of a very big cake. And ultimately what they sell is compute. If you look at the dollar of spend, um, you know, for one dollar of spend, maybe 98% of it goes to somebody else's margin. So if you, if you do AI, uh, as much as you can try to verticalize on, on the product, but not on the compute. So if you are, you know, if your business model implies, you know, buying a lot of tokens, it's a very hard circle to square to, you know, put that into 20 dollars, right? Um, a month. So, you know, I always say like, please, you know, look at it from that angle. And if you can try and avoid it.

AI assessment note: “do not resell compute if you can”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q When we think about that, it's rational if you think that efficiency and scaling laws continue to continue to place such emphasis on it. How do you think about model scaling and scaling laws coming into place?

A There's like a brute force approach to this. It is a very American approach more and more and more. But the thing is, you look at, for instance, the, uh, Uh, the XAI cluster. It's not a 100,000 GPUs. It is four times 25,000. So, you see, you're starting, you know, see some, because InfiniBand, and in the case, uh, uh, Rocky, which is, um, anyways, the technology they use to bridge their, uh, GPUs together. They, you have upper bounds, right? At, at some point, you're fighting physics. So you can push, it's like, you know, trying to get to the speed of light. You get, you know, a tiny, you know, as, as you approach it, the, the, the amount of energy you need, you know, is a lot higher and a lot higher and it grows and grows. So there's two approaches to that. One is two, I'm sorry, two, I would say counter to that would be that number one is we still scale, but there's a lot of waste and, uh, excess, you know, spending, Uh, um, on the engineering side, which is the deep seek approach, right? Very successful at that, mind you. They said, yeah, if we do this and this differently, then we get, you know, multiple sometimes, right? So virtually you increase your compute capacity because you're more efficient. And the other approach is Jan's, Jan Lequin's approach, which is, this is not scaling. And at some point you need, we need to look the problem in the face and, you know, Do some…

AI assessment note: “I think you can do more with less.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Can you just explain to us what PyTorch and Cuda are?

A Oh, yes, absolutely. Yeah, yeah. PyTorch is the ML framework that people use to build, actually train models, right? You can do inference with it, but by far the most successful framework for training is PyTorch, and PyTorch was very much built on top of CUDA, which is the Nvidia software, right? Uh, and the way, let's just say the strings of PyTorch make it ultimately, um, very, very bound to CUDA. So of course it runs on, you know, it runs on AMD, it runs on, you know, even Apple and so on, but there was always, you know, the, the tens of little details that not exactly run like, you know, you would expect, and there's work involved, but then also there's supply. Um, so Probably that's the number one thing. The second thing is there's a lot of, um, GPUs on the market and all of them, pretty much all of them are, are Nvidia. The reason being that if you think, you know, in layers and you say, all right, I'm going to buy, let's say GPUs and I'm going to sell them to folks to maybe not even do training, right? Just do inference. Then most likely, if you look at it that way, you'll end up buying Nvidia because Everybody will want to run an Nvidia because nobody knows really how to do whatever and they've trained on Nvidia. So they're like, I can just reuse my code and so on. Um, so there's like this self perpetuating, you know, uh, uh, circle of people just buy Nvidia because the…

AI assessment note: “PyTorch is the ML framework that people use to build... built on top of CUDA”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q So when I had Jonathan on, he was like, actually Nvidia have such a stronghold because they're one of the only buyers of HBM and that gives them this unique position. Actually, is being a sole buyer of HBM irrelevant if the world needs SRAM instead?

A No, you, you want HBM to, to, to be clear. Uh, no, SRAM, this will, this will not deliver. Uh, it's, it's, it's a dead end in terms of scaling SRAM means scaling the surface mean you get, you know, depreciating problems. It explodes everywhere, right? So, um, you need some SRAM, right? So we, you know, we'll have bigger amounts of SRAM, uh, in, in, into chips. And of course, bigger, what's called external memory into chips. The issue, um, with HBM, uh, is that it's still slow, and yes, maybe Nvidia has a stronghold, and they can prevent you from getting some, so that would be like, I, I call it the Nutella situation, in which, you know, Nutella, they owns 80% of the hazelnuts market, right? So yes, you can do, uh, a competitor, but who will you buy the nuts from, right? Um, so this is about, I mean, this is, you know, a bit above my league, but Uh, there will be a need for HBM. Uh, there will be a need for SRAM. I would say better, more dedicated architecture will be able to deliver these things. Um, and then there's like the next frontier after that, which is called compute in memory. There's two companies that, uh, that I know at least that, that, that are on that market. One is called rain rain dot AI. Um, Sam Altman is one of the investor, you know, There's no surprise. Uh, the other one is called Fractile. Uh, I think it's in the UK, actually. Um, so this is the next front…

AI assessment note: “No, you, you want HBM to, to, to be clear. Uh, no, SRAM”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q With increasing competitiveness within each of those layers, do we not see margin reduction?

A Absolutely. Yes. Yeah. But here's the, here's the problem though. So let's say you are on Google Cloud and you're on TPUs, right? Suddenly, you just remove that 90% chunk on, on, on the, you know, on the spend. The problem is, is that for multiple software reasons, which are, you know, which we are solving, um, at DML is that they're not really, I would say, a commercial success. They are very much successful inside of Google, but not much outside of Google, let's say, right? Amazon, same, is pushing very, very hard for their, you know, Tranium chips. So, My, I would say the future I see is that you use whatever, you know, your provider has, has, because you don't want to pay, you know, 90% outrageous margin, uh, and try to make, you know, a profit out of that.

AI assessment note: “Absolutely. Yes. Yeah. But here's the, here's the problem though.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q To what extent is that correct? Or actually, as Jonathan at Grok said in the show, you know, Nvidia is not meant for inference. Definitely not. And actually, that market won't be won by Nvidia.

A I mean, Technically speaking is right, but realistically speaking, I don't sure, I'm not sure I agree. The thing is these chips are on the market. They're here. I can, you know, out tab on Chrome and get one. That is something that, you know, I don't take lightly. Um, uh, availability that is right. So I think Nvidia is used to stay, uh, at least if not for the H-one hundred, you know, bubble bust. Uh, because these chips are going to be on the market and people will buy them and do inference, uh, with them. Remains to see, you know, the, um, the, the OPEX and the electricity, etc. But that is a complicated question. The thing is, uh, as far as I know, the, the, the, the only chips that are really, you know, um, uh, frontier on that sense are probably TPUs and then the upcoming chips. But the thing is, they're great chips. But they're not on the market. Or like there are outrageous prices, like millions of dollars, you know, to run a model.

AI assessment note: “Technically speaking is right, but realistically speaking, I don't sure, I'm not sure I agree.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Wow. Do you think NVIDIA owns both of those markets in five years time?

A Depends on the supply. I think that there's a shot that they don't. Because here's the thing, you know, even if we take, you know, same amount of, you know, let's imagine we get, we get, we have a new chip from Amazon that is the same amount. Oh, wait, we do. It's called Tranium. Um, and you know, why would I pay 90% margin of Nvidia if I can freely change to Tranium? My whole production is run on, runs on AWS anyways. So that may create Like if you run on the cloud and you're running on Nvidia, you're getting, you know, squeezed out of your money, most likely your investors money, but still, right? Um, so if you're on production on dedicated chips, uh, of course, you know, so maybe through commoditization, but you know, hey, I'm on AWS, I can just click and boom, it runs on AWS's chips. Who cares, right? I just run my model, uh, like I did, you know, uh, two minutes ago.

AI assessment note: “I think that there's a shot that they don't.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Which is why you don't, or are you saying that to get into one of these buy processes, you have to buy so much that it prohibits you?

A It's actually, it's actually both. So the buy-in is very high. So to make it worth it, you have to buy a lot. And if you buy a lot, this is, you know, what every, we talk to all of them. They, they always have the same questions and it's completely understandable. They say, this is great, but who's the customer? Because on the other side, let's take Amazon, for instance, with Tranium. Apple just came and said, Hey, we're going to buy a 100,000 of them. Oh, so you want to buy, you know, 10,000, you feel like the big shot, right? Yeah, but yeah, take, you know, go back to the queue because there's Apple before you, right? So they have to have very high commitments to make it, you know, you cannot be incrementally better. It's very hard and also very hard. I can give you a, I can give you one metric if you want. Um, I know for a fact that being seven times better and whatever, take whatever metric you want. Uh, whether it's spend, whether it's whatever. It's not enough to get people to switch. People will choose nothing over something. Right. I've like, I have stories. So this is a very hard market to enter into because you cannot also compete of incremental gains. Right. It's very hard. Right. So you have to convince a lot of people. Um, maybe you can go the, um, Middle East route in which, you know, they sprinkle everything and they, you know, evaluate everything, but you know, …

AI assessment note: “It's actually, it's actually both. So the buy-in is very high.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q I'm sorry. How does Microsoft buying all of AMD's supply make them not lose money on inference? Just help me understand that.

A Because I can give you like actual numbers. If you're on Eight H-One hundred. You can put two seven TB models on them because of the RAM. That's number one. Number two is, if you go from one GPU to two, you don't get twice the performance. Maybe you get 10% better performance. Yeah, that's the dirty secret nobody talks about. I'm talking inference, right? So, so you go from, let's say, a hundred to a 110 by doubling the amount of GPUs. That is insane. So you'd rather have two by one than one by two. With one machine of eight, eight, eight, 100, you kind of run two seven TBS model. If you do, you know, four GPUs and four GPUs. Right? Um, that's number one. If you run on AMD, well, there's enough memory inside the GPU to run one model per card. So you get, you know, eight GPUs, eight times the throughput, while on the other hand, you get eight GPUs, two, you know, two, maybe two and a half times the throughput. So that is, you know, a Forex right there, right? Just, you know, by virtue of this, of this. Um, and so that is, you know, the, the, the compute part, but if you look at all of these things, there are tremendous amount of, you know, we talked to companies who have chips are coming with Almost 300 gigabytes of memory on it, right? So that is, you know, a model, like, one chip per model. This is the best thing you want, right? Uh, if you're on seven TBs, right? Um, so, whic…

AI assessment note: “If you run on AMD, well, there's enough memory inside the GPU to run one model per card.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q When we think about that, it's rational if you think that efficiency and scaling laws continue to continue to place such emphasis on it. How do you think about model scaling and scaling laws coming into place?

A There's like a brute force approach to this. It is a very American approach more and more and more. But the thing is, you look at, for instance, the, uh, Uh, the XAI cluster. It's not a 100,000 GPUs. It is four times 25,000. So, you see, you're starting, you know, see some, because InfiniBand, and in the case, uh, uh, Rocky, which is, um, anyways, the technology they use to bridge their, uh, GPUs together. They, you have upper bounds, right? At, at some point, you're fighting physics. So you can push, it's like, you know, trying to get to the speed of light. You get, you know, a tiny, you know, as, as you approach it, the, the, the amount of energy you need, you know, is a lot higher and a lot higher and it grows and grows. So there's two approaches to that. One is two, I'm sorry, two, I would say counter to that would be that number one is we still scale, but there's a lot of waste and, uh, excess, you know, spending, Uh, um, on the engineering side, which is the deep seek approach, right? Very successful at that, mind you. They said, yeah, if we do this and this differently, then we get, you know, multiple sometimes, right? So virtually you increase your compute capacity because you're more efficient. And the other approach is Jan's, Jan Lequin's approach, which is, this is not scaling. And at some point you need, we need to look the problem in the face and, you know, Do some…

AI assessment note: “There's like a brute force approach to this... but I'm more of the, of these two approaches.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q So when I had Jonathan on, he was like, actually Nvidia have such a stronghold because they're one of the only buyers of HBM and that gives them this unique position. Actually, is being a sole buyer of HBM irrelevant if the world needs SRAM instead?

A No, you, you want HBM to, to, to be clear. Uh, no, SRAM, this will, this will not deliver. Uh, it's, it's, it's a dead end in terms of scaling SRAM means scaling the surface mean you get, you know, depreciating problems. It explodes everywhere, right? So, um, you need some SRAM, right? So we, you know, we'll have bigger amounts of SRAM, uh, in, in, into chips. And of course, bigger, what's called external memory into chips. The issue, um, with HBM, uh, is that it's still slow, and yes, maybe Nvidia has a stronghold, and they can prevent you from getting some, so that would be like, I, I call it the Nutella situation, in which, you know, Nutella, they owns 80% of the hazelnuts market, right? So yes, you can do, uh, a competitor, but who will you buy the nuts from, right? Um, so this is about, I mean, this is, you know, a bit above my league, but Uh, there will be a need for HBM. Uh, there will be a need for SRAM. I would say better, more dedicated architecture will be able to deliver these things. Um, and then there's like the next frontier after that, which is called compute in memory. There's two companies that, uh, that I know at least that, that, that are on that market. One is called rain rain dot AI. Um, Sam Altman is one of the investor, you know, There's no surprise. Uh, the other one is called Fractile. Uh, I think it's in the UK, actually. Um, so this is the next front…

AI assessment note: “No, you, you want HBM to, to, to be clear.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Can you just explain to us what PyTorch and Cuda are?

A Oh, yes, absolutely. Yeah, yeah. PyTorch is the ML framework that people use to build, actually train models, right? You can do inference with it, but by far the most successful framework for training is PyTorch, and PyTorch was very much built on top of CUDA, which is the Nvidia software, right? Uh, and the way, let's just say the strings of PyTorch make it ultimately, um, very, very bound to CUDA. So of course it runs on, you know, it runs on AMD, it runs on, you know, even Apple and so on, but there was always, you know, the, the tens of little details that not exactly run like, you know, you would expect, and there's work involved, but then also there's supply. Um, so Probably that's the number one thing. The second thing is there's a lot of, um, GPUs on the market and all of them, pretty much all of them are, are Nvidia. The reason being that if you think, you know, in layers and you say, all right, I'm going to buy, let's say GPUs and I'm going to sell them to folks to maybe not even do training, right? Just do inference. Then most likely, if you look at it that way, you'll end up buying Nvidia because Everybody will want to run an Nvidia because nobody knows really how to do whatever and they've trained on Nvidia. So they're like, I can just reuse my code and so on. Um, so there's like this self perpetuating, you know, uh, uh, circle of people just buy Nvidia because the…

AI assessment note: “PyTorch is the ML framework... built on top of CUDA, which is the Nvidia software”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Wow. Do you think NVIDIA owns both of those markets in five years time?

A Depends on the supply. I think that there's a shot that they don't. Because here's the thing, you know, even if we take, you know, same amount of, you know, let's imagine we get, we get, we have a new chip from Amazon that is the same amount. Oh, wait, we do. It's called Tranium. Um, and you know, why would I pay 90% margin of Nvidia if I can freely change to Tranium? My whole production is run on, runs on AWS anyways. So that may create Like if you run on the cloud and you're running on Nvidia, you're getting, you know, squeezed out of your money, most likely your investors money, but still, right? Um, so if you're on production on dedicated chips, uh, of course, you know, so maybe through commoditization, but you know, hey, I'm on AWS, I can just click and boom, it runs on AWS's chips. Who cares, right? I just run my model, uh, like I did, you know, uh, two minutes ago.

AI assessment note: “I think that there's a shot that they don't.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q With that realization, do you think we'll see Nvidia move up stack and also move into the cloud and model?

A They are. They are. They have a product called NIMH that does, you know, sort of does that. They are. The thing with Nvidia is that they spend a lot of energy making you care about stuff you shouldn't care about. And they were very successful. Like who gives a shit about CUDA? I'm sorry, but Uh, I don't want to, I don't want to care about that, right? I want to do my stuff, uh, uh, and NVIDIA got me into saying, hey, you should care about this because there's nothing else on the market. Well, that's not true, but ultimately this is the GPU I have in my machine, so, you know, off I go. If tomorrow that changes, why would I pay 90% margin on my compute? That's insane. This is why I believe it ultimately goes through the software. Uh, because the software, like if my entry, this is my entry point to the ecosystem. So if the software is, um, you know, abstracts away those endiosyncrasies as they do on CPUs, right? Then the, the providers will compete on specs and not on fake modes, uh, or circumstantial modes. So this is where I think, you know, the market is going and Of course, there's, there's the availability problem. There is, you know, if you, you know, piss off Jensen, you might need to kiss the ring, you know, uh, to get back in line. Right. Uh, um, but I mean, ultimately this is, I don't see this as being sustainable.

AI assessment note: “They are. They are. They have a product called NIMH that does”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Do you think it is a meaningful threat to OpenAI and ChatGPT. Bluntly, they still have the consumer loyalty, the consumer brand. To what extent is it actually a long-term threat?

A I'm not sure who is a threat to OpenAI at the moment. Um, here's why. You look at the numbers. I mean, we live in a bubble, right? We, you know, we follow every new episode, the whatever new model, whatever, who said what and so on. But, you know, I go to my mother and I Ask her, you know, do you know ChatGPT? And she says, yes. And do you know, I don't know, I don't want to dunk on anybody, but do you want to know some other model? And she says, what it is? What is it? Right? Even Gemini, right? Like Google, right? So they have a strong brand. They have a strong product. Um, arguably it is, I don't know, like, but there's a balance between the product and the models, honestly. This is Gary from Fluidstack actually, who told me that Is mental model in terms of model providers, ah, they'll be like car makers. Right? There's no win or tickle. Everybody will have their own. Um, because ultimately also human knowledge is, I mean, there's, you know, everybody has everything. Uh, so we're converging. Uh, but I like that analogy. Um, yes, Deep Seek made a very good, you know, made waves, but it was, it was, you know, waves that were amplified by the media and the narrative and, and the drama, right?

AI assessment note: “I'm not sure who is a threat to OpenAI at the moment.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Why are you bullish on it? And why is it so frontier?

A Cause it's Sian Lecun. It's hard to. Um, he's no bullshit, right? So, um, he explained to me how it worked and I was blown away. Simple as this. Uh, but it makes a lot of sense. Um, you know, we are creeped out because the machine talks back to us. That's it. But it's not a new thing, right? It used to, you know, this is not new technology when it came out. Well, like when it, when it exploded, it wasn't new technology. But suddenly it was talking back, and that freaked us out. Uh, and we got crazy on it, right? But language is one form of communication, uh, but it, it is ultimately a very narrow, uh, window into, you know, the world. Uh, we use it to describe the world arguably with some loss, right? Um, and so there's this, um, the JEPA approach is, long story short, is that You have essentially two things you want to do and you try and minimize the energy to do them. But, and, and from this understanding emerges, physics emerges and et cetera, because you're trying to minimize the amount of energy to go from one state to the other. And that actually makes sense. Like if you try and, you know, pick this air pod, you know, case, I'm not going to go round trip around the block to get it right. I just get it. And in my brain, it's wired to Just do the, the thing. If I If I go and, you know, talk to myself out loud, you know, put the, the, the hand down, move to the left and what…

AI assessment note: “he explained to me how it worked and I was blown away.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q before you said about AMD and I said, Hey, you know, I bought Nvidia and I bought AMD and Nvidia. Thanks Jensen. I've made a ton of money. And AMD, I think I'm up one percent, ah, versus the 20% gain I've had on video. My question, you said that AMD basically sold everything to Microsoft and meta and had a GTM problem. Can you just unpack that for me?

A So all, I would say chip makers have a GTM problem. Uh, it's all of them, whether, you know, it's Google, whether it's AMD, whether it's TenStorn. Right. The problem is, is that there's, I would say, Probably two fundamental problems. The number one is if you maintaining multiple stacks today is very, very, very hard. So you don't. So let's say I buy, you know, AMD, I want to buy AMD, right? That means I'm going to abandon Nvidia. Oh, crap. You know, I have a six year amortization plan on that. Oh, man, what do I do? So do I need to support both stacks? Unclear. Maybe until AMD tells me, hey, you know, you have, I don't know, let's say a thousand Nvidia GPUs. Um, you're about to buy a 100,000 of AMD. I mean, come on. Right. And I'm like, okay, that is, you know, makes it worth my while. Right. Um, so, but that is ultimately the fundamental problem is that the steps are very high. Right. I need to have a lot of incentives to buy into that ecosystem. So I need to buy a lot of them. Right? So if you're AMD, that is already a problem, right? Uh, but then Microsoft comes along and buys it all, makes, by the way, OpenAI, or at least on the inference side, puts OpenAI in the green because of the efficiency gains.

AI assessment note: “maintaining multiple stacks today is very, very, very hard. So you don't.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q What does everyone think they know about inference that they actually don't, or what does everyone get wrong about inference?

A Probably Not a lot of people are accustomed to what it entails to, to run production. So that inference is production and production is hard. Somebody has to wake up at night. And I used to be that guy, right? I don't want to do it again. Um, so production is hard. Thankfully, we have a lot of, uh, uh, software nowadays to do that a lot better, but there's not a lot of reuse because the AI field at least Is not really accustomed to that yet. It's changing. Uh, but you know, the discussions I had, you know, a year ago, and the discussions I had, you know, today are not the same. They're going to the right direction, but they're not there exactly yet. So probably that would be the number one thing. That is only, you know, um, uh, training code running only forward pass, right? This is not what it is.

AI assessment note: “training code running only forward pass, right? This is not what it is.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Do you think it is a meaningful threat to OpenAI and ChatGPT. Bluntly, they still have the consumer loyalty, the consumer brand. To what extent is it actually a long-term threat?

A I'm not sure who is a threat to OpenAI at the moment. Um, here's why. You look at the numbers. I mean, we live in a bubble, right? We, you know, we follow every new episode, the whatever new model, whatever, who said what and so on. But, you know, I go to my mother and I Ask her, you know, do you know ChatGPT? And she says, yes. And do you know, I don't know, I don't want to dunk on anybody, but do you want to know some other model? And she says, what it is? What is it? Right? Even Gemini, right? Like Google, right? So they have a strong brand. They have a strong product. Um, arguably it is, I don't know, like, but there's a balance between the product and the models, honestly. This is Gary from Fluidstack actually, who told me that Is mental model in terms of model providers, ah, they'll be like car makers. Right? There's no win or tickle. Everybody will have their own. Um, because ultimately also human knowledge is, I mean, there's, you know, everybody has everything. Uh, so we're converging. Uh, but I like that analogy. Um, yes, Deep Seek made a very good, you know, made waves, but it was, it was, you know, waves that were amplified by the media and the narrative and, and the drama, right?

AI assessment note: “I'm not sure who is a threat to OpenAI at the moment.”

Answered raw tape D 5 · C 4 · P 3 · Cm 4 4.05

Q Can I just ask, what's latent space reasoning?

A So the way models reason today is the reason it was in, in, in tokens. So it's as if, if you think to yourself, you would, you know, say out loud what you're thinking. So yes, it works, but it is a bit, you know, inefficient, right? And you lose, uh, information doing this. Latent space reasoning is this without going, I would say to English or whatever, right? So staying, you know, in what's called the latent space, which is where all the information of an LLM, let's take an LLM, an LLM lives. So this is very much how we, uh, how we, you know, work as humans. Uh, and we move toward what, what Yann Lequin calls, uh, energy based model in which We have different types of, um, uh, longer or shorter, I would say, thinking times, if you will. So that fundamentally, GPUs cannot deliver, deliver this, plain and simple at scale. So these chips can deliver the chips.

AI assessment note: “Latent space reasoning is this without going, I would say to English or whatever”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q At what point do we stop and say, Hey, there is a lot of wastage and we could do more better. How far away are we from that?

A I think until somebody does it. Deep Seek was a good, you know, wake up call, right? Uh, suddenly efficiency is in, right? Um, that's number one. And number two is until there's a new architecture that comes out and changes the game. So in the case of LLMs, for instance, you have these, what's called non transformer models that changes fundamentally the compute requirements. So that might be a frontier that, you know, completely obsoletes the transformers. Right. And if the trans, sorry, the transformers are the, um, I would say the building block by which current model work. Right. So the way they work is that for each, you know, token or syllable, if you will, the model will look at everything behind it. So you can see that as you add more text, you have more work to do. Right. Uh, so there are these new architectures, uh, um, that do not require this, uh, that might Change, you know, these things, uh, and probably shift the amount of compute needed to do training or to do inference. And then there's the new thing, which is, uh, Jan's, um, uh, thesis, which is the, um, uh, word model, right? As in LLMs are at that end, what we need is something that understands the world fundamentally, and this is the, it's, uh, uh, it's JEPA thesis, it's called. I'm very bullish on this, but it's, it's very frontier.

AI assessment note: “I think until somebody does it. Deep Seek was a good, you know, wake up call”

page 1 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.