The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Morin: Nvidia is far from the most efficient AI hardware platform
“But it's by far not the most efficient platform. And arguably, even in terms of software, it's not the best software platform.”
Morin: The market bubble around Nvidia H100 GPUs will burst
“There's going to be a need for inference. Very hard to say whether it will be worth, you know, everybody's money to do it on H 100. That is a bubble that I think will blow some time.”
Morin: GPUs are a clever workaround, not natively built for AI
“GPUs are, you know, are a good trick for AI, but they're not built for AI.”
Morin: AI chip oversupply will lead to GPUs selling at 30% value
“I very much worry there will be an oversupply of these chips. The problem is, is that, you know, remember, the chips are the collateral. So, you know, somewhere, you know, in the US or whatever, there's going to be a data center with like a thousand GPUs that …”
Morin: GPUs cannot deliver latent space AI reasoning at scale
“Fundamentally, GPUs cannot deliver, deliver this, plain and simple at scale.”
Morin: Nvidia may lose market dominance in AI inference and training
“I think that there's a shot that they don't.”
Morin: Google TPUs have more mature software and compute than Nvidia
“But AMD can do training, but it's also, but in terms of maturity, the, by far the most mature software and compute is TPUs, and then it's Nvidia.”
Morin: Stargate data center is inefficient vertical scaling for AI
“It is a vertical scaling. And as you know, my days are spent on efficiency. So I look at these things as being like, all right, this is a bigger, you know, this is an American car of AI. It's big. It consumes a lot of gas, but ultimately, you know, it's not a …”
Morin predicts AI compute will be 95% inference within five years
“In five years, I would say 95% inference, five percent training.”
Morin: Google is the sleeping giant of the AI race
“Google has, like, you know, Android, Google Docs, Whatever, they have everything, they can sprinkle everywhere. This is the sleeping giant in my mind.”
Morin: Running standalone AI model weights will eventually become obsolete
“Models in the sense of, you know, getting, you know, weights and running them is something that is ultimately going away because you know, in favor of like full blown backends, right? You feel like you're talking to a model, but ultimately you're talking to an…”
Morin: Switching from Nvidia to AMD offers 4x spend efficiency
“A simple example is if you know, switch from Nvidia to AMD on a seven TB model, you can get four times better efficiency, right? In terms of spend.”
Morin: Nvidia H100 costs 5x A100 price for 2x inference speed
“H 100 comes along and inference is it's worth five times the price. And it may be runs twice in terms of performance on inference. That is on training. It's a lot better, but on inference, it's like maybe twice as fast when it actually, when it came out, it ra…”
Morin: AI agents and reasoning will disrupt Nvidia's chip dominance
“Ultimately, the two things that could really, very much shake the industry, the chip industry, in my opinion, is our agents and reasoning.”
Morin: Nvidia Blackwell chips suffered surface bending causing cooling issues
“For Blackwell, they assembled two chips. But the surface was so big that the chip started to, you know , wave, like, I don't know the English word, but like, you know, started to bend a bit, which further perpetuated the problem because it then didn't make con…”
Morin: Nvidia won AI training via Mellanox interconnects, not raw compute
“The reason probably Nvidia won, at least in the training space, is because of Mellanox, right? Not because of the raw compute.”
Morin details profit margins across TSMC, Nvidia, and cloud providers
“Nvidia, like a TSMC sells you at 60% margin. Nvidia sells you at, you know, 90% margin. And on top of that, there's Amazon that takes, let's say a 30% margin.”
Morin: Google TPUs lack commercial success outside of Google
“They are very much successful inside of Google, but not much outside of Google, let's say, right?”
Morin: Cloud compute hoarding creates fake AI GPU scarcity
“So in, in the case of, you know, Amazon or Google, that would be buying reserved compute, which you're not going to use because if you buy it on demand, you will get tremendously ripped off. So that creates this like face scarcity of compute because that peopl…”
Morin: Nvidia Blackwell chip shipments are delayed and orders are canceled
“Blackwell is late and orders are getting canceled.”
Morin: Nvidia will remain dominant in AI inference due to availability
“The thing is these chips are on the market. They're here. I can, you know, out tab on Chrome and get one. That is something that, you know, I don't take lightly. Availability that is right. So I think Nvidia is used to stay at least if not for the H-one hundre…”
Morin: Microsoft's AMD chip deployments made OpenAI inference profitable
“Microsoft comes along and buys it all, makes, by the way, OpenAI, or at least on the inference side, puts OpenAI in the green because of the efficiency gains.”
Morin: Being 7x better on cost won't get customers off Nvidia
“I know for a fact that being seven times better and whatever, take whatever metric you want. Whether it's spend, whether it's whatever. It's not enough to get people to switch. People will choose nothing over something.”
Morin: Doubling GPUs in AI inference yields only 10% performance gain
“If you go from one GPU to two, you don't get twice the performance. Maybe you get 10% better performance. Yeah, that's the dirty secret nobody talks about. I'm talking inference, right? So, so you go from, let's say, a hundred to a 110 by doubling the amount o…”
Morin: AMD GPUs achieve 4x inference throughput over Nvidia setups
“If you run on AMD, well, there's enough memory inside the GPU to run one model per card. So you get, you know, eight GPUs, eight times the throughput, while on the other hand, you get eight GPUs, two, you know, two, maybe two and a half times the throughput. S…”
Morin: Nvidia gives zero discounts even on tens of thousands of GPUs
“I talked to a lot of people that build data centers and I tell them, you know, they're, mind you, these people like buy tens of thousands of GPUs. And I asked them, Hey, do you get at least a discount or something? And they're like, no. The only thing we get i…”
Morin: Distilled smaller AI models can outperform their larger base models
“Probably the most, I would say mind blowing thing about distillation is that sometimes the smaller models become better than the bigger model through distillation.”
Morin: Google DeepMind has abandoned model fine-tuning for context windows
“You talk to people at DeepMind And they don't even fine tune anymore. Because they have such, you know, what's called big context window”
Morin: Hardware bans will force China to out-innovate Western AI long-term
“They are constrained, so they are bound to Bound to do better. They can just not buy their way into better compute. So I think it hinders their success, but I think it's short term to think that way.”
Morin: Written-off narratives about Mistral AI's demise are baseless FUD
“They are very competent. So I don't know. I think it's easy to spread FUD. There's a lot of FUD going around, especially about regulation and everything. But here's the thing. I look around me and I don't see You know, what I read. I am hardly convinced about,…”
Morin: Talent and energy are the primary bottlenecks in AI
“What is ultimately the number, the, probably the two limiting factor today is talent. And energy. That's it.”
Morin: AI startups must avoid reselling compute and verticalize on product
“Probably the number one thing I would say is do not resell compute if you can. A lot of, you know, AI startups That are building on top of AI are trying to make a margin, you know, on top of a very big cake. And ultimately what they sell is compute. If you loo…”
Morin: Nvidia forces developers to care about unnecessary details like CUDA
“The thing with Nvidia is that they spend a lot of energy making you care about stuff you shouldn't care about, and they were very successful.”
Morin: Auto-scaling AI inference yields 5x to 10x spend efficiency
“And that's number, probably the number one thing that, you know, gives you a lot of efficiency in terms of spend. Like we're talking, you know, multiples, like, you know, five, you know, sometimes 10 X, you know, improvement.”
Morin: Etched and Visor will bring high-speed inference chips at lower prices
“So my bet is, I think there will be, you know, chips on the market that do that at much lower price. And there's two companies I see going in that direction. One is called Etched. And the other one is called Visor.”
Morin: Scaling SRAM is a dead end for AI hardware
“No, SRAM, this will not deliver. It's a dead end in terms of scaling SRAM means scaling the surface mean you get, you know, depreciating problems.”
Morin: Compute-in-memory will be the next frontier in AI hardware
“So this is the next frontier, and the idea is that instead of, like, transferring the data between external memory and the CPU and do the compute there, you actually, you know, bring the CPU to the memory and you do everything. It's very, you know, it's crazy …”
Morin: Apple purchased 100,000 Trainium AI chips from Amazon
“Let's take Amazon, for instance, with Tranium. Apple just came and said, Hey, we're going to buy a 100,000 of them.”
Morin: Bottom-up AI infrastructure strategies fail because developers do not care
“I think that if you are doing it bottom up, infra to applications, you will lose because nobody will care. As they don't today, right? If you look at TPUs, they're available, they're great. Nobody cares.”
Morin: Bullish on Yann LeCun's JEPA thesis over traditional LLMs
“As in LLMs are at that end, what we need is something that understands the world fundamentally, and this is the, it's JEPA thesis, it's called. I'm very bullish on this, but it's very frontier.”
Morin: AI model market will resemble car makers, not winner-take-all
“Is mental model in terms of model providers, ah, they'll be like car makers. Right? There's no win or tickle. Everybody will have their own.”
Morin: DeepSeek's impact was largely overblown by media narrative and drama
“Yes, Deep Seek made a very good, you know, made waves, but it was, you know, waves that were amplified by the media and the narrative and the drama, right?”
Morin: Founders should focus on success before worrying about European AI regulation
“No, I don't care. I have zero, I, this is something I makes me wonder sometimes. I understand the narrative and so on, but I am absolutely not fearful. Let's be successful first and then we'll talk about the politics.”
Morin: Nvidia intentionally smoothed H100 deliveries to prevent revenue spikes
“The supply of H 100 was actually a smooth out over the year so that they decided so that they didn't have like a big, you know, spike in deliveries and then a quarter less, right?”
Morin: Closed-source AI models are actually complex backend constellations
“At least if you look at close source model, they're not really models. They're more like backend, right? And there are a lot of tricks that you feel like you're talking to one model, but ultimately you're talking to a constellation, an assembly of backends tha…”
Morin: Future AI backend APIs will run locally in enterprise clouds
“The thing is, that API will be running locally, right? Locally, I mean, in your own, you know, cloud, you know, instances, and so on.”
Morin: Groq and Cerebras beat GPUs via on-chip data storage
“Actually, that's why Grok achieves, ah, not Grok, but Grok, Cerebras, and all these folks, they achieve very high performance single stream is because the data is right in the chip that doesn't have to get it from memory, which is slow, which GPU has to do.”
Morin: Interconnect dependency is the core difference between training and inference
“In terms of infra, probably the number one thing that is the number one difference between these two is the need for interconnect. So if you do, you know, production, you, if you can avoid to have interconnect between, you know, let's say a cluster of GPUs, of…”