The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Morin: The market bubble around Nvidia H100 GPUs will burst
“There's going to be a need for inference. Very hard to say whether it will be worth, you know, everybody's money to do it on H 100. That is a bubble that I think will blow some time.”
Morin: AI chip oversupply will lead to GPUs selling at 30% value
“I very much worry there will be an oversupply of these chips. The problem is, is that, you know, remember, the chips are the collateral. So, you know, somewhere, you know, in the US or whatever, there's going to be a data center with like a thousand GPUs that …”
Morin: GPUs cannot deliver latent space AI reasoning at scale
“Fundamentally, GPUs cannot deliver, deliver this, plain and simple at scale.”
Morin: Nvidia may lose market dominance in AI inference and training
“I think that there's a shot that they don't.”
Morin predicts AI compute will be 95% inference within five years
“In five years, I would say 95% inference, five percent training.”
Morin: Running standalone AI model weights will eventually become obsolete
“Models in the sense of, you know, getting, you know, weights and running them is something that is ultimately going away because you know, in favor of like full blown backends, right? You feel like you're talking to a model, but ultimately you're talking to an…”
Morin: Switching from Nvidia to AMD offers 4x spend efficiency
“A simple example is if you know, switch from Nvidia to AMD on a seven TB model, you can get four times better efficiency, right? In terms of spend.”
Morin: Nvidia H100 costs 5x A100 price for 2x inference speed
“H 100 comes along and inference is it's worth five times the price. And it may be runs twice in terms of performance on inference. That is on training. It's a lot better, but on inference, it's like maybe twice as fast when it actually, when it came out, it ra…”
Morin: Nvidia Blackwell chips suffered surface bending causing cooling issues
“For Blackwell, they assembled two chips. But the surface was so big that the chip started to, you know , wave, like, I don't know the English word, but like, you know, started to bend a bit, which further perpetuated the problem because it then didn't make con…”
Morin details profit margins across TSMC, Nvidia, and cloud providers
“Nvidia, like a TSMC sells you at 60% margin. Nvidia sells you at, you know, 90% margin. And on top of that, there's Amazon that takes, let's say a 30% margin.”
Morin: Google TPUs lack commercial success outside of Google
“They are very much successful inside of Google, but not much outside of Google, let's say, right?”
Morin: Nvidia Blackwell chip shipments are delayed and orders are canceled
“Blackwell is late and orders are getting canceled.”
Morin: Nvidia will remain dominant in AI inference due to availability
“The thing is these chips are on the market. They're here. I can, you know, out tab on Chrome and get one. That is something that, you know, I don't take lightly. Availability that is right. So I think Nvidia is used to stay at least if not for the H-one hundre…”
Morin: Microsoft's AMD chip deployments made OpenAI inference profitable
“Microsoft comes along and buys it all, makes, by the way, OpenAI, or at least on the inference side, puts OpenAI in the green because of the efficiency gains.”
Morin: Doubling GPUs in AI inference yields only 10% performance gain
“If you go from one GPU to two, you don't get twice the performance. Maybe you get 10% better performance. Yeah, that's the dirty secret nobody talks about. I'm talking inference, right? So, so you go from, let's say, a hundred to a 110 by doubling the amount o…”
Morin: AMD GPUs achieve 4x inference throughput over Nvidia setups
“If you run on AMD, well, there's enough memory inside the GPU to run one model per card. So you get, you know, eight GPUs, eight times the throughput, while on the other hand, you get eight GPUs, two, you know, two, maybe two and a half times the throughput. S…”
Morin: Nvidia gives zero discounts even on tens of thousands of GPUs
“I talked to a lot of people that build data centers and I tell them, you know, they're, mind you, these people like buy tens of thousands of GPUs. And I asked them, Hey, do you get at least a discount or something? And they're like, no. The only thing we get i…”
Morin: Distilled smaller AI models can outperform their larger base models
“Probably the most, I would say mind blowing thing about distillation is that sometimes the smaller models become better than the bigger model through distillation.”
Morin: Google DeepMind has abandoned model fine-tuning for context windows
“You talk to people at DeepMind And they don't even fine tune anymore. Because they have such, you know, what's called big context window”
Morin: Hardware bans will force China to out-innovate Western AI long-term
“They are constrained, so they are bound to Bound to do better. They can just not buy their way into better compute. So I think it hinders their success, but I think it's short term to think that way.”
Morin: Auto-scaling AI inference yields 5x to 10x spend efficiency
“And that's number, probably the number one thing that, you know, gives you a lot of efficiency in terms of spend. Like we're talking, you know, multiples, like, you know, five, you know, sometimes 10 X, you know, improvement.”
Morin: Etched and Visor will bring high-speed inference chips at lower prices
“So my bet is, I think there will be, you know, chips on the market that do that at much lower price. And there's two companies I see going in that direction. One is called Etched. And the other one is called Visor.”
Morin: Compute-in-memory will be the next frontier in AI hardware
“So this is the next frontier, and the idea is that instead of, like, transferring the data between external memory and the CPU and do the compute there, you actually, you know, bring the CPU to the memory and you do everything. It's very, you know, it's crazy …”
Morin: Apple purchased 100,000 Trainium AI chips from Amazon
“Let's take Amazon, for instance, with Tranium. Apple just came and said, Hey, we're going to buy a 100,000 of them.”
Morin: AI model market will resemble car makers, not winner-take-all
“Is mental model in terms of model providers, ah, they'll be like car makers. Right? There's no win or tickle. Everybody will have their own.”
Morin: Nvidia intentionally smoothed H100 deliveries to prevent revenue spikes
“The supply of H 100 was actually a smooth out over the year so that they decided so that they didn't have like a big, you know, spike in deliveries and then a quarter less, right?”
Morin: Future AI backend APIs will run locally in enterprise clouds
“The thing is, that API will be running locally, right? Locally, I mean, in your own, you know, cloud, you know, instances, and so on.”
Morin: Groq and Cerebras beat GPUs via on-chip data storage
“Actually, that's why Grok achieves, ah, not Grok, but Grok, Cerebras, and all these folks, they achieve very high performance single stream is because the data is right in the chip that doesn't have to get it from memory, which is slow, which GPU has to do.”
Morin: xAI cluster is four 25k GPU networks, not 100k unified
“The XAI cluster. It's not a 100,000 GPUs. It is four times 25,000.”
Morin: China's domestic AI chips currently match Nvidia's A100 capability
“They're a bit late in terms of, you know, ASIC. There are like A-one-hundred level”
Morin: Latency reasoning is 2025's fundamental AI infrastructure shift
“Latency reasoning. Definitely. This year. So, you know, as I was saying, like, the shift from throughput, so how speed my answers to how long it takes for my answer complete to appear. That is probably one of the fundamental, like this year, right?”