Assertion Not checkable as stated
Patel: SemiAnalysis tracks all 1,500 global semiconductor fabs
“We track all 1500 fabs in the world. For your purposes, only 50 of them matter, but like, you know, all 1500 fabs around the world.”
Assertion Not checkable as stated
Patel: NVIDIA holds 98% of AI workloads outside Google, 70% overall
“So I would say if you ignored Google, it would be over 98%. But then when you bring Google into the mix, it's actually more like 70 because Google is really that large a percentage of AI workloads especially production workloads.”
Assertion Supported
Patel: Google has run transformers in Search workloads since 2018
“Google was running transformers even in their search workload since 2018, 2019. The advent of BERT, which was one of the most most well-known, most popular transformers before we got to the GPT madness is, has been their, in their production search workloads f…”
Assertion Not checkable as stated
Patel: Apple rents Google's TPU silicon, but GCP AI rentals remain GPUs
“While they do have some customers for their internal silicon externally, such as Apple the vast majority of their external rental business for AI in terms of cloud business is still GPUs.”
Opinion
Patel: Every semiconductor company except NVIDIA is terrible at software
“I would say every semiconductor company in the world sucks at software except for NVIDIA, right?”
Assertion Not checkable as stated
Patel: NVIDIA ships chips from design to deployment faster than competitors
“They get chips out faster than other people from
Thought design to deployed.”
Assertion Open · timeframe Dec 2024
Patel: Google built rack-scale AI systems with Broadcom in 2018 before NVIDIA
“Google actually did this alongside Broadcom you know, and they did it before Nvidia, right? You know, today everyone's freaking out about, or not freaking out, but like everyone's like very excited about Nvidia's Blackwell system, right? It is a rack Of GPUs. …”
Insight
Patel: Semiconductor companies lack the engineers needed for rack-scale systems
“Building a chip is one thing, but building many chips that connect together, cooling them appropriately, networking them together, making sure that it's reliable at that scale is, is a whole host of problems that semiconductor companies don't have the engineer…”
Opinion
Patel: NVIDIA's primary differentiation comes from deep supply chain integration
“I would say for differentiating, NVIDIA has primarily focused on supply chain things, which, you know, might sound like, oh, well like, yeah, they're just like ordering stuff. No, no, no, no. You have to work deeply with the supply chain to build the next gene…”
Assertion Not checkable as stated
Patel: NVIDIA's Jensen Huang only plans 12 to 18 months ahead
“Well, the funny thing is a lot of people at NVIDIA will say Jensen doesn't plan more than a year or year and a half out. Because they change things and they'll deploy them out that fast, right? No semi, every other semiconductor company takes years to deploy, …”
Insight
Patel: NVIDIA's inference moat relies on hardware rather than software
“NVIDIA's moat in, in inference is actually A lot smaller on software but it's a lot bigger on, hey, they just have the best hardware.”
Assertion Partly supported
Patel: NVIDIA is cutting Blackwell margins to compete with custom ASICs
“Like with Blackwell, not only is it way, way, way faster, anywhere from 10 to 15 times on really large models for inference, because they've optimized it for very large language models, they've also decided, hey, we're gonna cut our margin, too, somewhat, beca…”
Prediction Not checkable as stated
Patel: Tanking LLM delivery costs will induce demand for compute
“The cost for delivering LLMs is Is tanking, which is going to induce demand, right?”
Assertion Supported
Patel: Microsoft runs GPT models on AMD hardware for inference
“Microsoft has deployed GPT-style models on On other competitors' hardware, such as AMD they've, and some of their own, but mostly AMD, and so they can wring that out with software because they can spend hundreds of engineers, dozens of engineers' hours hundred…”
Assertion Supported
Gerstner: Jensen projected $1T in new AI and CPU replacement workloads
“And for the first time, he said, not only are we going to have a trillion dollars of new AI workloads over the course of the next four years he said, but we're also going to have a trillion dollars of CPU replacement, of data center replacement workloads over …”
Assertion Supported
Patel: IBM mainframes increase in volume and revenue every cycle
“IBM mainframe sell more volume and revenue every single cycle.”
Assertion Not checkable as stated
Patel: Most AWS data center CPUs are Intel chips from 2015-2020
“The plurality of Amazon's CPUs in their data centers Are 24 core Intel CPUs from, that were manufactured from 2015 to twenty-twenty.”
Insight
Patel: Consolidating legacy CPU servers frees up power for AI workloads
“If I just replace, like, six servers with one, I've basically invented power out of thin air, right? I mean, like, you know, in effect, because these old servers, which are six plus years old, or even, you know, they can just be deprecated and put, so with Cap…”
Opinion
Gerstner: Data centers and power, not GPUs, are the real bottleneck
“What I think it was more a assessment on the real bottleneck, which is data centers and power as opposed to GPUs because GPUs have come online.”
Assertion Not checkable as stated
Patel: AI training has barely tapped vast video data reserves
“We have barely, barely, barely tapped video data, right? So there is a significant amount of data that's not tapped, it's just video data Is so much more information than written data.”
Insight
Patel: Synthetic data generation enables continued AI scaling despite data limits
“You can create data out of thin air almost, right? In certain domains, right? And so this is the whole, the debate around scaling laws is how can we create data?”
Assertion Not checkable as stated
Patel: AI industry is in early days of synthetic data
“Where have we gone on synthetic data? Oh, we're still like very early days, right? We've spent tens of millions of dollars maybe on synthetic data.”
Insight
Patel: Synthetic training only works in functionally verifiable domains like math
“We can't teach it what good art is. Because we have no way to functionally prove what good art is. We can teach it to write really good software. We can teach it how to do mathematical proofs. We can teach it how to engineer systems, because there are, while t…”
Prediction Not checkable as stated
Patel: AI models may improve faster over the next 6-12 months
“We may actually see models improve faster in the next six months to a year than we saw them improve in the last year. Because there's this new axis of synthetic data generation and the amount of compute we can throw at it is, we're still right here in the scal…”
Assertion Not checkable as stated
Patel: AI labs have tapped out human-written internet text data
“Humans post on the internet every day, and we've already tapped that out, right? Kind of more or less on a text.”
Opinion
Patel: Wall Street hyperscaler CapEx estimates are far too low
“I think when you look at the streets estimates for capex, they're all far too low.”
Insight
Patel: Multi-gigawatt buildouts disprove claims that AI scaling is over
“Why is Mark Zuckerberg building a two gigawatt data center in Louisiana? Why is Amazon building these multi gigawatt data centers? Why is Google, why is Microsoft building multiple gigawatt data centers? Plus buying billions and billions of dollars of fiber to…”
Insight
Patel: Modern AI training requires more inference compute than weight updates
“In fact, there's more inference in training than there is updating the model weights, because you have to generate hundreds of possibilities And then, oh, you only train on a couple of them, right?”
Insight
Patel: AI pre-training gains are becoming logarithmically more expensive
“So, the whole paradigm of training, you know, pre-training is, is, is not slowing down. It's just, it's logarithmically more expensive each, for each generation, for each incremental improvement.”
Assertion Supported
Gerstner: NVIDIA trades around 30 P/E compared to Cisco's 2000 peak
“Cisco, you know, 2000 and we'll show it on the pod, but you know, they peaked at like a 120 PE. Right. And yeah, you know, if you look at the fall off that occurred in revenue and in EBITDA, you know, and then it had 70% compression in the priced earnings mult…”
Assertion Supported
Patel: Cisco was fueled by speculative debt whereas NVIDIA is cash-backed
“Cisco's revenue, a lot of it was funded through private-slash-credit investments into building out telecom infrastructure, right? When we look at NVIDIA's revenue sources, very little of it is private-slash-credit, right? And in some cases, yes, it's private-s…”
Assertion Contradicted
Patel: Inflation-adjusted private capital at dot-com peak exceeded today's levels
“At the peak of the dot com, you know, especially once you inflation adjust it the private capital entering the space was much larger than it is today”
Opinion
Gurley: Corporate America invests more in AI than during the internet boom
“I think corporate America is investing more in AI and with more conviction than they did even in the internet wave also.”
Assertion Supported
Patel: GPT-4 cost hundreds of millions to train, generates billions in revenue
“Hundreds of millions of dollars to train GPT-IV. And it's generating billions of dollars of revenue.”
Insight
Patel: Reasoning models like OpenAI o1 increase compute costs by 50x
“When I do this with O-one, right, because it's doing that thinking phase of 10,000... It spends a lot of memory on generating this KV cache and reading this KV cache constantly. Now the maximum batch size, i.e. Concurrent users I can have, is a fraction of tha…”
Insight
Patel: Expensive reasoning queries unlock pricing power by automating new human tasks
“Yes, the queries are expensive, but they're nothing close to the human, right? And so each level of productivity gain I get each level of capabilities jump is a whole new class of tasks that it can do And therefore I can charge for that. Right. So this is the …”
Prediction Held up
Patel: Google and Anthropic will both release reasoning models soon
“There's a Google model that is doing reasoning right now, and it's not released yet, but it's gonna be released soon enough, right? Anthropic is going to release a reasoning model.”
Prediction Not checkable as stated
Patel: Reasoning models will see humongous performance gains within a year
“And so this, the performance improvements we'll get out of these models is, is humongous, right? In, in the coming, you know, six months to a year in certain benchmarks where you have functional verifiers.”
Assertion Not checkable as stated
Patel: Microsoft earns 50% to 70% gross margins on OpenAI models
“Microsoft's earning 50 to 70% gross margins on OpenAI models, and that's with the profit share they get to get, or the share that they give OpenAI, right?”
Assertion Not checkable as stated
Patel: Anthropic showed ~70% gross margins in its most recent round
“Or, you know, Anthropic, similarly, in their most recent round, they were showing, like, 70% gross margins.”
Assertion Contradicted
Patel: A quarter of Nvidia shipments could provide Llama-7B to humanity
“If you take one quarter of NVIDIA shipments and you said all of them are going to inference Lama-Seven-B, you can give every single person on earth a hundred tokens per minute, right? Or sorry, a hundred tokens per second. You give every single person on earth…”
Assertion Not checkable as stated
Patel: SK Hynix memory is NVIDIA's largest COGS item, surpassing TSMC
“When you look at the cost of goods sold of NVIDIA their highest cost of goods sold is not TSMC, which is a Thing that people don't realize. It's actually HBM memory primarily SK Hynix.”
Assertion Partly supported
Patel: Samsung holds almost zero share in HBM memory, especially at NVIDIA
“In HBM, Samsung has almost no share, right? Especially at NVIDIA”
Assertion Not checkable as stated
Patel: Standard high-end server memory yields higher gross margins than HBM
“The gross margins on HBM have not been fantastic. They've been good, but they haven't been fantastic. Actually, regular memory, high-end, like, server memory that is not HBM is actually higher gross margin than HBM.”
Assertion Partly supported
Patel: AWS Trainium2 offers highest HBM capacity and bandwidth per dollar
“Their whole thing at reInvent, if you really talk to them when they announced Trinium II and our whole post about it and our analysis of it is, like, supply chain-wise, this is, looks, you know, you squint your eyes, this looks like an Amazon Basics TPU, right…”
Opinion
Patel: AMD lacks software talent and refuses to fund internal GPU clusters
“AMD is really good, but they're missing software. AMD has no clue how to do software, I think. They've got very few developers on it. They won't spend the money to build a GPU cluster for themselves so that they can develop software.”
Prediction Not checkable as stated
Patel: AMD will see less AI revenue from Microsoft and Meta in 2025
“Yes, I think they'll have they'll have a lot less success with Microsoft than they did this year. And they'll have less success than they did with Meta than they did this year.”
Assertion Supported
Patel: Google TPU clusters scale up to 8,000 chips today
“Nvidia's talking about GB 200, NVL 72, TPUs go to 8000 today, right?”
Assertion Not checkable as stated
Patel: Initial NVIDIA GPU cloud deployments suffer a 5% failure rate
“Google's brought in a level of reliability that NVIDIA GPUs don't have. You know, the dirty secret is to go ask people what the reliability rate of GPUs is in the cloud or in a deployment. It's like, oh God, it is not, they're reliable-ish, like, but like, esp…”
Assertion Not checkable as stated
Patel: Apple accounts for over 70% of Google's TPU rental revenue
“There's only one company accounts for over 70% of Google's revenue from TPUs as far as I understand, and that's Apple.”
Assertion Not checkable as stated
Patel: Amazon Trainium 2 costs ~$5,000 per chip versus $30,000+ for NVIDIA
“You're not paying, you know, north of 30,000, you know, 40,000 dollars per chip for the server, you're paying, you know, significantly less, right? 5000 dollars per chip, right?”
Assertion Supported
Patel: Amazon and Anthropic are building a 400,000-chip Trainium supercomputer
“Amazon and Anthropic have decided to, you know, make a 400,000 tranium server supercomputer, right? 400,000 chips, right?”
Assertion Supported
Patel: Broadcom holds custom ASIC wins with Meta, OpenAI, and Apple
“Broadcom does have multiple custom ASIC wins, right? It's not just Google here. Meta's, Meta's ramping up mostly still for recommendation systems, but their custom chips are gonna get better. You know, there's other players like OpenAI who are making a chip, r…”
Assertion Supported
Patel: Broadcom is building an NVSwitch competitor for AMD and others
“Broadcom is making a Competitors to that, that they will cede to the market, right? Multiple companies will be using that. Not just, you know, AMD will be using that competitor to NVSwitch, but they're not making it themselves because they don't have the skill…”
Prediction Not checkable as stated
Patel: Google TPU purchases will pause for six months over space limits
“Like in the next six months there is a bit of a slowdown in Google TPU purchases because they have no data center space. They want more. They just literally have no data center space to put them.”
Prediction Held up
Patel: Hyperscalers will significantly boost capex in 2025, lifting chip suppliers
“The plans for hyperscalers are pretty firm on, they're gonna spend a crap load more next year, right? And therefore, the ecosystem of networking players, of ASIC vendors, of systems vendors is gonna do well, whether it be NVIDIA or Marvell or Broadcom or AMD o…”
Prediction Not checkable as stated
Patel: Only 5 to 10 of the 80 GPU NeoClouds will survive
“80 NeoClouds are not going to survive. Maybe five to 10 will. And that's because five of those are sovereign, right? And then the other five are like actually like market competitive.”
Assertion Not checkable as stated
Patel: Hyperscalers represent 50% to 60% of AI chip revenue
“Roughly you can say hyperscalers are 50 ish percent of revenue, 50 to 60%, and the rest of it is Neo cloud slash sovereign AI.”
Assertion Supported
Patel: Nvidia Blackwell costs over twice as much to manufacture as Hopper
“The cost to make Blackwell is north of two X that of the cost to make Hopper, right?”
Prediction Not checkable as stated
Patel: Meta and Microsoft may take free cash flows close to zero
“I think Meta and Microsoft may even take their free cash flows close to zero and just spent.”
Insight
Gerstner: AI infrastructure capex growth requires matching 30% inference revenue growth
“So, you know, if you think that infrastructure expenses are going to grow at 30% a year, then I think you have to believe that the underlying inference revenues, right, both on the consumer side and the enterprise side are going to grow somewhere in that range…”
Assertion Not checkable as stated
Gerstner: Lower-tier AI model makers are abandoning the compute arms race
“I think you already see some of these smaller second and third tier models, changing business model, falling aside, no longer engaged in the arms race.”