Insight certainty 4/5 debate potential 3/5

Patel: Reasoning models like OpenAI o1 increase compute costs by 50x

Dylan Patel · AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod · Dec 23, 2024 · at 49:24

Dylan Patel, founder of SemiAnalysis, breaks down the server-level memory overhead and economics of inference-time compute in reasoning models.

0:00 / 0:44exact quote · 44.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“When I do this with O-one, right, because it's doing that thinking phase of 10,000... It spends a lot of memory on generating this KV cache and reading this KV cache constantly. Now the maximum batch size, i.e. Concurrent users I can have, is a fraction of that. One-fourth to one-fifth the number of users can currently use this server. So not only do I need to generate 10 X as many tokens, each token that's generated is four to five X less users. So the cost increase is, is stupendous when you think about a single user cost increase for a single token to be generated is four to five X, but then I'm generating 10 X as many tokens. So you could argue the cost increases, 50 X for an O-one style model on input to output.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dylan Patel

Opinion
Patel: Every semiconductor company except NVIDIA is terrible at software
“I would say every semiconductor company in the world sucks at software except for NVIDIA, right?”
Dylan Patel Dec 23, 2024 ▶ 7:07 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
Opinion
Patel: AMD lacks software talent and refuses to fund internal GPU clusters
“AMD is really good, but they're missing software. AMD has no clue how to do software, I think. They've got very few developers on it. They won't spend the money to build a GPU cluster for themselves so that they can develop software.”
Dylan Patel Dec 23, 2024 ▶ 1:07:36 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
Assertion Not checkable as stated
Patel: Initial NVIDIA GPU cloud deployments suffer a 5% failure rate
“Google's brought in a level of reliability that NVIDIA GPUs don't have. You know, the dirty secret is to go ask people what the reliability rate of GPUs is in the cloud or in a deployment. It's like, oh God, it is not, they're reliable-ish, like, but like, esp…”
Dylan Patel Dec 23, 2024 ▶ 1:11:54 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
Prediction Not checkable as stated
Patel: Meta and Microsoft may take free cash flows close to zero
“I think Meta and Microsoft may even take their free cash flows close to zero and just spent.”
Dylan Patel Dec 23, 2024 ▶ 1:24:40 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
Assertion Not checkable as stated
Patel: NVIDIA's Jensen Huang only plans 12 to 18 months ahead
“Well, the funny thing is a lot of people at NVIDIA will say Jensen doesn't plan more than a year or year and a half out. Because they change things and they'll deploy them out that fast, right? No semi, every other semiconductor company takes years to deploy, …”
Dylan Patel Dec 23, 2024 ▶ 12:58 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
Insight
Patel: NVIDIA's inference moat relies on hardware rather than software
“NVIDIA's moat in, in inference is actually A lot smaller on software but it's a lot bigger on, hey, they just have the best hardware.”
Dylan Patel Dec 23, 2024 ▶ 13:53 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 40 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.