“When I do this with O-one, right, because it's doing that thinking phase of 10,000... It spends a lot of memory on generating this KV cache and reading this KV cache constantly. Now the maximum batch size, i.e. Concurrent users I can have, is a fraction of that. One-fourth to one-fifth the number of users can currently use this server. So not only do I need to generate 10 X as many tokens, each token that's generated is four to five X less users. So the cost increase is, is stupendous when you think about a single user cost increase for a single token to be generated is four to five X, but then I'm generating 10 X as many tokens. So you could argue the cost increases, 50 X for an O-one style model on input to output.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Dylan Patel
Opinion
Patel: Every semiconductor company except NVIDIA is terrible at software
“I would say every semiconductor company in the world sucks at software except for NVIDIA, right?”
Patel: AMD lacks software talent and refuses to fund internal GPU clusters
“AMD is really good, but they're missing software. AMD has no clue how to do software, I think. They've got very few developers on it. They won't spend the money to build a GPU cluster for themselves so that they can develop software.”
“Google's brought in a level of reliability that NVIDIA GPUs don't have. You know, the dirty secret is to go ask people what the reliability rate of GPUs is in the cloud or in a deployment. It's like, oh God, it is not, they're reliable-ish, like, but like, esp…”
Patel: NVIDIA's Jensen Huang only plans 12 to 18 months ahead
“Well, the funny thing is a lot of people at NVIDIA will say Jensen doesn't plan more than a year or year and a half out. Because they change things and they'll deploy them out that fast, right? No semi, every other semiconductor company takes years to deploy, …”
This entire site, over 40 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.