Sep 1, 2023 · 15m · a16z

The True Cost of Compute

Guido Appenzeller · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this final installment of a three-part AI hardware series from a16z, host and guest Guido Appenzeller explore the economics, mathematical formulas, and strategic implications of AI compute. They unpack the real-world costs of training and running large language models, evaluating whether compute scale acts as a temporary barrier or a permanent competitive moat.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 2.8 Guest teaching 4.8 Guest disagreement 0.2 The host pushing back 0.4
05100:0010:001:19–3:46 · The host as informed peer 2/10 Guido Appenzeller's Background in Infrastructure The host provides background framing and cites Guido's recent article regarding startup compute expenditures exceeding 80% of capital raised. Guido explains how startup hardware spending normalizes as products mature into enterprise software. The tone is entirely collaborative and educational.3:46–5:58 · The host as informed peer 2/10 Mathematical Formulas for Model Compute Requirements The host introduces key parameters like batch size and learning rate, while Guido breaks down the core mathematical approximations for transformer models (2x parameters for inference, 6x for training). Guido leads the technical explanation entirely without friction.5:58–8:12 · The host as informed peer 2/10 Practical Training Cost Breakdown for GPT-3 The host sets up the GPT-3 case study to ground the formulas for listeners. Guido details the napkin math that scales raw GPU hours into tens of millions of dollars once multi-year cloud reservation commitments are factored in.8:12–12:57 · The host as informed peer 5/10 Economic Differences Between Inference and Training The host demonstrates strong topic familiarity by recalling hardware limitations from previous episodes and asking if massive compute costs create an unbreakable moat for incumbents. Guido agrees while adding nuance about data limits and Chinchilla scaling bounds.12:57–14:57 · The host as informed peer 3/10 Scale of Training Data and Episode Conclusion The host closes out the series, featuring a synthetic AI voice clip explaining data scale via library book comparisons. The host supplements this with data on Llama 2's two-trillion token training set.1:19–3:46 · Guest teaching 4/10 Guido Appenzeller's Background in Infrastructure The host provides background framing and cites Guido's recent article regarding startup compute expenditures exceeding 80% of capital raised. Guido explains how startup hardware spending normalizes as products mature into enterprise software. The tone is entirely collaborative and educational.3:46–5:58 · Guest teaching 6/10 Mathematical Formulas for Model Compute Requirements The host introduces key parameters like batch size and learning rate, while Guido breaks down the core mathematical approximations for transformer models (2x parameters for inference, 6x for training). Guido leads the technical explanation entirely without friction.5:58–8:12 · Guest teaching 7/10 Practical Training Cost Breakdown for GPT-3 The host sets up the GPT-3 case study to ground the formulas for listeners. Guido details the napkin math that scales raw GPU hours into tens of millions of dollars once multi-year cloud reservation commitments are factored in.8:12–12:57 · Guest teaching 6/10 Economic Differences Between Inference and Training The host demonstrates strong topic familiarity by recalling hardware limitations from previous episodes and asking if massive compute costs create an unbreakable moat for incumbents. Guido agrees while adding nuance about data limits and Chinchilla scaling bounds.12:57–14:57 · Guest teaching 1/10 Scale of Training Data and Episode Conclusion The host closes out the series, featuring a synthetic AI voice clip explaining data scale via library book comparisons. The host supplements this with data on Llama 2's two-trillion token training set.1:19–3:46 · Guest disagreement 0/10 Guido Appenzeller's Background in Infrastructure The host provides background framing and cites Guido's recent article regarding startup compute expenditures exceeding 80% of capital raised. Guido explains how startup hardware spending normalizes as products mature into enterprise software. The tone is entirely collaborative and educational.3:46–5:58 · Guest disagreement 0/10 Mathematical Formulas for Model Compute Requirements The host introduces key parameters like batch size and learning rate, while Guido breaks down the core mathematical approximations for transformer models (2x parameters for inference, 6x for training). Guido leads the technical explanation entirely without friction.5:58–8:12 · Guest disagreement 0/10 Practical Training Cost Breakdown for GPT-3 The host sets up the GPT-3 case study to ground the formulas for listeners. Guido details the napkin math that scales raw GPU hours into tens of millions of dollars once multi-year cloud reservation commitments are factored in.8:12–12:57 · Guest disagreement 1/10 Economic Differences Between Inference and Training The host demonstrates strong topic familiarity by recalling hardware limitations from previous episodes and asking if massive compute costs create an unbreakable moat for incumbents. Guido agrees while adding nuance about data limits and Chinchilla scaling bounds.12:57–14:57 · Guest disagreement 0/10 Scale of Training Data and Episode Conclusion The host closes out the series, featuring a synthetic AI voice clip explaining data scale via library book comparisons. The host supplements this with data on Llama 2's two-trillion token training set.1:19–3:46 · The host pushing back 0/10 Guido Appenzeller's Background in Infrastructure The host provides background framing and cites Guido's recent article regarding startup compute expenditures exceeding 80% of capital raised. Guido explains how startup hardware spending normalizes as products mature into enterprise software. The tone is entirely collaborative and educational.3:46–5:58 · The host pushing back 0/10 Mathematical Formulas for Model Compute Requirements The host introduces key parameters like batch size and learning rate, while Guido breaks down the core mathematical approximations for transformer models (2x parameters for inference, 6x for training). Guido leads the technical explanation entirely without friction.5:58–8:12 · The host pushing back 0/10 Practical Training Cost Breakdown for GPT-3 The host sets up the GPT-3 case study to ground the formulas for listeners. Guido details the napkin math that scales raw GPU hours into tens of millions of dollars once multi-year cloud reservation commitments are factored in.8:12–12:57 · The host pushing back 2/10 Economic Differences Between Inference and Training The host demonstrates strong topic familiarity by recalling hardware limitations from previous episodes and asking if massive compute costs create an unbreakable moat for incumbents. Guido agrees while adding nuance about data limits and Chinchilla scaling bounds.12:57–14:57 · The host pushing back 0/10 Scale of Training Data and Episode Conclusion The host closes out the series, featuring a synthetic AI voice clip explaining data scale via library book comparisons. The host supplements this with data on Llama 2's two-trillion token training set.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 9:22 Qualifying host's claim on low-cost hardware

Guido gently qualifies the host's assertion that cheap chips cannot be used for training, clarifying that sophisticated software makes it technically possible even if overhead cancels out savings.

Hardest push from the host ▶ 9:58 Challenging incumbent dominance narrative

The host actively probes whether astronomical compute costs lock in victory for heavily capitalized incumbents or if startups can still compete.

Biggest teaching moment ▶ 7:30 Explaining how reservation commitments scale costs

Guido demonstrates why theoretical $1 million compute estimates multiply tenfold in practice due to cloud provider two-year reservation requirements.

The host holds their own ▶ 9:13 Synthesizing hardware limitations across training and inference

The host showcases domain expertise by accurately noting that lower-performance chips cannot easily be combined to train large models due to interconnect bottlenecks.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Guido Appenzeller's Background in Infrastructure 2400 The host provides background framing and cites Guido's recent article regarding startup compute expenditures exceeding 80% of capital raised. Guido explains how startup hardware spending normalizes as products mature into enterprise software. The tone is entirely collaborative and educational.
Mathematical Formulas for Model Compute Requirements 2600 The host introduces key parameters like batch size and learning rate, while Guido breaks down the core mathematical approximations for transformer models (2x parameters for inference, 6x for training). Guido leads the technical explanation entirely without friction.
Practical Training Cost Breakdown for GPT-3 2700 The host sets up the GPT-3 case study to ground the formulas for listeners. Guido details the napkin math that scales raw GPU hours into tens of millions of dollars once multi-year cloud reservation commitments are factored in.
Economic Differences Between Inference and Training 5612 The host demonstrates strong topic familiarity by recalling hardware limitations from previous episodes and asking if massive compute costs create an unbreakable moat for incumbents. Guido agrees while adding nuance about data limits and Chinchilla scaling bounds.
Scale of Training Data and Episode Conclusion 3100 The host closes out the series, featuring a synthetic AI voice clip explaining data scale via library book comparisons. The host supplements this with data on Llama 2's two-trillion token training set.

Statements from this episode (12)

Opinion
Appenzeller: AI training is humanity's most complex computational problem
“There's very few computational problems that complex that mankind has, you know, undertaken.”
Guido Appenzeller Sep 1, 2023 ▶ 0:00
Assertion Not checkable as stated
Appenzeller: AI startups run lean teams to afford compute capacity
“We've specifically seen that with founders that want to train their own models, right, which is extremely expensive, and you have to spend a large chunk of your funding just on compute capacity, right, and they often run with very small lean teams, right, that…”
Guido Appenzeller Sep 1, 2023 ▶ 2:54
Prediction Not checkable as stated
Appenzeller: AI compute share of startup capital will decline over time
“Over time, I would expect this to normalize a little bit, and it's mostly as you go more from the core technology that you're building in very early days towards more a complete product offering, right, there's just a lot more Boxes to check and, you know, fea…”
Guido Appenzeller Sep 1, 2023 ▶ 3:07
Insight
Appenzeller: Transformer inference takes 2x parameter FLOPs, training takes 6x
“In a transformer You can sort of approximate the inference time as twice the number of parameters floating point operations, right? And the training time is about six times the number of parameters.”
Guido Appenzeller Sep 1, 2023 ▶ 4:33
Insight
Appenzeller: Naive AI training runs at under 10% GPU utilization
“If you naively implement it, you can probably go below 10% utilization, but, you know, you can probably get into the tens of percent with a little bit of with a little bit of work.”
Guido Appenzeller Sep 1, 2023 ▶ 5:39
Assertion Supported
Appenzeller: Renting an Nvidia A100 GPU costs $1 to $4 hourly
“Renting an A-Hundred costs you between, I want to say between one and four dollars probably, right? Depending on, on who you rent it from.”
Guido Appenzeller Sep 1, 2023 ▶ 7:04
Assertion Not checkable as stated
Appenzeller: Training large language models costs tens of millions of dollars
“Training one of these large language models today is, it's not a 100,000 dollar thing, it's probably millions of dollars thing. Practically speaking, what we're seeing in industry is that it's actually more of a tens of millions of dollars thing.”
Guido Appenzeller Sep 1, 2023 ▶ 7:39
Assertion Not checkable as stated
Appenzeller: LLM inference costs between a tenth and hundredth of a cent
“You know, if you run the numbers, like a large language model, you actually at a fraction of a cent, like a 10th of a cent or a 100th of a cent, somewhere in that ballpark. For the inference.”
Guido Appenzeller Sep 1, 2023 ▶ 8:33
Assertion Not checkable as stated
Appenzeller: Training open-source LLMs requires $2M to $10M in compute
“You need to find a couple of million or ten million dollars of compute capacity to do it, and that makes it so much harder, right?”
Guido Appenzeller Sep 1, 2023 ▶ 10:50
Prediction Not checkable as stated
Appenzeller: AI model training costs will plateau as human data runs out
“Expectation at the moment is that the cost for training these models, you know, may actually sort of top out or even go down a little bit, you know, as the chips get faster, but we don't discover new training material as quickly.”
Guido Appenzeller Sep 1, 2023 ▶ 12:13
Opinion
Appenzeller: Capital moats in AI are shallow speed bumps
“If that assumption is true, I think this means that the moat that's created by these large capital investments is actually not particularly deep, right? It's more of a speed bump than Then, you know, something that prevents new entrants.”
Guido Appenzeller Sep 1, 2023 ▶ 12:35
Prediction Not checkable as stated
Appenzeller: Well-funded startups will drive future LLM innovation
“Today, training a large language model is something that is definitely within reach for a well-funded startup, right? So, and for that reason, we expect to see more innovation in that area in the future.”
Guido Appenzeller Sep 1, 2023 ▶ 12:47
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.