Mar 5, 2025 · 27m · a16z
DeepSeek, Reasoning Models, and the Future of LLMs
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the a16z podcast, hosts Guido Appenzeller and Marco Mascorro analyze DeepSeek's technical breakthroughs, explaining how reasoning models, reinforcement learning, mixture-of-experts architectures, and test-time compute scaling are reshaping the landscape of artificial intelligence.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
In a completely collaborative episode, Marco provides the closest nuance reframe by clarifying that R1 was not a sudden standalone discovery, but rather a compilation of existing techniques developed across previous models.
Hardest push from the host ▶ 18:55 Host qualifying public training cost figuresGuido pushes back on taking the headline $5.5M training cost at face value, pointing out using an aircraft analogy that final execution runs represent only a fraction of total R&D compute expenditure.
Biggest teaching moment ▶ 25:05 Explaining distillation superiority on smaller architecturesMarco educates the audience on DeepSeek's paper finding that applying RL directly to small models like Llama 7B yields minimal gains, whereas distilling reasoning traces from R1 delivers massive performance jumps.
The host holds their own ▶ 16:19 Fermi estimation of synthetic data economic valueGuido demonstrates strong domain authority by calculating live Fermi estimates of PhD labor costs to demonstrate that computer-generated synthetic reasoning traces saved $60M compared to human annotation.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| Reasoning Models vs Classic LLMs Output Comparison | 5 | 4 | 0 | 0 | Guido introduces the performance shift toward reasoning models using an orbital mechanics prompt example to illustrate step-by-step thinking. Marco explains the mechanics of traditional pre-training, SFT, and RLHF, with Guido framing the interview questions collaboratively. | |
| Evolution of DeepSeek Models and GRPO | 4 | 5 | 0 | 0 | Marco outlines how DeepSeek R1 compiled prior innovations like Multi-head Latent Attention and GRPO from DeepSeek Math. Guido adds context on how verifiable math and code rewards allow self-learning without human feedback. | |
| DeepSeek-R1-Zero and Pure Reinforcement Learning | 6 | 5 | 0 | 0 | Guido demonstrates technical knowledge by contrasting DeepSeek V3's 256 mixture-of-experts architecture with Mixtral's 8 experts. Marco elaborates on how pure RL in R1-Zero produced raw reasoning alongside language-mixing flaws. | |
| DeepSeek-R1 Full Training Pipeline | 5 | 4 | 0 | 0 | Guido breaks down the training flowchart and computes that reasoning responses reached up to 10,000 tokens or roughly 20 pages. Marco explains the role of thinking tokens and post-training alignment steps. | |
| Synthetic Data Generation and Multi-Stage Alignment | 6 | 5 | 0 | 0 | Guido performs a back-of-the-envelope calculation showing that synthetically generating 600,000 reasoning traces replaced an estimated $60M in human PhD labor. Marco details the multi-stage alignment pipeline involving cold start data and preference tuning. | |
| Training Economics and Architectural Innovations | 6 | 6 | 0 | 1 | Guido contrasts DeepSeek's reported $5.5M training cost with Llama benchmark costs (~$3M), pointing out that final run costs omit broader R&D experimentation budgets. Marco explains technical innovations in MLA, decoupled RoPE, and GRPO efficiency. | |
| Test-Time Compute, Model Distillation, and Future Outlook | 6 | 5 | 0 | 0 | Guido analyzes test-time compute shifts, quoting Jensen Huang's scaling curves and mentioning local execution via Ollama. Marco highlights key findings showing model distillation outperforms direct RL on smaller base models like Llama 7B. |