Mar 5, 2025 · 27m · a16z

DeepSeek, Reasoning Models, and the Future of LLMs

Guido Appenzeller · 12m spoken Marco Mascorro · 12m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, hosts Guido Appenzeller and Marco Mascorro analyze DeepSeek's technical breakthroughs, explaining how reasoning models, reinforcement learning, mixture-of-experts architectures, and test-time compute scaling are reshaping the landscape of artificial intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 5.4 Guest teaching 4.9 Guest disagreement 0.0 The host pushing back 0.1
05100:0010:0020:000:48–4:48 · The host as informed peer 5/10 Reasoning Models vs Classic LLMs Output Comparison Guido introduces the performance shift toward reasoning models using an orbital mechanics prompt example to illustrate step-by-step thinking. Marco explains the mechanics of traditional pre-training, SFT, and RLHF, with Guido framing the interview questions collaboratively.4:48–7:04 · The host as informed peer 4/10 Evolution of DeepSeek Models and GRPO Marco outlines how DeepSeek R1 compiled prior innovations like Multi-head Latent Attention and GRPO from DeepSeek Math. Guido adds context on how verifiable math and code rewards allow self-learning without human feedback.7:04–9:50 · The host as informed peer 6/10 DeepSeek-R1-Zero and Pure Reinforcement Learning Guido demonstrates technical knowledge by contrasting DeepSeek V3's 256 mixture-of-experts architecture with Mixtral's 8 experts. Marco elaborates on how pure RL in R1-Zero produced raw reasoning alongside language-mixing flaws.9:50–12:39 · The host as informed peer 5/10 DeepSeek-R1 Full Training Pipeline Guido breaks down the training flowchart and computes that reasoning responses reached up to 10,000 tokens or roughly 20 pages. Marco explains the role of thinking tokens and post-training alignment steps.12:39–17:54 · The host as informed peer 6/10 Synthetic Data Generation and Multi-Stage Alignment Guido performs a back-of-the-envelope calculation showing that synthetically generating 600,000 reasoning traces replaced an estimated $60M in human PhD labor. Marco details the multi-stage alignment pipeline involving cold start data and preference tuning.17:54–21:37 · The host as informed peer 6/10 Training Economics and Architectural Innovations Guido contrasts DeepSeek's reported $5.5M training cost with Llama benchmark costs (~$3M), pointing out that final run costs omit broader R&D experimentation budgets. Marco explains technical innovations in MLA, decoupled RoPE, and GRPO efficiency.21:37–27:06 · The host as informed peer 6/10 Test-Time Compute, Model Distillation, and Future Outlook Guido analyzes test-time compute shifts, quoting Jensen Huang's scaling curves and mentioning local execution via Ollama. Marco highlights key findings showing model distillation outperforms direct RL on smaller base models like Llama 7B.0:48–4:48 · Guest teaching 4/10 Reasoning Models vs Classic LLMs Output Comparison Guido introduces the performance shift toward reasoning models using an orbital mechanics prompt example to illustrate step-by-step thinking. Marco explains the mechanics of traditional pre-training, SFT, and RLHF, with Guido framing the interview questions collaboratively.4:48–7:04 · Guest teaching 5/10 Evolution of DeepSeek Models and GRPO Marco outlines how DeepSeek R1 compiled prior innovations like Multi-head Latent Attention and GRPO from DeepSeek Math. Guido adds context on how verifiable math and code rewards allow self-learning without human feedback.7:04–9:50 · Guest teaching 5/10 DeepSeek-R1-Zero and Pure Reinforcement Learning Guido demonstrates technical knowledge by contrasting DeepSeek V3's 256 mixture-of-experts architecture with Mixtral's 8 experts. Marco elaborates on how pure RL in R1-Zero produced raw reasoning alongside language-mixing flaws.9:50–12:39 · Guest teaching 4/10 DeepSeek-R1 Full Training Pipeline Guido breaks down the training flowchart and computes that reasoning responses reached up to 10,000 tokens or roughly 20 pages. Marco explains the role of thinking tokens and post-training alignment steps.12:39–17:54 · Guest teaching 5/10 Synthetic Data Generation and Multi-Stage Alignment Guido performs a back-of-the-envelope calculation showing that synthetically generating 600,000 reasoning traces replaced an estimated $60M in human PhD labor. Marco details the multi-stage alignment pipeline involving cold start data and preference tuning.17:54–21:37 · Guest teaching 6/10 Training Economics and Architectural Innovations Guido contrasts DeepSeek's reported $5.5M training cost with Llama benchmark costs (~$3M), pointing out that final run costs omit broader R&D experimentation budgets. Marco explains technical innovations in MLA, decoupled RoPE, and GRPO efficiency.21:37–27:06 · Guest teaching 5/10 Test-Time Compute, Model Distillation, and Future Outlook Guido analyzes test-time compute shifts, quoting Jensen Huang's scaling curves and mentioning local execution via Ollama. Marco highlights key findings showing model distillation outperforms direct RL on smaller base models like Llama 7B.0:48–4:48 · Guest disagreement 0/10 Reasoning Models vs Classic LLMs Output Comparison Guido introduces the performance shift toward reasoning models using an orbital mechanics prompt example to illustrate step-by-step thinking. Marco explains the mechanics of traditional pre-training, SFT, and RLHF, with Guido framing the interview questions collaboratively.4:48–7:04 · Guest disagreement 0/10 Evolution of DeepSeek Models and GRPO Marco outlines how DeepSeek R1 compiled prior innovations like Multi-head Latent Attention and GRPO from DeepSeek Math. Guido adds context on how verifiable math and code rewards allow self-learning without human feedback.7:04–9:50 · Guest disagreement 0/10 DeepSeek-R1-Zero and Pure Reinforcement Learning Guido demonstrates technical knowledge by contrasting DeepSeek V3's 256 mixture-of-experts architecture with Mixtral's 8 experts. Marco elaborates on how pure RL in R1-Zero produced raw reasoning alongside language-mixing flaws.9:50–12:39 · Guest disagreement 0/10 DeepSeek-R1 Full Training Pipeline Guido breaks down the training flowchart and computes that reasoning responses reached up to 10,000 tokens or roughly 20 pages. Marco explains the role of thinking tokens and post-training alignment steps.12:39–17:54 · Guest disagreement 0/10 Synthetic Data Generation and Multi-Stage Alignment Guido performs a back-of-the-envelope calculation showing that synthetically generating 600,000 reasoning traces replaced an estimated $60M in human PhD labor. Marco details the multi-stage alignment pipeline involving cold start data and preference tuning.17:54–21:37 · Guest disagreement 0/10 Training Economics and Architectural Innovations Guido contrasts DeepSeek's reported $5.5M training cost with Llama benchmark costs (~$3M), pointing out that final run costs omit broader R&D experimentation budgets. Marco explains technical innovations in MLA, decoupled RoPE, and GRPO efficiency.21:37–27:06 · Guest disagreement 0/10 Test-Time Compute, Model Distillation, and Future Outlook Guido analyzes test-time compute shifts, quoting Jensen Huang's scaling curves and mentioning local execution via Ollama. Marco highlights key findings showing model distillation outperforms direct RL on smaller base models like Llama 7B.0:48–4:48 · The host pushing back 0/10 Reasoning Models vs Classic LLMs Output Comparison Guido introduces the performance shift toward reasoning models using an orbital mechanics prompt example to illustrate step-by-step thinking. Marco explains the mechanics of traditional pre-training, SFT, and RLHF, with Guido framing the interview questions collaboratively.4:48–7:04 · The host pushing back 0/10 Evolution of DeepSeek Models and GRPO Marco outlines how DeepSeek R1 compiled prior innovations like Multi-head Latent Attention and GRPO from DeepSeek Math. Guido adds context on how verifiable math and code rewards allow self-learning without human feedback.7:04–9:50 · The host pushing back 0/10 DeepSeek-R1-Zero and Pure Reinforcement Learning Guido demonstrates technical knowledge by contrasting DeepSeek V3's 256 mixture-of-experts architecture with Mixtral's 8 experts. Marco elaborates on how pure RL in R1-Zero produced raw reasoning alongside language-mixing flaws.9:50–12:39 · The host pushing back 0/10 DeepSeek-R1 Full Training Pipeline Guido breaks down the training flowchart and computes that reasoning responses reached up to 10,000 tokens or roughly 20 pages. Marco explains the role of thinking tokens and post-training alignment steps.12:39–17:54 · The host pushing back 0/10 Synthetic Data Generation and Multi-Stage Alignment Guido performs a back-of-the-envelope calculation showing that synthetically generating 600,000 reasoning traces replaced an estimated $60M in human PhD labor. Marco details the multi-stage alignment pipeline involving cold start data and preference tuning.17:54–21:37 · The host pushing back 1/10 Training Economics and Architectural Innovations Guido contrasts DeepSeek's reported $5.5M training cost with Llama benchmark costs (~$3M), pointing out that final run costs omit broader R&D experimentation budgets. Marco explains technical innovations in MLA, decoupled RoPE, and GRPO efficiency.21:37–27:06 · The host pushing back 0/10 Test-Time Compute, Model Distillation, and Future Outlook Guido analyzes test-time compute shifts, quoting Jensen Huang's scaling curves and mentioning local execution via Ollama. Marco highlights key findings showing model distillation outperforms direct RL on smaller base models like Llama 7B.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 5:21 Gentle reframe of R1 as iterative rather than a sudden breakthrough

In a completely collaborative episode, Marco provides the closest nuance reframe by clarifying that R1 was not a sudden standalone discovery, but rather a compilation of existing techniques developed across previous models.

Hardest push from the host ▶ 18:55 Host qualifying public training cost figures

Guido pushes back on taking the headline $5.5M training cost at face value, pointing out using an aircraft analogy that final execution runs represent only a fraction of total R&D compute expenditure.

Biggest teaching moment ▶ 25:05 Explaining distillation superiority on smaller architectures

Marco educates the audience on DeepSeek's paper finding that applying RL directly to small models like Llama 7B yields minimal gains, whereas distilling reasoning traces from R1 delivers massive performance jumps.

The host holds their own ▶ 16:19 Fermi estimation of synthetic data economic value

Guido demonstrates strong domain authority by calculating live Fermi estimates of PhD labor costs to demonstrate that computer-generated synthetic reasoning traces saved $60M compared to human annotation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Reasoning Models vs Classic LLMs Output Comparison 5400 Guido introduces the performance shift toward reasoning models using an orbital mechanics prompt example to illustrate step-by-step thinking. Marco explains the mechanics of traditional pre-training, SFT, and RLHF, with Guido framing the interview questions collaboratively.
Evolution of DeepSeek Models and GRPO 4500 Marco outlines how DeepSeek R1 compiled prior innovations like Multi-head Latent Attention and GRPO from DeepSeek Math. Guido adds context on how verifiable math and code rewards allow self-learning without human feedback.
DeepSeek-R1-Zero and Pure Reinforcement Learning 6500 Guido demonstrates technical knowledge by contrasting DeepSeek V3's 256 mixture-of-experts architecture with Mixtral's 8 experts. Marco elaborates on how pure RL in R1-Zero produced raw reasoning alongside language-mixing flaws.
DeepSeek-R1 Full Training Pipeline 5400 Guido breaks down the training flowchart and computes that reasoning responses reached up to 10,000 tokens or roughly 20 pages. Marco explains the role of thinking tokens and post-training alignment steps.
Synthetic Data Generation and Multi-Stage Alignment 6500 Guido performs a back-of-the-envelope calculation showing that synthetically generating 600,000 reasoning traces replaced an estimated $60M in human PhD labor. Marco details the multi-stage alignment pipeline involving cold start data and preference tuning.
Training Economics and Architectural Innovations 6601 Guido contrasts DeepSeek's reported $5.5M training cost with Llama benchmark costs (~$3M), pointing out that final run costs omit broader R&D experimentation budgets. Marco explains technical innovations in MLA, decoupled RoPE, and GRPO efficiency.
Test-Time Compute, Model Distillation, and Future Outlook 6500 Guido analyzes test-time compute shifts, quoting Jensen Huang's scaling curves and mentioning local execution via Ollama. Marco highlights key findings showing model distillation outperforms direct RL on smaller base models like Llama 7B.

Statements from this episode (17)

Assertion Supported
Mascorro: DeepSeek consistently open-sources its model weights and training techniques
“So, one of the good things about DeepSeek is basically they open source their weights, their techniques, and how they build these models, and they've been doing that for a while.”
Marco Mascorro Mar 5, 2025 ▶ 0:19
Prediction Not checkable as stated
Appenzeller: All future state-of-the-art AI models will use reasoning techniques
“And I think looking forward, From now on, pretty much any state of the art model will use some of those techniques, and we've seen this already, you know, from models from OpenAI and models from Google that are structurally very, very similar, and this has hug…”
Guido Appenzeller Mar 5, 2025 ▶ 0:31
Assertion Supported
Appenzeller: Reasoning models now dominate top AI model rankings
“If you look at the slide here that shows the current ranking of one of the best AI models that we have today, you'll see that pretty much the whole top of the rankings has been taken over by reasoning models.”
Guido Appenzeller Mar 5, 2025 ▶ 0:51
Insight
Appenzeller: Small distilled models use reasoning to overcome limited memory capacity
“On the right side, we have a distilled version of DeepSeq R-one. So this is a very, very small model. It can't actually answer this directly from memory, but what it does, it starts reasoning. And if you read the text, right, it really starts to hustle. It's t…”
Guido Appenzeller Mar 5, 2025 ▶ 1:42
Insight
Mascorro: High-quality LLMs can be built purely with SFT data
“Now, the reality is like you can get to really good models purely with like SFT data.”
Marco Mascorro Mar 5, 2025 ▶ 4:02
Insight
Mascorro: DeepSeek-R1 proved reinforcement learning improves models without human feedback
“And I think the big thing in, in R-one, or generally with these reasoning models is, We were doing before there was a human in the loop always, right? Like when we have this SFT training and these other techniques that we're doing after like RLHF and having R …”
Marco Mascorro Mar 5, 2025 ▶ 6:42
Assertion Supported
Mascorro: DeepSeek-R1-Zero Improved Math Scores but Struggled with Readability and Language Switching
“R one zero, which in a way was a very interesting model because it showed that it improved in some reasoning benchmarks and math benchmarks. But eventually didn't do really well on other things, right? Like it was switching between languages. I think that was …”
Marco Mascorro Mar 5, 2025 ▶ 7:53
Assertion Supported
Mascorro: DeepSeek-V3 features 256 experts, far exceeding typical open-source models
“We talk about it as 256 experts, which is a large, a relative large number of experts in terms of at least open source models that we've seen out there.”
Marco Mascorro Mar 5, 2025 ▶ 9:00
Opinion
Appenzeller: DeepSeek-R1-Zero is arguably better at reasoning than the final R1
“DeepSeq R-one-zero is actually a very, very good reasoning model. It's arguably better in reasoning than the final DeepSeq R-one.”
Guido Appenzeller Mar 5, 2025 ▶ 10:54
Assertion Supported
Mascorro: DeepSeek-R1 post-training used two SFT and two RL phases
“So basically the way they did that, trying to fix R one zero, is it added a couple more phases in the post-training. That included two supervised fine tuning phases and two reinforcement learning phases. And these reinforcement learning phases, they were a lar…”
Marco Mascorro Mar 5, 2025 ▶ 11:25
Opinion
Appenzeller: DeepSeek-R1-Zero was the quantum leap in performance
“So basically, R-one-zero, that was the quantum leap in model performance.”
Guido Appenzeller Mar 5, 2025 ▶ 13:41
Assertion Not checkable as stated
Appenzeller: Generating human reasoning traces for AI training would cost $60 million
“You know, for some of the math problems, you definitely want somebody with a graduate degree. If, I mean, we saw that the average answer was 20 pages. How much do I have to pay a math PhD to generate 20 pages of text, right, if this is a I don't know, a hundre…”
Guido Appenzeller Mar 5, 2025 ▶ 16:20
Assertion Supported
Appenzeller: DeepSeek spent $5.5 million at market rates to train V3
“They quoted 5.5 million dollars, I think, you know, at market rates to train it.”
Guido Appenzeller Mar 5, 2025 ▶ 18:14
Assertion Not checkable as stated
Appenzeller: Pre-training Meta's Llama models cost slightly over $3 million
“We ran the cost for some Lama models last year, and I think we ended up with, you know, a little over three million dollars, right?”
Guido Appenzeller Mar 5, 2025 ▶ 18:31
Insight
Appenzeller: The final training run is a small fraction of total costs
“The final test run is often not the majority of the money that you spend, right? You need many test runs that don't work well. You're highly paid PhDs that, that, that do the do the training probably highly paid. They need sort of infrastructure to experiment …”
Guido Appenzeller Mar 5, 2025 ▶ 18:47
Assertion Not checkable as stated
Appenzeller: Switching to reasoning models would increase inference compute needs 20x
“Very roughly, if we all, if everybody would switch tomorrow from whatever they have today to a reasoning model, we would need 20 times more inference.”
Guido Appenzeller Mar 5, 2025 ▶ 22:30
Assertion Supported
Mascorro: Distillations from DeepSeek-R1 Outperformed Direct RL on Smaller Models
“So it turns out in their experiments, they took Lama's EV and some of these are QN models, and they basically apply RL straight the same way they did it with R one on these base models. And it turns out that it improved in some fields, but it was not a signifi…”
Marco Mascorro Mar 5, 2025 ▶ 25:26
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.