Jan 24, 2025 · 23m · latent-space

The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1

Shawn Wang · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Latent Space Podcast episode, the Bespoke Labs team discusses their 48-hour sprint distilling DeepSeek-R1 into Qwen, demonstrating how open reasoning traces, streamlined data curation, and targeted post-training allow compact 7B models to rival frontier reasoning systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.8% of the talking time here. How this is scored →

The hosts as informed peer 5.6 Guest teaching 5.0 Guest disagreement 0.1 The hosts pushing back 1.3
05100:0010:0020:001:24–4:23 · The hosts as informed peer 5/10 The 48-Hour Sprint: Distilling DeepSeek R1 Swyx sets up the context of the 48-hour sprint following DeepSeek R1's release. Mahesh and Ryan detail the rapid iteration using their Curator tool to distill R1 into Qwen.4:23–6:47 · The hosts as informed peer 4/10 Defining Model Distillation: From Logits to Synthetic Data Alessio asks for clarification on distillation terminology. Mahesh provides a clear technical breakdown ranging from Hinton's 2016 logit distillation to synthetic data fine-tuning.6:48–10:24 · The hosts as informed peer 4/10 Open Reasoning Traces and Autoregressive Backtracking The conversation covers the opacity of OpenAI o1 traces versus open R1 traces. Trung and Mahesh explain how autoregressive generation performs implicit backtracking without separate search algorithms.10:25–15:08 · The hosts as informed peer 8/10 Frontier Search Architectures vs. Distilled Small Models Swyx challenges the idea that search is entirely discarded, drawing distinctions between frontier training/inference architectures and distilled mini models while referencing PRMs and MCTS.15:09–17:56 · The hosts as informed peer 5/10 Pipeline Differences: Coherence and Ablating Reannotation Swyx probes why Bespoke's 7B model succeeded where Sky-T1 failed. Trung explains their specific pipeline ablation of removing the GPT-4o-mini reannotation step because R1 reasoning traces were already coherent.17:58–21:08 · The hosts as informed peer 8/10 Enabling Student Models to Surpass Teacher Models Swyx demonstrates deep background knowledge citing the STaR paper and Cosine's SWE-bench results, while Mahesh shares how their 7B MiniCheck model outperformed Llama 3 405B.21:09–23:14 · The hosts as informed peer 5/10 Data Curation Principles and the Bespoke Curator Platform Mahesh outlines the three fundamental axes of data curation—quantity, quality, and diversity—while Swyx touches on FineWeb-Edu synthetic textbook approaches.1:24–4:23 · Guest teaching 4/10 The 48-Hour Sprint: Distilling DeepSeek R1 Swyx sets up the context of the 48-hour sprint following DeepSeek R1's release. Mahesh and Ryan detail the rapid iteration using their Curator tool to distill R1 into Qwen.4:23–6:47 · Guest teaching 6/10 Defining Model Distillation: From Logits to Synthetic Data Alessio asks for clarification on distillation terminology. Mahesh provides a clear technical breakdown ranging from Hinton's 2016 logit distillation to synthetic data fine-tuning.6:48–10:24 · Guest teaching 6/10 Open Reasoning Traces and Autoregressive Backtracking The conversation covers the opacity of OpenAI o1 traces versus open R1 traces. Trung and Mahesh explain how autoregressive generation performs implicit backtracking without separate search algorithms.10:25–15:08 · Guest teaching 4/10 Frontier Search Architectures vs. Distilled Small Models Swyx challenges the idea that search is entirely discarded, drawing distinctions between frontier training/inference architectures and distilled mini models while referencing PRMs and MCTS.15:09–17:56 · Guest teaching 6/10 Pipeline Differences: Coherence and Ablating Reannotation Swyx probes why Bespoke's 7B model succeeded where Sky-T1 failed. Trung explains their specific pipeline ablation of removing the GPT-4o-mini reannotation step because R1 reasoning traces were already coherent.17:58–21:08 · Guest teaching 5/10 Enabling Student Models to Surpass Teacher Models Swyx demonstrates deep background knowledge citing the STaR paper and Cosine's SWE-bench results, while Mahesh shares how their 7B MiniCheck model outperformed Llama 3 405B.21:09–23:14 · Guest teaching 4/10 Data Curation Principles and the Bespoke Curator Platform Mahesh outlines the three fundamental axes of data curation—quantity, quality, and diversity—while Swyx touches on FineWeb-Edu synthetic textbook approaches.1:24–4:23 · Guest disagreement 0/10 The 48-Hour Sprint: Distilling DeepSeek R1 Swyx sets up the context of the 48-hour sprint following DeepSeek R1's release. Mahesh and Ryan detail the rapid iteration using their Curator tool to distill R1 into Qwen.4:23–6:47 · Guest disagreement 0/10 Defining Model Distillation: From Logits to Synthetic Data Alessio asks for clarification on distillation terminology. Mahesh provides a clear technical breakdown ranging from Hinton's 2016 logit distillation to synthetic data fine-tuning.6:48–10:24 · Guest disagreement 0/10 Open Reasoning Traces and Autoregressive Backtracking The conversation covers the opacity of OpenAI o1 traces versus open R1 traces. Trung and Mahesh explain how autoregressive generation performs implicit backtracking without separate search algorithms.10:25–15:08 · Guest disagreement 1/10 Frontier Search Architectures vs. Distilled Small Models Swyx challenges the idea that search is entirely discarded, drawing distinctions between frontier training/inference architectures and distilled mini models while referencing PRMs and MCTS.15:09–17:56 · Guest disagreement 0/10 Pipeline Differences: Coherence and Ablating Reannotation Swyx probes why Bespoke's 7B model succeeded where Sky-T1 failed. Trung explains their specific pipeline ablation of removing the GPT-4o-mini reannotation step because R1 reasoning traces were already coherent.17:58–21:08 · Guest disagreement 0/10 Enabling Student Models to Surpass Teacher Models Swyx demonstrates deep background knowledge citing the STaR paper and Cosine's SWE-bench results, while Mahesh shares how their 7B MiniCheck model outperformed Llama 3 405B.21:09–23:14 · Guest disagreement 0/10 Data Curation Principles and the Bespoke Curator Platform Mahesh outlines the three fundamental axes of data curation—quantity, quality, and diversity—while Swyx touches on FineWeb-Edu synthetic textbook approaches.1:24–4:23 · The hosts pushing back 0/10 The 48-Hour Sprint: Distilling DeepSeek R1 Swyx sets up the context of the 48-hour sprint following DeepSeek R1's release. Mahesh and Ryan detail the rapid iteration using their Curator tool to distill R1 into Qwen.4:23–6:47 · The hosts pushing back 0/10 Defining Model Distillation: From Logits to Synthetic Data Alessio asks for clarification on distillation terminology. Mahesh provides a clear technical breakdown ranging from Hinton's 2016 logit distillation to synthetic data fine-tuning.6:48–10:24 · The hosts pushing back 0/10 Open Reasoning Traces and Autoregressive Backtracking The conversation covers the opacity of OpenAI o1 traces versus open R1 traces. Trung and Mahesh explain how autoregressive generation performs implicit backtracking without separate search algorithms.10:25–15:08 · The hosts pushing back 7/10 Frontier Search Architectures vs. Distilled Small Models Swyx challenges the idea that search is entirely discarded, drawing distinctions between frontier training/inference architectures and distilled mini models while referencing PRMs and MCTS.15:09–17:56 · The hosts pushing back 1/10 Pipeline Differences: Coherence and Ablating Reannotation Swyx probes why Bespoke's 7B model succeeded where Sky-T1 failed. Trung explains their specific pipeline ablation of removing the GPT-4o-mini reannotation step because R1 reasoning traces were already coherent.17:58–21:08 · The hosts pushing back 1/10 Enabling Student Models to Surpass Teacher Models Swyx demonstrates deep background knowledge citing the STaR paper and Cosine's SWE-bench results, while Mahesh shares how their 7B MiniCheck model outperformed Llama 3 405B.21:09–23:14 · The hosts pushing back 0/10 Data Curation Principles and the Bespoke Curator Platform Mahesh outlines the three fundamental axes of data curation—quantity, quality, and diversity—while Swyx touches on FineWeb-Edu synthetic textbook approaches.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 26.7% · guest 73.3%0:00 · the hosts 26.7% · guest 73.3%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 42% · guest 58%9:00 · the hosts 42% · guest 58%12:00 · the hosts 16.6% · guest 83.4%12:00 · the hosts 16.6% · guest 83.4%15:00 · the hosts 2.9% · guest 97.1%15:00 · the hosts 2.9% · guest 97.1%18:00 · the hosts 31.4% · guest 68.6%18:00 · the hosts 31.4% · guest 68.6%21:00 · the hosts 15.5% · guest 84.5%21:00 · the hosts 15.5% · guest 84.5%
Sharpest disagreement ▶ 11:18 Mahesh and Ryan qualifying inference search distinctions

In a very collegial episode, this moment shows the guests gently pushing back against overgeneralizing search assumptions between training and inference.

Hardest push from the hosts ▶ 10:25 Swyx challenging the no-search narrative

Swyx explicitly refuses the simple framing that search algorithms are dead, suggesting Q* could have been misunderstood and distinguishing frontier models from distilled minis.

Biggest teaching moment ▶ 9:29 Trung on emergent search and the Bitter Lesson

Trung educates the hosts on how R1 eliminated the need for complex driver search algorithms by learning implicit backtracking through reinforcement learning.

The host holds their own ▶ 19:18 Swyx demonstrating prior knowledge with STaR and SWE-bench

Swyx demonstrates formidable domain expertise by linking the interviewees' distillation findings back to foundational papers and Cosine's prior SWE-bench breakthroughs.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The 48-Hour Sprint: Distilling DeepSeek R1 5400 Swyx sets up the context of the 48-hour sprint following DeepSeek R1's release. Mahesh and Ryan detail the rapid iteration using their Curator tool to distill R1 into Qwen.
Defining Model Distillation: From Logits to Synthetic Data 4600 Alessio asks for clarification on distillation terminology. Mahesh provides a clear technical breakdown ranging from Hinton's 2016 logit distillation to synthetic data fine-tuning.
Open Reasoning Traces and Autoregressive Backtracking 4600 The conversation covers the opacity of OpenAI o1 traces versus open R1 traces. Trung and Mahesh explain how autoregressive generation performs implicit backtracking without separate search algorithms.
Frontier Search Architectures vs. Distilled Small Models 8417 Swyx challenges the idea that search is entirely discarded, drawing distinctions between frontier training/inference architectures and distilled mini models while referencing PRMs and MCTS.
Pipeline Differences: Coherence and Ablating Reannotation 5601 Swyx probes why Bespoke's 7B model succeeded where Sky-T1 failed. Trung explains their specific pipeline ablation of removing the GPT-4o-mini reannotation step because R1 reasoning traces were already coherent.
Enabling Student Models to Surpass Teacher Models 8501 Swyx demonstrates deep background knowledge citing the STaR paper and Cosine's SWE-bench results, while Mahesh shares how their 7B MiniCheck model outperformed Llama 3 405B.
Data Curation Principles and the Bespoke Curator Platform 5400 Mahesh outlines the three fundamental axes of data curation—quantity, quality, and diversity—while Swyx touches on FineWeb-Edu synthetic textbook approaches.

Statements from this episode (3)

Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Shawn Wang Jan 24, 2025 ▶ 10:32
Opinion
OpenAI likely reaches frontier capabilities via search, then distills into mini models
“The only way you reach the frontier with the full size models of O-one and O-three is with that stuff. And then you can distill to the minis, the O-one mini, O-three mini. So in my writeup, I said like, maybe this is the formula for O-one mini, O-three mini. T…”
Shawn Wang Jan 24, 2025 ▶ 11:05
Assertion Supported
Cosine's fine-tuned GPT-4o scored higher on SWE-bench than OpenAI's original o1
“And then we also on this podcast, we interviewed Cosign that actually fine tune four O to, on three bench, sorry to achieve data on three bench. And that score was actually higher than O one when it came out.”
Shawn Wang Jan 24, 2025 ▶ 19:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.