Jan 24, 2025 · 23m · latent-space
The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Latent Space Podcast episode, the Bespoke Labs team discusses their 48-hour sprint distilling DeepSeek-R1 into Qwen, demonstrating how open reasoning traces, streamlined data curation, and targeted post-training allow compact 7B models to rival frontier reasoning systems.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.8% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
In a very collegial episode, this moment shows the guests gently pushing back against overgeneralizing search assumptions between training and inference.
Hardest push from the hosts ▶ 10:25 Swyx challenging the no-search narrativeSwyx explicitly refuses the simple framing that search algorithms are dead, suggesting Q* could have been misunderstood and distinguishing frontier models from distilled minis.
Biggest teaching moment ▶ 9:29 Trung on emergent search and the Bitter LessonTrung educates the hosts on how R1 eliminated the need for complex driver search algorithms by learning implicit backtracking through reinforcement learning.
The host holds their own ▶ 19:18 Swyx demonstrating prior knowledge with STaR and SWE-benchSwyx demonstrates formidable domain expertise by linking the interviewees' distillation findings back to foundational papers and Cosine's prior SWE-bench breakthroughs.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| The 48-Hour Sprint: Distilling DeepSeek R1 | 5 | 4 | 0 | 0 | Swyx sets up the context of the 48-hour sprint following DeepSeek R1's release. Mahesh and Ryan detail the rapid iteration using their Curator tool to distill R1 into Qwen. | |
| Defining Model Distillation: From Logits to Synthetic Data | 4 | 6 | 0 | 0 | Alessio asks for clarification on distillation terminology. Mahesh provides a clear technical breakdown ranging from Hinton's 2016 logit distillation to synthetic data fine-tuning. | |
| Open Reasoning Traces and Autoregressive Backtracking | 4 | 6 | 0 | 0 | The conversation covers the opacity of OpenAI o1 traces versus open R1 traces. Trung and Mahesh explain how autoregressive generation performs implicit backtracking without separate search algorithms. | |
| Frontier Search Architectures vs. Distilled Small Models | 8 | 4 | 1 | 7 | Swyx challenges the idea that search is entirely discarded, drawing distinctions between frontier training/inference architectures and distilled mini models while referencing PRMs and MCTS. | |
| Pipeline Differences: Coherence and Ablating Reannotation | 5 | 6 | 0 | 1 | Swyx probes why Bespoke's 7B model succeeded where Sky-T1 failed. Trung explains their specific pipeline ablation of removing the GPT-4o-mini reannotation step because R1 reasoning traces were already coherent. | |
| Enabling Student Models to Surpass Teacher Models | 8 | 5 | 0 | 1 | Swyx demonstrates deep background knowledge citing the STaR paper and Cosine's SWE-bench results, while Mahesh shares how their 7B MiniCheck model outperformed Llama 3 405B. | |
| Data Curation Principles and the Bespoke Curator Platform | 5 | 4 | 0 | 0 | Mahesh outlines the three fundamental axes of data curation—quantity, quality, and diversity—while Swyx touches on FineWeb-Edu synthetic textbook approaches. |