Oct 20, 2023 · 1h 24m · latent-space
The End of Finetuning — with Jeremy Howard of Fast.ai
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Fast.ai co-founder Jeremy Howard discusses the evolution of deep learning, arguing that traditional fine-tuning must be replaced by continued pre-training while passionately advocating for open-source AI democratization against centralized gatekeeping.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jeremy immediately shuts down the skeptical framing regarding fine-tuning's inability to absorb new knowledge, emphatically declaring that fine-tuning is literally identical to continued pre-training.
Hardest push from the hosts ▶ 1:05:01 Swix presses Jeremy on whether models can actually internalize new factsSwix directly challenges Jeremy's dismissal of RAG by voicing community skepticism about whether fine-tuning can genuinely encode new information without hallucinations.
Biggest teaching moment ▶ 44:24 Deconstructing modern fine-tuning and catastrophic forgettingJeremy dismantles the prevailing industry consensus on fine-tuning, explaining how his own pioneer paper ULMFiT is now misused and why Code Llama broke by ignoring data mixing.
The host holds their own ▶ 17:48 Alessio cites Alec Radford's 2018 acknowledgment of ULMFiTAlessio demonstrates sharp domain command by digging up and quoting Alec Radford's exact 2018 launch tweet proving OpenAI's direct lineage from Jeremy's paper.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Introductions, Philosophy Background, and Early Ventures | 4 | 1 | 1 | 1 | Swix introduces Jeremy Howard and touches on his unusual educational and professional trajectory, including his philosophy degree, McKinsey career, and early startups. The dynamic is cordial, biographical, and respectful with minimal tension. | |
| The Founding and Democratizing Mission of Fast.ai | 5 | 4 | 3 | 1 | Jeremy details the founding thesis of Fast.ai in 2016, pushing back against the elitist culture of deep learning requiring elite PhD credentials. Alessio frames the technical context of AWD-LSTM and parameter scaling. | |
| The Chinese Room Experiment and Developing ULMFiT | 4 | 7 | 4 | 1 | Jeremy provides an extensive conceptual lecture connecting Searle's Chinese Room argument to language models and the development of ULMFiT. He recounts how established NLP researchers insisted general pre-training and transfer learning were theoretically impossible. | |
| How ULMFiT Influenced OpenAI's GPT Architecture | 5 | 6 | 3 | 2 | Alessio asks why prior researchers missed general language pre-training, whether due to setup difficulty or complacency. Jeremy explains how Alec Radford directly credited ULMFiT for inspiring GPT-1 after an initial discussion. | |
| The Shift to Zero-Shot Learning and the Big Iron Distraction | 6 | 5 | 5 | 1 | Alessio cites Radford's original tweet acknowledging ULMFiT. Jeremy critiques the industry's subsequent multi-year obsession with zero-shot/few-shot prompting and 'big iron' compute, arguing it distracted from practical transfer learning. | |
| Choosing Public Good and Democratization Over Tech Elitism | 4 | 3 | 3 | 1 | Alessio asks why Jeremy prioritized open democratization over traditional startup fundraising after his early commercial exits. Jeremy explains his philosophical motivations to prevent powerful AI tech from concentrating among elites. | |
| Fast.ai's Software Ecosystem and the DawnBench Triumph | 5 | 5 | 4 | 1 | Swix summarizes the Fast.ai software ecosystem, prompting Jeremy to describe his research methodology. Jeremy explains beating Google's TPU team in Stanford's DawnBench competition using progressive resizing on commodity hardware. | |
| Single-Epoch Memorization in LLMs and Catastrophic Forgetting | 5 | 7 | 5 | 2 | Jeremy recounts discovering sudden loss drops ('clunks') at epoch boundaries during fine-tuning. He describes his frustration when open-source Discords dismissed the anomaly as normal rather than investigating single-epoch dataset memorization. | |
| The Death of Fine-Tuning: Moving to Continued Pre-Training | 4 | 8 | 6 | 1 | Jeremy boldly declares that his own invention of 3-step fine-tuning is obsolete and damaging modern models like Code Llama. He redefines the entire paradigm as continuous pre-training on diverse data mixtures to prevent catastrophic forgetting. | |
| Navigating Open-Source AI Discords and Online Communities | 5 | 4 | 3 | 2 | The hosts inquire about participating in elite open-source AI communities. Jeremy discusses the necessity of gated private channels to filter out low-effort speculation, emphasizing that genuine builders who ship work easily get access. | |
| Chris Lattner, Swift for TensorFlow, and the Creation of Mojo | 5 | 6 | 4 | 2 | Alessio brings up Modular and Mojo. Jeremy shares his history collaborating with Chris Lattner on Swift for TensorFlow, recognizing TensorFlow 2 as an inevitable failure, and advising Lattner to build an independent language ecosystem. | |
| RAG Limitations, Fine-Tuning Knowledge, and Small Models | 6 | 8 | 7 | 5 | Jeremy criticizes RAG as an inefficient hack. Swix challenges Jeremy on whether fine-tuning can truly incorporate new factual knowledge, prompting Jeremy to forcefully reject the distinction between fine-tuning and pre-training. | |
| The Future of AI Education, Technical Debt, and Flash Attention | 6 | 5 | 4 | 2 | Alessio and Jeremy discuss the future of AI education and system-level efficiencies. Alessio brings up Tri Dao's Flash Attention insights, while Jeremy highlights the enormous accumulated technical debt across current LLM development. | |
| Lightning Round and the Imperative of Open AI Access | 4 | 4 | 3 | 1 | In the lightning round, Jeremy discusses empirical training dynamics and delivers an impassioned closing argument for democratized access to AI over centralized elite control. |