Oct 20, 2023 · 1h 24m · latent-space

The End of Finetuning — with Jeremy Howard of Fast.ai

Jeremy Howard · 1h 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Fast.ai co-founder Jeremy Howard discusses the evolution of deep learning, arguing that traditional fine-tuning must be replaced by continued pre-training while passionately advocating for open-source AI democratization against centralized gatekeeping.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.9 Guest teaching 5.2 Guest disagreement 3.9 The hosts pushing back 1.6
05100:0020:0040:001:00:001:20:000:11–4:07 · The hosts as informed peer 4/10 Introductions, Philosophy Background, and Early Ventures Swix introduces Jeremy Howard and touches on his unusual educational and professional trajectory, including his philosophy degree, McKinsey career, and early startups. The dynamic is cordial, biographical, and respectful with minimal tension.4:07–9:33 · The hosts as informed peer 5/10 The Founding and Democratizing Mission of Fast.ai Jeremy details the founding thesis of Fast.ai in 2016, pushing back against the elitist culture of deep learning requiring elite PhD credentials. Alessio frames the technical context of AWD-LSTM and parameter scaling.9:34–15:10 · The hosts as informed peer 4/10 The Chinese Room Experiment and Developing ULMFiT Jeremy provides an extensive conceptual lecture connecting Searle's Chinese Room argument to language models and the development of ULMFiT. He recounts how established NLP researchers insisted general pre-training and transfer learning were theoretically impossible.15:10–17:47 · The hosts as informed peer 5/10 How ULMFiT Influenced OpenAI's GPT Architecture Alessio asks why prior researchers missed general language pre-training, whether due to setup difficulty or complacency. Jeremy explains how Alec Radford directly credited ULMFiT for inspiring GPT-1 after an initial discussion.17:48–22:17 · The hosts as informed peer 6/10 The Shift to Zero-Shot Learning and the Big Iron Distraction Alessio cites Radford's original tweet acknowledging ULMFiT. Jeremy critiques the industry's subsequent multi-year obsession with zero-shot/few-shot prompting and 'big iron' compute, arguing it distracted from practical transfer learning.22:17–28:58 · The hosts as informed peer 4/10 Choosing Public Good and Democratization Over Tech Elitism Alessio asks why Jeremy prioritized open democratization over traditional startup fundraising after his early commercial exits. Jeremy explains his philosophical motivations to prevent powerful AI tech from concentrating among elites.28:58–35:37 · The hosts as informed peer 5/10 Fast.ai's Software Ecosystem and the DawnBench Triumph Swix summarizes the Fast.ai software ecosystem, prompting Jeremy to describe his research methodology. Jeremy explains beating Google's TPU team in Stanford's DawnBench competition using progressive resizing on commodity hardware.35:39–44:24 · The hosts as informed peer 5/10 Single-Epoch Memorization in LLMs and Catastrophic Forgetting Jeremy recounts discovering sudden loss drops ('clunks') at epoch boundaries during fine-tuning. He describes his frustration when open-source Discords dismissed the anomaly as normal rather than investigating single-epoch dataset memorization.44:24–48:38 · The hosts as informed peer 4/10 The Death of Fine-Tuning: Moving to Continued Pre-Training Jeremy boldly declares that his own invention of 3-step fine-tuning is obsolete and damaging modern models like Code Llama. He redefines the entire paradigm as continuous pre-training on diverse data mixtures to prevent catastrophic forgetting.48:38–54:46 · The hosts as informed peer 5/10 Navigating Open-Source AI Discords and Online Communities The hosts inquire about participating in elite open-source AI communities. Jeremy discusses the necessity of gated private channels to filter out low-effort speculation, emphasizing that genuine builders who ship work easily get access.54:46–1:02:46 · The hosts as informed peer 5/10 Chris Lattner, Swift for TensorFlow, and the Creation of Mojo Alessio brings up Modular and Mojo. Jeremy shares his history collaborating with Chris Lattner on Swift for TensorFlow, recognizing TensorFlow 2 as an inevitable failure, and advising Lattner to build an independent language ecosystem.1:02:48–1:09:27 · The hosts as informed peer 6/10 RAG Limitations, Fine-Tuning Knowledge, and Small Models Jeremy criticizes RAG as an inefficient hack. Swix challenges Jeremy on whether fine-tuning can truly incorporate new factual knowledge, prompting Jeremy to forcefully reject the distinction between fine-tuning and pre-training.1:09:28–1:17:06 · The hosts as informed peer 6/10 The Future of AI Education, Technical Debt, and Flash Attention Alessio and Jeremy discuss the future of AI education and system-level efficiencies. Alessio brings up Tri Dao's Flash Attention insights, while Jeremy highlights the enormous accumulated technical debt across current LLM development.1:17:08–1:24:21 · The hosts as informed peer 4/10 Lightning Round and the Imperative of Open AI Access In the lightning round, Jeremy discusses empirical training dynamics and delivers an impassioned closing argument for democratized access to AI over centralized elite control.0:11–4:07 · Guest teaching 1/10 Introductions, Philosophy Background, and Early Ventures Swix introduces Jeremy Howard and touches on his unusual educational and professional trajectory, including his philosophy degree, McKinsey career, and early startups. The dynamic is cordial, biographical, and respectful with minimal tension.4:07–9:33 · Guest teaching 4/10 The Founding and Democratizing Mission of Fast.ai Jeremy details the founding thesis of Fast.ai in 2016, pushing back against the elitist culture of deep learning requiring elite PhD credentials. Alessio frames the technical context of AWD-LSTM and parameter scaling.9:34–15:10 · Guest teaching 7/10 The Chinese Room Experiment and Developing ULMFiT Jeremy provides an extensive conceptual lecture connecting Searle's Chinese Room argument to language models and the development of ULMFiT. He recounts how established NLP researchers insisted general pre-training and transfer learning were theoretically impossible.15:10–17:47 · Guest teaching 6/10 How ULMFiT Influenced OpenAI's GPT Architecture Alessio asks why prior researchers missed general language pre-training, whether due to setup difficulty or complacency. Jeremy explains how Alec Radford directly credited ULMFiT for inspiring GPT-1 after an initial discussion.17:48–22:17 · Guest teaching 5/10 The Shift to Zero-Shot Learning and the Big Iron Distraction Alessio cites Radford's original tweet acknowledging ULMFiT. Jeremy critiques the industry's subsequent multi-year obsession with zero-shot/few-shot prompting and 'big iron' compute, arguing it distracted from practical transfer learning.22:17–28:58 · Guest teaching 3/10 Choosing Public Good and Democratization Over Tech Elitism Alessio asks why Jeremy prioritized open democratization over traditional startup fundraising after his early commercial exits. Jeremy explains his philosophical motivations to prevent powerful AI tech from concentrating among elites.28:58–35:37 · Guest teaching 5/10 Fast.ai's Software Ecosystem and the DawnBench Triumph Swix summarizes the Fast.ai software ecosystem, prompting Jeremy to describe his research methodology. Jeremy explains beating Google's TPU team in Stanford's DawnBench competition using progressive resizing on commodity hardware.35:39–44:24 · Guest teaching 7/10 Single-Epoch Memorization in LLMs and Catastrophic Forgetting Jeremy recounts discovering sudden loss drops ('clunks') at epoch boundaries during fine-tuning. He describes his frustration when open-source Discords dismissed the anomaly as normal rather than investigating single-epoch dataset memorization.44:24–48:38 · Guest teaching 8/10 The Death of Fine-Tuning: Moving to Continued Pre-Training Jeremy boldly declares that his own invention of 3-step fine-tuning is obsolete and damaging modern models like Code Llama. He redefines the entire paradigm as continuous pre-training on diverse data mixtures to prevent catastrophic forgetting.48:38–54:46 · Guest teaching 4/10 Navigating Open-Source AI Discords and Online Communities The hosts inquire about participating in elite open-source AI communities. Jeremy discusses the necessity of gated private channels to filter out low-effort speculation, emphasizing that genuine builders who ship work easily get access.54:46–1:02:46 · Guest teaching 6/10 Chris Lattner, Swift for TensorFlow, and the Creation of Mojo Alessio brings up Modular and Mojo. Jeremy shares his history collaborating with Chris Lattner on Swift for TensorFlow, recognizing TensorFlow 2 as an inevitable failure, and advising Lattner to build an independent language ecosystem.1:02:48–1:09:27 · Guest teaching 8/10 RAG Limitations, Fine-Tuning Knowledge, and Small Models Jeremy criticizes RAG as an inefficient hack. Swix challenges Jeremy on whether fine-tuning can truly incorporate new factual knowledge, prompting Jeremy to forcefully reject the distinction between fine-tuning and pre-training.1:09:28–1:17:06 · Guest teaching 5/10 The Future of AI Education, Technical Debt, and Flash Attention Alessio and Jeremy discuss the future of AI education and system-level efficiencies. Alessio brings up Tri Dao's Flash Attention insights, while Jeremy highlights the enormous accumulated technical debt across current LLM development.1:17:08–1:24:21 · Guest teaching 4/10 Lightning Round and the Imperative of Open AI Access In the lightning round, Jeremy discusses empirical training dynamics and delivers an impassioned closing argument for democratized access to AI over centralized elite control.0:11–4:07 · Guest disagreement 1/10 Introductions, Philosophy Background, and Early Ventures Swix introduces Jeremy Howard and touches on his unusual educational and professional trajectory, including his philosophy degree, McKinsey career, and early startups. The dynamic is cordial, biographical, and respectful with minimal tension.4:07–9:33 · Guest disagreement 3/10 The Founding and Democratizing Mission of Fast.ai Jeremy details the founding thesis of Fast.ai in 2016, pushing back against the elitist culture of deep learning requiring elite PhD credentials. Alessio frames the technical context of AWD-LSTM and parameter scaling.9:34–15:10 · Guest disagreement 4/10 The Chinese Room Experiment and Developing ULMFiT Jeremy provides an extensive conceptual lecture connecting Searle's Chinese Room argument to language models and the development of ULMFiT. He recounts how established NLP researchers insisted general pre-training and transfer learning were theoretically impossible.15:10–17:47 · Guest disagreement 3/10 How ULMFiT Influenced OpenAI's GPT Architecture Alessio asks why prior researchers missed general language pre-training, whether due to setup difficulty or complacency. Jeremy explains how Alec Radford directly credited ULMFiT for inspiring GPT-1 after an initial discussion.17:48–22:17 · Guest disagreement 5/10 The Shift to Zero-Shot Learning and the Big Iron Distraction Alessio cites Radford's original tweet acknowledging ULMFiT. Jeremy critiques the industry's subsequent multi-year obsession with zero-shot/few-shot prompting and 'big iron' compute, arguing it distracted from practical transfer learning.22:17–28:58 · Guest disagreement 3/10 Choosing Public Good and Democratization Over Tech Elitism Alessio asks why Jeremy prioritized open democratization over traditional startup fundraising after his early commercial exits. Jeremy explains his philosophical motivations to prevent powerful AI tech from concentrating among elites.28:58–35:37 · Guest disagreement 4/10 Fast.ai's Software Ecosystem and the DawnBench Triumph Swix summarizes the Fast.ai software ecosystem, prompting Jeremy to describe his research methodology. Jeremy explains beating Google's TPU team in Stanford's DawnBench competition using progressive resizing on commodity hardware.35:39–44:24 · Guest disagreement 5/10 Single-Epoch Memorization in LLMs and Catastrophic Forgetting Jeremy recounts discovering sudden loss drops ('clunks') at epoch boundaries during fine-tuning. He describes his frustration when open-source Discords dismissed the anomaly as normal rather than investigating single-epoch dataset memorization.44:24–48:38 · Guest disagreement 6/10 The Death of Fine-Tuning: Moving to Continued Pre-Training Jeremy boldly declares that his own invention of 3-step fine-tuning is obsolete and damaging modern models like Code Llama. He redefines the entire paradigm as continuous pre-training on diverse data mixtures to prevent catastrophic forgetting.48:38–54:46 · Guest disagreement 3/10 Navigating Open-Source AI Discords and Online Communities The hosts inquire about participating in elite open-source AI communities. Jeremy discusses the necessity of gated private channels to filter out low-effort speculation, emphasizing that genuine builders who ship work easily get access.54:46–1:02:46 · Guest disagreement 4/10 Chris Lattner, Swift for TensorFlow, and the Creation of Mojo Alessio brings up Modular and Mojo. Jeremy shares his history collaborating with Chris Lattner on Swift for TensorFlow, recognizing TensorFlow 2 as an inevitable failure, and advising Lattner to build an independent language ecosystem.1:02:48–1:09:27 · Guest disagreement 7/10 RAG Limitations, Fine-Tuning Knowledge, and Small Models Jeremy criticizes RAG as an inefficient hack. Swix challenges Jeremy on whether fine-tuning can truly incorporate new factual knowledge, prompting Jeremy to forcefully reject the distinction between fine-tuning and pre-training.1:09:28–1:17:06 · Guest disagreement 4/10 The Future of AI Education, Technical Debt, and Flash Attention Alessio and Jeremy discuss the future of AI education and system-level efficiencies. Alessio brings up Tri Dao's Flash Attention insights, while Jeremy highlights the enormous accumulated technical debt across current LLM development.1:17:08–1:24:21 · Guest disagreement 3/10 Lightning Round and the Imperative of Open AI Access In the lightning round, Jeremy discusses empirical training dynamics and delivers an impassioned closing argument for democratized access to AI over centralized elite control.0:11–4:07 · The hosts pushing back 1/10 Introductions, Philosophy Background, and Early Ventures Swix introduces Jeremy Howard and touches on his unusual educational and professional trajectory, including his philosophy degree, McKinsey career, and early startups. The dynamic is cordial, biographical, and respectful with minimal tension.4:07–9:33 · The hosts pushing back 1/10 The Founding and Democratizing Mission of Fast.ai Jeremy details the founding thesis of Fast.ai in 2016, pushing back against the elitist culture of deep learning requiring elite PhD credentials. Alessio frames the technical context of AWD-LSTM and parameter scaling.9:34–15:10 · The hosts pushing back 1/10 The Chinese Room Experiment and Developing ULMFiT Jeremy provides an extensive conceptual lecture connecting Searle's Chinese Room argument to language models and the development of ULMFiT. He recounts how established NLP researchers insisted general pre-training and transfer learning were theoretically impossible.15:10–17:47 · The hosts pushing back 2/10 How ULMFiT Influenced OpenAI's GPT Architecture Alessio asks why prior researchers missed general language pre-training, whether due to setup difficulty or complacency. Jeremy explains how Alec Radford directly credited ULMFiT for inspiring GPT-1 after an initial discussion.17:48–22:17 · The hosts pushing back 1/10 The Shift to Zero-Shot Learning and the Big Iron Distraction Alessio cites Radford's original tweet acknowledging ULMFiT. Jeremy critiques the industry's subsequent multi-year obsession with zero-shot/few-shot prompting and 'big iron' compute, arguing it distracted from practical transfer learning.22:17–28:58 · The hosts pushing back 1/10 Choosing Public Good and Democratization Over Tech Elitism Alessio asks why Jeremy prioritized open democratization over traditional startup fundraising after his early commercial exits. Jeremy explains his philosophical motivations to prevent powerful AI tech from concentrating among elites.28:58–35:37 · The hosts pushing back 1/10 Fast.ai's Software Ecosystem and the DawnBench Triumph Swix summarizes the Fast.ai software ecosystem, prompting Jeremy to describe his research methodology. Jeremy explains beating Google's TPU team in Stanford's DawnBench competition using progressive resizing on commodity hardware.35:39–44:24 · The hosts pushing back 2/10 Single-Epoch Memorization in LLMs and Catastrophic Forgetting Jeremy recounts discovering sudden loss drops ('clunks') at epoch boundaries during fine-tuning. He describes his frustration when open-source Discords dismissed the anomaly as normal rather than investigating single-epoch dataset memorization.44:24–48:38 · The hosts pushing back 1/10 The Death of Fine-Tuning: Moving to Continued Pre-Training Jeremy boldly declares that his own invention of 3-step fine-tuning is obsolete and damaging modern models like Code Llama. He redefines the entire paradigm as continuous pre-training on diverse data mixtures to prevent catastrophic forgetting.48:38–54:46 · The hosts pushing back 2/10 Navigating Open-Source AI Discords and Online Communities The hosts inquire about participating in elite open-source AI communities. Jeremy discusses the necessity of gated private channels to filter out low-effort speculation, emphasizing that genuine builders who ship work easily get access.54:46–1:02:46 · The hosts pushing back 2/10 Chris Lattner, Swift for TensorFlow, and the Creation of Mojo Alessio brings up Modular and Mojo. Jeremy shares his history collaborating with Chris Lattner on Swift for TensorFlow, recognizing TensorFlow 2 as an inevitable failure, and advising Lattner to build an independent language ecosystem.1:02:48–1:09:27 · The hosts pushing back 5/10 RAG Limitations, Fine-Tuning Knowledge, and Small Models Jeremy criticizes RAG as an inefficient hack. Swix challenges Jeremy on whether fine-tuning can truly incorporate new factual knowledge, prompting Jeremy to forcefully reject the distinction between fine-tuning and pre-training.1:09:28–1:17:06 · The hosts pushing back 2/10 The Future of AI Education, Technical Debt, and Flash Attention Alessio and Jeremy discuss the future of AI education and system-level efficiencies. Alessio brings up Tri Dao's Flash Attention insights, while Jeremy highlights the enormous accumulated technical debt across current LLM development.1:17:08–1:24:21 · The hosts pushing back 1/10 Lightning Round and the Imperative of Open AI Access In the lightning round, Jeremy discusses empirical training dynamics and delivers an impassioned closing argument for democratized access to AI over centralized elite control.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:05:18 Vehement rejection of RAG superiority over fine-tuning knowledge

Jeremy immediately shuts down the skeptical framing regarding fine-tuning's inability to absorb new knowledge, emphatically declaring that fine-tuning is literally identical to continued pre-training.

Hardest push from the hosts ▶ 1:05:01 Swix presses Jeremy on whether models can actually internalize new facts

Swix directly challenges Jeremy's dismissal of RAG by voicing community skepticism about whether fine-tuning can genuinely encode new information without hallucinations.

Biggest teaching moment ▶ 44:24 Deconstructing modern fine-tuning and catastrophic forgetting

Jeremy dismantles the prevailing industry consensus on fine-tuning, explaining how his own pioneer paper ULMFiT is now misused and why Code Llama broke by ignoring data mixing.

The host holds their own ▶ 17:48 Alessio cites Alec Radford's 2018 acknowledgment of ULMFiT

Alessio demonstrates sharp domain command by digging up and quoting Alec Radford's exact 2018 launch tweet proving OpenAI's direct lineage from Jeremy's paper.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions, Philosophy Background, and Early Ventures 4111 Swix introduces Jeremy Howard and touches on his unusual educational and professional trajectory, including his philosophy degree, McKinsey career, and early startups. The dynamic is cordial, biographical, and respectful with minimal tension.
The Founding and Democratizing Mission of Fast.ai 5431 Jeremy details the founding thesis of Fast.ai in 2016, pushing back against the elitist culture of deep learning requiring elite PhD credentials. Alessio frames the technical context of AWD-LSTM and parameter scaling.
The Chinese Room Experiment and Developing ULMFiT 4741 Jeremy provides an extensive conceptual lecture connecting Searle's Chinese Room argument to language models and the development of ULMFiT. He recounts how established NLP researchers insisted general pre-training and transfer learning were theoretically impossible.
How ULMFiT Influenced OpenAI's GPT Architecture 5632 Alessio asks why prior researchers missed general language pre-training, whether due to setup difficulty or complacency. Jeremy explains how Alec Radford directly credited ULMFiT for inspiring GPT-1 after an initial discussion.
The Shift to Zero-Shot Learning and the Big Iron Distraction 6551 Alessio cites Radford's original tweet acknowledging ULMFiT. Jeremy critiques the industry's subsequent multi-year obsession with zero-shot/few-shot prompting and 'big iron' compute, arguing it distracted from practical transfer learning.
Choosing Public Good and Democratization Over Tech Elitism 4331 Alessio asks why Jeremy prioritized open democratization over traditional startup fundraising after his early commercial exits. Jeremy explains his philosophical motivations to prevent powerful AI tech from concentrating among elites.
Fast.ai's Software Ecosystem and the DawnBench Triumph 5541 Swix summarizes the Fast.ai software ecosystem, prompting Jeremy to describe his research methodology. Jeremy explains beating Google's TPU team in Stanford's DawnBench competition using progressive resizing on commodity hardware.
Single-Epoch Memorization in LLMs and Catastrophic Forgetting 5752 Jeremy recounts discovering sudden loss drops ('clunks') at epoch boundaries during fine-tuning. He describes his frustration when open-source Discords dismissed the anomaly as normal rather than investigating single-epoch dataset memorization.
The Death of Fine-Tuning: Moving to Continued Pre-Training 4861 Jeremy boldly declares that his own invention of 3-step fine-tuning is obsolete and damaging modern models like Code Llama. He redefines the entire paradigm as continuous pre-training on diverse data mixtures to prevent catastrophic forgetting.
Navigating Open-Source AI Discords and Online Communities 5432 The hosts inquire about participating in elite open-source AI communities. Jeremy discusses the necessity of gated private channels to filter out low-effort speculation, emphasizing that genuine builders who ship work easily get access.
Chris Lattner, Swift for TensorFlow, and the Creation of Mojo 5642 Alessio brings up Modular and Mojo. Jeremy shares his history collaborating with Chris Lattner on Swift for TensorFlow, recognizing TensorFlow 2 as an inevitable failure, and advising Lattner to build an independent language ecosystem.
RAG Limitations, Fine-Tuning Knowledge, and Small Models 6875 Jeremy criticizes RAG as an inefficient hack. Swix challenges Jeremy on whether fine-tuning can truly incorporate new factual knowledge, prompting Jeremy to forcefully reject the distinction between fine-tuning and pre-training.
The Future of AI Education, Technical Debt, and Flash Attention 6542 Alessio and Jeremy discuss the future of AI education and system-level efficiencies. Alessio brings up Tri Dao's Flash Attention insights, while Jeremy highlights the enormous accumulated technical debt across current LLM development.
Lightning Round and the Imperative of Open AI Access 4431 In the lightning round, Jeremy discusses empirical training dynamics and delivers an impassioned closing argument for democratized access to AI over centralized elite control.

Statements from this episode (25)

Assertion Supported
Howard: Enlitic was the first company to use deep learning in medicine
“Yeah, I was actually the first company to use deep learning in medicine, so I kind of founded the field.”
Jeremy Howard Oct 20, 2023 ▶ 3:28
Insight
Howard: Transfer learning reduces deep learning compute and data needs
“There's this thing which nobody knows about, nobody talks about, called transfer learning, where you take somebody else's model where they already figured out, like, how to Detect edges, and gradients, and corners, and text, and whatever else, and then you can…”
Jeremy Howard Oct 20, 2023 ▶ 7:53
Insight
Howard: Accurate next-token prediction forces models to learn world models and causality
“I thought, okay, so if I do this at a much bigger scale, using all of Wikipedia, what would it need to be able to do to finish a sentence in Wikipedia effectively, to do it quite accurately, quite often? I thought, geez, it would actually have to know a lot ab…”
Jeremy Howard Oct 20, 2023 ▶ 12:37
Assertion Partly supported
Howard: Modern LLMs like ChatGPT still follow ULMFiT's three-step training framework
“So we generated this three-step system. So step one was train a language model on a big corpus. Step two was fine-tune a language model on a more curated corpus, and step three was further fine-tune that model on a task. And of course, that's what everybody st…”
Jeremy Howard Oct 20, 2023 ▶ 14:04
Assertion Not checkable as stated
Howard: NLP researchers unanimously insisted transfer learning would never work for language
“Like, the reason it took me so long to try it was because I asked all my friends in NLP if this could work, and everybody said, no, it definitely won't work. It wasn't like, oh, maybe. Everybody was like, it definitely won't work. NLP is much more complicated …”
Jeremy Howard Oct 20, 2023 ▶ 14:41
Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Jeremy Howard Oct 20, 2023 ▶ 15:41
Opinion
Howard: AI research went backwards for years pursuing zero-shot learning
“And so I actually feel like we kind of went backwards for years and not to be honest, I mean, I'm a bit sad about this now, but I kind of got so disappointed and dissuaded by like, It felt like these bigger lab, much bigger labs, you know, like fast.ai had onl…”
Jeremy Howard Oct 20, 2023 ▶ 19:29
Insight
Howard: Focusing on massive compute infrastructure is a distraction in AI
“It might seem like a recent thing, but actually throughout my 30 years in data science, the attention's always been on, you know, the big iron results. So when I first started, everybody was talking about data warehouses, and it was all about Teradata, and it'…”
Jeremy Howard Oct 20, 2023 ▶ 21:22
Assertion Not checkable as stated
Howard: Tesla and OpenAI Scholars used Fast.ai courses for deep learning training
“Andre Capathy grabbed me when I saw him at NeurIpes a few years ago, and he's like, I have to tell you, thanks to the fast AI courses, when people come to Tesla, and they need to know more about deep learning, we always send them to your course. And the OpenAI…”
Jeremy Howard Oct 20, 2023 ▶ 28:15
Disclosure
Howard: Fast.ai was never more than two people and is now solo
“That's just one of, like, three things we do, is the course, you know, and it's only ever been at most two people, either me and Rachel, or me and Sylvain. Nowadays it's just me.”
Jeremy Howard Oct 20, 2023 ▶ 28:31
Insight
Jeremy Howard: AI research artifacts should be software and courses, not papers
“To me the main artifact shouldn't be papers, because papers are things read by a small exclusive group of people, you know, to me the main artifacts should be, like, something teaching you people, here's how to use this insight, and here's software you can use…”
Jeremy Howard Oct 20, 2023 ▶ 32:00
Assertion Partly supported
Jeremy Howard: Fast.ai won Stanford's DawnBench in 10 days using progressive resizing
“We only found out about this 10 days before the competition finished but, you know, we basically got together an emergency bunch of our students, and Rachel and I, and sat for the next 10 days, and just tried to crunch through, and Try to use all of our best i…”
Jeremy Howard Oct 20, 2023 ▶ 34:21
Assertion Supported
Howard: Experiments show LLMs can memorize full datasets in one epoch
“And so we ran a bunch of experiments, and all of them supported the hypothesis that it was memorizing the data set in a single thing at once.”
Jeremy Howard Oct 20, 2023 ▶ 41:14
Insight
Howard: Practitioners should track validation accuracy over validation loss
“Don't, I keep telling people, don't track validation loss, track validation accuracy because at least that, that will still be useful.”
Jeremy Howard Oct 20, 2023 ▶ 42:23
Opinion
Howard: Meta 'blew it' on Code Llama due to catastrophic forgetting
“So Code Llama was a, I think it was like a five hundred billion token fine-tuning of Llama II using code. And also prose about code that Meta did. And honestly, they kind of blew it. Because Code Llama is good at coding, but it's bad at everything else.”
Jeremy Howard Oct 20, 2023 ▶ 43:26
Opinion
Jeremy Howard: The three-step ULMFiT fine-tuning approach is wrong and obsolete
“Even though I originally created the three-step approach that everybody now does, my view is it's actually wrong, and we shouldn't use it.”
Jeremy Howard Oct 20, 2023 ▶ 44:27
Insight
Howard: There is no fine-tuning, only continued pre-training
“To me, the right way to do this is to fine, fine-tune language models, is to actually throw away the idea of fine-tuning. There's no such thing. There's only continued pre-training.”
Jeremy Howard Oct 20, 2023 ▶ 45:27
Insight
Alignment tax happens because models are fine-tuned instead of continued pre-trained
“ULM fit is the wrong approach and that's why we're seeing a lot of these You know, so-called alignment tax, and this view of like, oh, a model can't both code and do other things. You know, I think it's actually because people are training them wrong.”
Jeremy Howard Oct 20, 2023 ▶ 46:32
Assertion Not checkable as stated
Howard: Open-source AI Discords host substantive conversations in private channels
“Nearly all the discords, nearly all of the conversation happens in private channels”
Jeremy Howard Oct 20, 2023 ▶ 49:23
Opinion
Howard: TensorFlow 2 was a failure that Google avoided internally
“Then in the end, you know, Google didn't follow through, which is fair enough, like, asking everybody to, you know, learn a new programming language is going to be tough, but, like, it was very obvious, very, very obvious at that time that TensorFlow II was go…”
Jeremy Howard Oct 20, 2023 ▶ 59:34
Assertion Not checkable as stated
Howard: JAX was a grassroots Google reaction against TensorFlow 2
“But I mean, in the meantime, I will say, you know, Google now does have a backup plan. You know, they have JAX, which was never a strategy. It was just a bunch of people who also recognized TensorFlow two as shit, and they just decided to build something else.”
Jeremy Howard Oct 20, 2023 ▶ 1:01:58
Opinion
Howard: RAG is an inefficient hack compared to fine-tuning
“RAG is like such a inefficient hack, really, isn't it? It's like, You know, segment up my data in some somewhat arbitrary way, embed it, ask questions about that, you know, hope that my embedding, you know, model embeds questions in the same embedding space as…”
Jeremy Howard Oct 20, 2023 ▶ 1:03:36
Assertion Supported
Howard: Microsoft's phi-1.5 lacks world knowledge due to synthetic training data
“Fi-one-point-five has never read Wikipedia, for example, so it doesn't know who Tom Cruise is, you know it doesn't know who anybody is, he doesn't know about any movies, it doesn't really know anything about anything, like, because it was never, it's never rea…”
Jeremy Howard Oct 20, 2023 ▶ 1:07:21
Insight
Howard: Rapid LLM race created massive technical debt and optimization opportunities
“There's a whole lot of technical debt everywhere, you know, nobody's really figured this stuff out because everybody's been so busy building what we know works as quickly as possible. So, yeah, I think there's a huge amount of opportunity to, you know, I think…”
Jeremy Howard Oct 20, 2023 ▶ 1:14:48
Prediction Not checkable as stated
Howard: Mojo-like languages will unlock thousands of FlashAttention-scale breakthroughs
“There is a thousand flash attentions out there for us to build. You just got to make it easy for us to build them. So like Triton definitely helps. But it's still, Not easy. You know, it still requires kind of really understanding the VPU architecture, writing…”
Jeremy Howard Oct 20, 2023 ▶ 1:16:28
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.