Opinion certainty 4/5 debate potential 4/5

Howard: RAG is an inefficient hack compared to fine-tuning

Jeremy Howard · The End of Finetuning — with Jeremy Howard of Fast.ai · Oct 20, 2023 · at 1:03:36

Fast.ai co-founder Jeremy Howard discusses AI developer workflows and critiques the limitations of retrieval-augmented generation (RAG).

0:00 / 0:45exact quote · 45.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“RAG is like such a inefficient hack, really, isn't it? It's like, You know, segment up my data in some somewhat arbitrary way, embed it, ask questions about that, you know, hope that my embedding, you know, model embeds questions in the same embedding space as the paragraphs, which obviously is not going to, if your question is like, if I've got a whole bunch of archive papers, embeddings, and I asked like, what are all the ways in which we can make inference more efficient? Like, the only paragraphs it'll find is like if there's a review paper that says here's a list of ways to make, you know inference more efficient.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Jeremy Howard

Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Jeremy Howard Oct 20, 2023 ▶ 15:41 The End of Finetuning — with Jeremy Howard of Fast.ai
Opinion
Howard: Meta 'blew it' on Code Llama due to catastrophic forgetting
“So Code Llama was a, I think it was like a five hundred billion token fine-tuning of Llama II using code. And also prose about code that Meta did. And honestly, they kind of blew it. Because Code Llama is good at coding, but it's bad at everything else.”
Jeremy Howard Oct 20, 2023 ▶ 43:26 The End of Finetuning — with Jeremy Howard of Fast.ai
Opinion
Howard: TensorFlow 2 was a failure that Google avoided internally
“Then in the end, you know, Google didn't follow through, which is fair enough, like, asking everybody to, you know, learn a new programming language is going to be tough, but, like, it was very obvious, very, very obvious at that time that TensorFlow II was go…”
Jeremy Howard Oct 20, 2023 ▶ 59:34 The End of Finetuning — with Jeremy Howard of Fast.ai
Assertion Not checkable as stated
Howard: JAX was a grassroots Google reaction against TensorFlow 2
“But I mean, in the meantime, I will say, you know, Google now does have a backup plan. You know, they have JAX, which was never a strategy. It was just a bunch of people who also recognized TensorFlow two as shit, and they just decided to build something else.”
Jeremy Howard Oct 20, 2023 ▶ 1:01:58 The End of Finetuning — with Jeremy Howard of Fast.ai
Assertion Not checkable as stated
Howard: Answer.AI runs fully in-house stack without AWS or Google Cloud
“This group of, which has averaged about 10 to 12 people, currently nine, I think, have built a pretty Transformational and complex piece of software, which we can do a quick demo of later if you're interested. Using a complete web application development platf…”
Jeremy Howard Oct 2, 2025 ▶ 4:28 The antidote to AI fatigue — Answer.ai Solveit
Insight
Howard: Correcting LLM errors in chat history degrades subsequent model answers
“The autoregressive nature of language models means that if they make a mistake, and you correct it, and then say, no, that was a mistake, please do it this way instead. The more often you do that, the worse the dialogue answers get. Because it's in the trainin…”
Jeremy Howard Oct 2, 2025 ▶ 16:17 The antidote to AI fatigue — Answer.ai Solveit
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.