Howard: RAG is an inefficient hack compared to fine-tuning
Jeremy Howard · The End of Finetuning — with Jeremy Howard of Fast.ai · Oct 20, 2023 · at 1:03:36
Fast.ai co-founder Jeremy Howard discusses AI developer workflows and critiques the limitations of retrieval-augmented generation (RAG).
“RAG is like such a inefficient hack, really, isn't it? It's like, You know, segment up my data in some somewhat arbitrary way, embed it, ask questions about that, you know, hope that my embedding, you know, model embeds questions in the same embedding space as the paragraphs, which obviously is not going to, if your question is like, if I've got a whole bunch of archive papers, embeddings, and I asked like, what are all the ways in which we can make inference more efficient? Like, the only paragraphs it'll find is like if there's a review paper that says here's a list of ways to make, you know inference more efficient.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →