Oct 12, 2023 · 1h 13m · latent-space

RAG is a hack - with Jerry Liu of LlamaIndex

Jerry Liu · 53m spoken Shawn Wang · 9m spoken Alessio Fanelli · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space podcast, LlamaIndex creator Jerry Liu discusses the origins, architecture, and evolution of LlamaIndex, breaks down the technical trade-offs between RAG and fine-tuning, and shares critical strategies for building and evaluating production-grade LLM applications.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.4% of the talking time here. How this is scored →

The hosts as informed peer 5.3 Guest teaching 5.1 Guest disagreement 2.3 The hosts pushing back 2.8
05100:0015:0030:0045:001:00:005:29–11:42 · The hosts as informed peer 4/10 The Origins and Evolution of GPT Tree Index The hosts inquire about Jerry's origin story at Robust Intelligence and the transition from GPT Tree Index to LlamaIndex. Jerry explains the conceptual limitations of bottom-up tree structures and why early GPT-3 was insufficient for multi-hop reasoning.11:42–20:58 · The hosts as informed peer 5/10 Founding LlamaIndex, Moats, and Company Growth Alessio asks whether expanding context windows eliminate the need for data structuring. Jerry breaks down the economic and latency curves of network transfer costs over petabytes of enterprise data versus targeted retrieval.21:00–34:42 · The hosts as informed peer 6/10 The Core Debate: Why RAG is a Hack vs Fine-Tuning Jerry delivers his core thesis that RAG is an unoptimized software engineering hack compared to holistic neural net training, while Swyx pushes back on explainability, audit trails, and access control as permanent advantages of retrieval over weight-based memorization.34:43–45:44 · The hosts as informed peer 7/10 Building RAG from Scratch and LlamaIndex Architecture Swyx brings up analogies to 'Kubernetes the Hard Way' and suggests universal hyperparameter optimization for chunk sizes, but Jerry clarifies why domain variance makes universal chunking impossible and outlines query-side linear transform techniques.45:45–54:14 · The hosts as informed peer 4/10 Advanced Retrieval, SEC Insights, and Commercial Platform Alessio and Swyx explore practical production challenges like document sunsetting and citation streaming, prompting Jerry to share how building SEC Insights surfaced the engineering hurdles of citation bubbling and full-stack orchestration.54:14–1:04:31 · The hosts as informed peer 6/10 The Data Ecosystem, LLMs, and Vector Databases Swyx brings up the competitive landscape of vector stores and jokingly asks if Jerry's preference is just Postgres. Jerry clarifies the need for hybrid expressivity combining SQL relational filtering with dense semantic retrieval.1:04:32–1:09:22 · The hosts as informed peer 5/10 Evaluation Frameworks and Agent Benchmarks Swyx expresses skepticism over LLM-as-a-judge startups due to output variance, while Jerry pushes back that automated fine-tuned evaluation models are inevitable given that human annotation cannot scale.1:09:23–1:13:06 · The hosts as informed peer 5/10 Lightning Round and Key Takeaways In the lightning round, Jerry predicts that long-term personalization will live in dynamic weight updates rather than external vector stores, leading Swyx to challenge model weights as an unreliable storage medium.5:29–11:42 · Guest teaching 5/10 The Origins and Evolution of GPT Tree Index The hosts inquire about Jerry's origin story at Robust Intelligence and the transition from GPT Tree Index to LlamaIndex. Jerry explains the conceptual limitations of bottom-up tree structures and why early GPT-3 was insufficient for multi-hop reasoning.11:42–20:58 · Guest teaching 4/10 Founding LlamaIndex, Moats, and Company Growth Alessio asks whether expanding context windows eliminate the need for data structuring. Jerry breaks down the economic and latency curves of network transfer costs over petabytes of enterprise data versus targeted retrieval.21:00–34:42 · Guest teaching 7/10 The Core Debate: Why RAG is a Hack vs Fine-Tuning Jerry delivers his core thesis that RAG is an unoptimized software engineering hack compared to holistic neural net training, while Swyx pushes back on explainability, audit trails, and access control as permanent advantages of retrieval over weight-based memorization.34:43–45:44 · Guest teaching 5/10 Building RAG from Scratch and LlamaIndex Architecture Swyx brings up analogies to 'Kubernetes the Hard Way' and suggests universal hyperparameter optimization for chunk sizes, but Jerry clarifies why domain variance makes universal chunking impossible and outlines query-side linear transform techniques.45:45–54:14 · Guest teaching 5/10 Advanced Retrieval, SEC Insights, and Commercial Platform Alessio and Swyx explore practical production challenges like document sunsetting and citation streaming, prompting Jerry to share how building SEC Insights surfaced the engineering hurdles of citation bubbling and full-stack orchestration.54:14–1:04:31 · Guest teaching 5/10 The Data Ecosystem, LLMs, and Vector Databases Swyx brings up the competitive landscape of vector stores and jokingly asks if Jerry's preference is just Postgres. Jerry clarifies the need for hybrid expressivity combining SQL relational filtering with dense semantic retrieval.1:04:32–1:09:22 · Guest teaching 6/10 Evaluation Frameworks and Agent Benchmarks Swyx expresses skepticism over LLM-as-a-judge startups due to output variance, while Jerry pushes back that automated fine-tuned evaluation models are inevitable given that human annotation cannot scale.1:09:23–1:13:06 · Guest teaching 4/10 Lightning Round and Key Takeaways In the lightning round, Jerry predicts that long-term personalization will live in dynamic weight updates rather than external vector stores, leading Swyx to challenge model weights as an unreliable storage medium.5:29–11:42 · Guest disagreement 1/10 The Origins and Evolution of GPT Tree Index The hosts inquire about Jerry's origin story at Robust Intelligence and the transition from GPT Tree Index to LlamaIndex. Jerry explains the conceptual limitations of bottom-up tree structures and why early GPT-3 was insufficient for multi-hop reasoning.11:42–20:58 · Guest disagreement 2/10 Founding LlamaIndex, Moats, and Company Growth Alessio asks whether expanding context windows eliminate the need for data structuring. Jerry breaks down the economic and latency curves of network transfer costs over petabytes of enterprise data versus targeted retrieval.21:00–34:42 · Guest disagreement 4/10 The Core Debate: Why RAG is a Hack vs Fine-Tuning Jerry delivers his core thesis that RAG is an unoptimized software engineering hack compared to holistic neural net training, while Swyx pushes back on explainability, audit trails, and access control as permanent advantages of retrieval over weight-based memorization.34:43–45:44 · Guest disagreement 2/10 Building RAG from Scratch and LlamaIndex Architecture Swyx brings up analogies to 'Kubernetes the Hard Way' and suggests universal hyperparameter optimization for chunk sizes, but Jerry clarifies why domain variance makes universal chunking impossible and outlines query-side linear transform techniques.45:45–54:14 · Guest disagreement 1/10 Advanced Retrieval, SEC Insights, and Commercial Platform Alessio and Swyx explore practical production challenges like document sunsetting and citation streaming, prompting Jerry to share how building SEC Insights surfaced the engineering hurdles of citation bubbling and full-stack orchestration.54:14–1:04:31 · Guest disagreement 3/10 The Data Ecosystem, LLMs, and Vector Databases Swyx brings up the competitive landscape of vector stores and jokingly asks if Jerry's preference is just Postgres. Jerry clarifies the need for hybrid expressivity combining SQL relational filtering with dense semantic retrieval.1:04:32–1:09:22 · Guest disagreement 3/10 Evaluation Frameworks and Agent Benchmarks Swyx expresses skepticism over LLM-as-a-judge startups due to output variance, while Jerry pushes back that automated fine-tuned evaluation models are inevitable given that human annotation cannot scale.1:09:23–1:13:06 · Guest disagreement 2/10 Lightning Round and Key Takeaways In the lightning round, Jerry predicts that long-term personalization will live in dynamic weight updates rather than external vector stores, leading Swyx to challenge model weights as an unreliable storage medium.5:29–11:42 · The hosts pushing back 2/10 The Origins and Evolution of GPT Tree Index The hosts inquire about Jerry's origin story at Robust Intelligence and the transition from GPT Tree Index to LlamaIndex. Jerry explains the conceptual limitations of bottom-up tree structures and why early GPT-3 was insufficient for multi-hop reasoning.11:42–20:58 · The hosts pushing back 2/10 Founding LlamaIndex, Moats, and Company Growth Alessio asks whether expanding context windows eliminate the need for data structuring. Jerry breaks down the economic and latency curves of network transfer costs over petabytes of enterprise data versus targeted retrieval.21:00–34:42 · The hosts pushing back 4/10 The Core Debate: Why RAG is a Hack vs Fine-Tuning Jerry delivers his core thesis that RAG is an unoptimized software engineering hack compared to holistic neural net training, while Swyx pushes back on explainability, audit trails, and access control as permanent advantages of retrieval over weight-based memorization.34:43–45:44 · The hosts pushing back 3/10 Building RAG from Scratch and LlamaIndex Architecture Swyx brings up analogies to 'Kubernetes the Hard Way' and suggests universal hyperparameter optimization for chunk sizes, but Jerry clarifies why domain variance makes universal chunking impossible and outlines query-side linear transform techniques.45:45–54:14 · The hosts pushing back 1/10 Advanced Retrieval, SEC Insights, and Commercial Platform Alessio and Swyx explore practical production challenges like document sunsetting and citation streaming, prompting Jerry to share how building SEC Insights surfaced the engineering hurdles of citation bubbling and full-stack orchestration.54:14–1:04:31 · The hosts pushing back 3/10 The Data Ecosystem, LLMs, and Vector Databases Swyx brings up the competitive landscape of vector stores and jokingly asks if Jerry's preference is just Postgres. Jerry clarifies the need for hybrid expressivity combining SQL relational filtering with dense semantic retrieval.1:04:32–1:09:22 · The hosts pushing back 4/10 Evaluation Frameworks and Agent Benchmarks Swyx expresses skepticism over LLM-as-a-judge startups due to output variance, while Jerry pushes back that automated fine-tuned evaluation models are inevitable given that human annotation cannot scale.1:09:23–1:13:06 · The hosts pushing back 3/10 Lightning Round and Key Takeaways In the lightning round, Jerry predicts that long-term personalization will live in dynamic weight updates rather than external vector stores, leading Swyx to challenge model weights as an unreliable storage medium.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 56.1% · guest 43.9%0:00 · the hosts 56.1% · guest 43.9%3:00 · the hosts 20.9% · guest 79.1%3:00 · the hosts 20.9% · guest 79.1%6:00 · the hosts 28.6% · guest 71.4%6:00 · the hosts 28.6% · guest 71.4%9:00 · the hosts 11.1% · guest 88.9%9:00 · the hosts 11.1% · guest 88.9%12:00 · the hosts 10.8% · guest 89.2%12:00 · the hosts 10.8% · guest 89.2%15:00 · the hosts 14.1% · guest 85.9%15:00 · the hosts 14.1% · guest 85.9%18:00 · the hosts 23.9% · guest 76.1%18:00 · the hosts 23.9% · guest 76.1%21:00 · the hosts 30.4% · guest 69.6%21:00 · the hosts 30.4% · guest 69.6%24:00 · the hosts 9.9% · guest 90.1%24:00 · the hosts 9.9% · guest 90.1%27:00 · the hosts 20.8% · guest 79.2%27:00 · the hosts 20.8% · guest 79.2%30:00 · the hosts 5.5% · guest 94.5%30:00 · the hosts 5.5% · guest 94.5%33:00 · the hosts 27.5% · guest 72.5%33:00 · the hosts 27.5% · guest 72.5%36:00 · the hosts 34.4% · guest 65.6%36:00 · the hosts 34.4% · guest 65.6%39:00 · the hosts 10% · guest 90%39:00 · the hosts 10% · guest 90%42:00 · the hosts 15% · guest 85%42:00 · the hosts 15% · guest 85%45:00 · the hosts 23.2% · guest 76.8%45:00 · the hosts 23.2% · guest 76.8%48:00 · the hosts 16.5% · guest 83.5%48:00 · the hosts 16.5% · guest 83.5%51:00 · the hosts 15.8% · guest 84.2%51:00 · the hosts 15.8% · guest 84.2%54:00 · the hosts 27.9% · guest 72.1%54:00 · the hosts 27.9% · guest 72.1%57:00 · the hosts 37.4% · guest 62.6%57:00 · the hosts 37.4% · guest 62.6%1:00:00 · the hosts 25.7% · guest 74.3%1:00:00 · the hosts 25.7% · guest 74.3%1:03:00 · the hosts 9.7% · guest 90.3%1:03:00 · the hosts 9.7% · guest 90.3%1:06:00 · the hosts 4.3% · guest 95.7%1:06:00 · the hosts 4.3% · guest 95.7%1:09:00 · the hosts 22.2% · guest 77.8%1:09:00 · the hosts 22.2% · guest 77.8%1:12:00 · the hosts 2.8% · guest 97.2%1:12:00 · the hosts 2.8% · guest 97.2%
Sharpest disagreement ▶ 24:38 Jerry calls RAG an unoptimized hack

Jerry contrarianly undermines his own category's purity by framing current RAG architectures as cobbled-together algorithmic software hacks rather than proper end-to-end machine learning optimization.

Hardest push from the hosts ▶ 27:50 Swyx presses on explainability and source attribution

Swyx refuses the idea that end-to-end model training can fully replace RAG, pointing out that enterprise trust strictly requires direct citation links and access control that neural net weights cannot provide.

Biggest teaching moment ▶ 43:25 Query-side linear transforms over frozen document embeddings

Jerry explains how production pipelines avoid massive re-indexing costs during embedding fine-tuning by applying learned linear transforms solely to the query vector.

The host holds their own ▶ 37:40 Swyx draws parallels to Kubernetes and React from scratch

Swyx demonstrates deep developer tooling experience by connecting LlamaIndex's low-level tutorials to Kelsey Hightower's 'Kubernetes the Hard Way' and his own React-from-scratch frameworks.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Origins and Evolution of GPT Tree Index 4512 The hosts inquire about Jerry's origin story at Robust Intelligence and the transition from GPT Tree Index to LlamaIndex. Jerry explains the conceptual limitations of bottom-up tree structures and why early GPT-3 was insufficient for multi-hop reasoning.
Founding LlamaIndex, Moats, and Company Growth 5422 Alessio asks whether expanding context windows eliminate the need for data structuring. Jerry breaks down the economic and latency curves of network transfer costs over petabytes of enterprise data versus targeted retrieval.
The Core Debate: Why RAG is a Hack vs Fine-Tuning 6744 Jerry delivers his core thesis that RAG is an unoptimized software engineering hack compared to holistic neural net training, while Swyx pushes back on explainability, audit trails, and access control as permanent advantages of retrieval over weight-based memorization.
Building RAG from Scratch and LlamaIndex Architecture 7523 Swyx brings up analogies to 'Kubernetes the Hard Way' and suggests universal hyperparameter optimization for chunk sizes, but Jerry clarifies why domain variance makes universal chunking impossible and outlines query-side linear transform techniques.
Advanced Retrieval, SEC Insights, and Commercial Platform 4511 Alessio and Swyx explore practical production challenges like document sunsetting and citation streaming, prompting Jerry to share how building SEC Insights surfaced the engineering hurdles of citation bubbling and full-stack orchestration.
The Data Ecosystem, LLMs, and Vector Databases 6533 Swyx brings up the competitive landscape of vector stores and jokingly asks if Jerry's preference is just Postgres. Jerry clarifies the need for hybrid expressivity combining SQL relational filtering with dense semantic retrieval.
Evaluation Frameworks and Agent Benchmarks 5634 Swyx expresses skepticism over LLM-as-a-judge startups due to output variance, while Jerry pushes back that automated fine-tuned evaluation models are inevitable given that human annotation cannot scale.
Lightning Round and Key Takeaways 5423 In the lightning round, Jerry predicts that long-term personalization will live in dynamic weight updates rather than external vector stores, leading Swyx to challenge model weights as an unreliable storage medium.

Statements from this episode (14)

Insight
Liu: Avoid pre-GPT-4 models for tasks requiring complex reasoning
“Like, I'm one of the first to say, like, you know, you shouldn't use anything pre-GPT-IV for anything that requires, like, complex reasoning because it's just going to gonna be unreliable. Okay, disregarding stuff like fine-tuning.”
Jerry Liu Oct 12, 2023 ▶ 9:28
Insight
Liu: RAG will remain essential for managing LLM cost-performance trade-offs
“There's always going to be some curve regardless of like the performance of the best performing models of like cost versus performance. And so what RAG does is it does provide extra data points along that access because you kind of control the amount of contex…”
Jerry Liu Oct 12, 2023 ▶ 15:01
Disclosure
Liu: LlamaIndex closed its Greylock seed round in under a week
“What we really wanted to do was because for us, like, time was of the essence, like, we wanted to ship very quickly and still kind of build mindshare in this space, we just kept the fundraising process very efficient. I think we basically did it in, like, a we…”
Jerry Liu Oct 12, 2023 ▶ 16:12
Assertion Supported
Wang: LlamaIndex reached 600,000 monthly downloads by September 2023
“And since you're fundraising Post, which was in June. And now it's September, so it's been about three months. You've actually gained 50% in, in terms of stars and followers. You've three X your download count to 600,000 a month and your discord membership has…”
Shawn Wang Oct 12, 2023 ▶ 19:46
Insight
Liu: RAG is fundamentally just an algorithmic prompt-stuffing hack
“RAG is basically just a hack, but it turns out it's a very good hack because what is RAG? RAG is you keep the model fixed, and you just figure out a good way to, like, stuff stuff into the prompt of the language model. Everything that we're doing nowadays in t…”
Jerry Liu Oct 12, 2023 ▶ 24:45
Prediction Open · timeframe Oct 2028
Liu: Developers will eventually fine-tune new factual knowledge into LLMs
“That's one of those things where I think long-term, you definitely can. I think some people say you can't. I disagree. I think you definitely can. Just right now, I haven't gotten into work yet.”
Jerry Liu Oct 12, 2023 ▶ 29:53
Opinion
Liu: Security and access control are not P0 for enterprise apps
“I think users have asked for it, but I don't think that's like a P zero. Like, I think the P zero is more on, like, can we get this thing working before we expand this to, like, more users within the org.”
Jerry Liu Oct 12, 2023 ▶ 34:32
Assertion Not checkable as stated
Liu: 90% of users ask how to improve LLM app performance
“Honestly, 90% of users I talk to have questions about how to improve the performance of their app.”
Jerry Liu Oct 12, 2023 ▶ 37:08
Insight
Liu: Chain-of-thought query decomposition produces superior retrieval results
“Another example here is actually LLM based reasoning, like LLM based chain of thought reasoning. You can take a question, break it down into smaller components and use that to actually send to your retrieval system. And that gives you better results since kind…”
Jerry Liu Oct 12, 2023 ▶ 47:53
Disclosure
Liu: LlamaIndex is de-emphasizing direct revenue generation in 2023
“I think where our revenue focus this year is kind of is less emphasized. Like, it's more just about, like, can we build some managed offering that, like, provides complimentary value to what the open source library provides.”
Jerry Liu Oct 12, 2023 ▶ 53:23
Insight
Liu: Improving vector store lookup algorithms offers low marginal gains
“I don't think the delta on, like, improving the vector store, like, embedding lookup algorithm is that high. I think this stuff has been mostly solved or at least there's just a lot of other stuff you can do to try to improve the overall performance.”
Jerry Liu Oct 12, 2023 ▶ 1:01:34
Prediction Not checkable as stated
Liu: Fine-tuned LLMs will become the primary scalable evaluation solution
“No, but these models will get better and you'll probably fine tune a model to be a better judge. I think that's probably what's going to happen. So I'm like reasonably bullish on this because I don't think there's really a good alternative beyond you just huma…”
Jerry Liu Oct 12, 2023 ▶ 1:06:00
Prediction Not checkable as stated
Liu: The final state for personalized AI memory will not be RAG
“I think a lot of people have thoughts about that, but, like, for what it's worth, I don't think the final state will be ragged. I think it will be some, like, fancy algorithm or architecture where you, like, bake it into, like, the architecture of the model it…”
Jerry Liu Oct 12, 2023 ▶ 1:10:29
Insight
Liu: Developers should build RAG from scratch before using framework abstractions
“Building, like, RAG from scratch. I mean, I think everybody should do it, I think. Like, I would check out the guide if you guys haven't already, and I think it's in our docs, but instead of just using you know, either the kind of, like the retriever query eng…”
Jerry Liu Oct 12, 2023 ▶ 1:12:32
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.