Mar 6, 2025 · 50m · mad

Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI

Douwe Kiela · 35m spoken Matt Turck · 12m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Douwe Kiela, CEO of Contextual AI and lead author of the seminal 2020 RAG paper, to discuss frontier AI models, the evolution of Retrieval-Augmented Generation, and the transition toward enterprise-grade agentic AI systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 25% of the talking time here. How this is scored →

Matt as informed peer 3.7 Guest teaching 4.6 Guest disagreement 0.8 Matt pushing back 1.4
05100:0015:0030:0045:001:45–7:22 · Matt as informed peer 4/10 Assessing Frontier Model Releases, AI Benchmarks, and Test-Time Compute Matt demonstrates solid familiarity with recent frontier model releases, scaling law discussions, and test-time compute. Douwe gently reframes the dichotomy around unsupervised versus reasoning models, explaining that chain-of-thought capabilities mean models like GPT-4o are already reasoning models.7:22–13:52 · Matt as informed peer 4/10 DeepSeek Impact, Training Cost Realities, and Geopolitical Narratives Douwe rejects dominant media narratives around DeepSeek, explaining that the reported six million dollar training cost figure misleadingly omits the 100x prior R&D exploration compute. He also re-educates on model hallucination, noting it is ill-defined and often a desirable feature for creative writing tasks.13:52–20:41 · Matt as informed peer 2/10 Origins and Evolution of Retrieval-Augmented Generation (RAG) Douwe shares the story of inventing RAG at Meta FAIR using early vector search tools like FAISS. Matt acts as an engaging guide, asking about Meta FAIR's open research culture and early reactions to transformer papers.20:41–30:55 · Matt as informed peer 4/10 Naive RAG vs. Advanced RAG and the Long-Context Debate Matt brings up industry debates on fine-tuning and long-context windows replacing RAG. Douwe dismantles these false dichotomies, using a vivid Harry Potter analogy to demonstrate the extreme computational inefficiency of processing massive context windows for basic facts.30:56–33:34 · Matt as informed peer 3/10 Document Extraction and Layout Understanding in RAG Systems Douwe details how layout segmentation models process visual structures in enterprise PDFs like tables and charts. Matt asks whether document extraction issues stem from human decisions or technical bottlenecks.33:34–35:54 · Matt as informed peer 4/10 Enterprise Vector Databases vs. On-Demand API Tool Calling Matt asks whether future architectures will shift from pre-indexed vector databases to live API tool calling. Douwe explains how on-demand tool calling complements vector search by preserving native enterprise access control.35:54–38:04 · Matt as informed peer 3/10 Contextual Language Models and Grounding to Eliminate Hallucinations Douwe outlines Contextual Language Models (CLMs) fine-tuned on top of Llama to strictly ground responses and admit when context is missing. Matt and Douwe highlight Meta's massive contributions to open-source infrastructure like PyTorch and Llama.38:04–40:40 · Matt as informed peer 4/10 Model Alignment and Preference Optimization: From RLHF to KTO and APO Matt references specific preference alignment techniques like KTO and APO. Douwe breaks down why moving away from paired RLHF preferences allows enterprises to optimize models from direct feedback on as few as 100 samples.40:40–44:34 · Matt as informed peer 5/10 Building Specialized RAG Agents for Enterprise Applications Matt presents a thoughtful mental model of retrieval as a tool within broader agentic workflows. Douwe validates this framing, adding that virtually all high-value enterprise agents are fundamentally specialized RAG agents.44:34–49:35 · Matt as informed peer 4/10 The Role of Synthetic Data in End-to-End RAG Optimization Matt asks about practical enterprise adoption barriers and customer expectations regarding AI accuracy. Douwe explains that 100 percent LLM accuracy is impossible, forcing enterprise deployments to manage non-deterministic risks using audit trails and claim verification.1:45–7:22 · Guest teaching 4/10 Assessing Frontier Model Releases, AI Benchmarks, and Test-Time Compute Matt demonstrates solid familiarity with recent frontier model releases, scaling law discussions, and test-time compute. Douwe gently reframes the dichotomy around unsupervised versus reasoning models, explaining that chain-of-thought capabilities mean models like GPT-4o are already reasoning models.7:22–13:52 · Guest teaching 6/10 DeepSeek Impact, Training Cost Realities, and Geopolitical Narratives Douwe rejects dominant media narratives around DeepSeek, explaining that the reported six million dollar training cost figure misleadingly omits the 100x prior R&D exploration compute. He also re-educates on model hallucination, noting it is ill-defined and often a desirable feature for creative writing tasks.13:52–20:41 · Guest teaching 4/10 Origins and Evolution of Retrieval-Augmented Generation (RAG) Douwe shares the story of inventing RAG at Meta FAIR using early vector search tools like FAISS. Matt acts as an engaging guide, asking about Meta FAIR's open research culture and early reactions to transformer papers.20:41–30:55 · Guest teaching 6/10 Naive RAG vs. Advanced RAG and the Long-Context Debate Matt brings up industry debates on fine-tuning and long-context windows replacing RAG. Douwe dismantles these false dichotomies, using a vivid Harry Potter analogy to demonstrate the extreme computational inefficiency of processing massive context windows for basic facts.30:56–33:34 · Guest teaching 4/10 Document Extraction and Layout Understanding in RAG Systems Douwe details how layout segmentation models process visual structures in enterprise PDFs like tables and charts. Matt asks whether document extraction issues stem from human decisions or technical bottlenecks.33:34–35:54 · Guest teaching 4/10 Enterprise Vector Databases vs. On-Demand API Tool Calling Matt asks whether future architectures will shift from pre-indexed vector databases to live API tool calling. Douwe explains how on-demand tool calling complements vector search by preserving native enterprise access control.35:54–38:04 · Guest teaching 4/10 Contextual Language Models and Grounding to Eliminate Hallucinations Douwe outlines Contextual Language Models (CLMs) fine-tuned on top of Llama to strictly ground responses and admit when context is missing. Matt and Douwe highlight Meta's massive contributions to open-source infrastructure like PyTorch and Llama.38:04–40:40 · Guest teaching 5/10 Model Alignment and Preference Optimization: From RLHF to KTO and APO Matt references specific preference alignment techniques like KTO and APO. Douwe breaks down why moving away from paired RLHF preferences allows enterprises to optimize models from direct feedback on as few as 100 samples.40:40–44:34 · Guest teaching 4/10 Building Specialized RAG Agents for Enterprise Applications Matt presents a thoughtful mental model of retrieval as a tool within broader agentic workflows. Douwe validates this framing, adding that virtually all high-value enterprise agents are fundamentally specialized RAG agents.44:34–49:35 · Guest teaching 5/10 The Role of Synthetic Data in End-to-End RAG Optimization Matt asks about practical enterprise adoption barriers and customer expectations regarding AI accuracy. Douwe explains that 100 percent LLM accuracy is impossible, forcing enterprise deployments to manage non-deterministic risks using audit trails and claim verification.1:45–7:22 · Guest disagreement 2/10 Assessing Frontier Model Releases, AI Benchmarks, and Test-Time Compute Matt demonstrates solid familiarity with recent frontier model releases, scaling law discussions, and test-time compute. Douwe gently reframes the dichotomy around unsupervised versus reasoning models, explaining that chain-of-thought capabilities mean models like GPT-4o are already reasoning models.7:22–13:52 · Guest disagreement 4/10 DeepSeek Impact, Training Cost Realities, and Geopolitical Narratives Douwe rejects dominant media narratives around DeepSeek, explaining that the reported six million dollar training cost figure misleadingly omits the 100x prior R&D exploration compute. He also re-educates on model hallucination, noting it is ill-defined and often a desirable feature for creative writing tasks.13:52–20:41 · Guest disagreement 0/10 Origins and Evolution of Retrieval-Augmented Generation (RAG) Douwe shares the story of inventing RAG at Meta FAIR using early vector search tools like FAISS. Matt acts as an engaging guide, asking about Meta FAIR's open research culture and early reactions to transformer papers.20:41–30:55 · Guest disagreement 2/10 Naive RAG vs. Advanced RAG and the Long-Context Debate Matt brings up industry debates on fine-tuning and long-context windows replacing RAG. Douwe dismantles these false dichotomies, using a vivid Harry Potter analogy to demonstrate the extreme computational inefficiency of processing massive context windows for basic facts.30:56–33:34 · Guest disagreement 0/10 Document Extraction and Layout Understanding in RAG Systems Douwe details how layout segmentation models process visual structures in enterprise PDFs like tables and charts. Matt asks whether document extraction issues stem from human decisions or technical bottlenecks.33:34–35:54 · Guest disagreement 0/10 Enterprise Vector Databases vs. On-Demand API Tool Calling Matt asks whether future architectures will shift from pre-indexed vector databases to live API tool calling. Douwe explains how on-demand tool calling complements vector search by preserving native enterprise access control.35:54–38:04 · Guest disagreement 0/10 Contextual Language Models and Grounding to Eliminate Hallucinations Douwe outlines Contextual Language Models (CLMs) fine-tuned on top of Llama to strictly ground responses and admit when context is missing. Matt and Douwe highlight Meta's massive contributions to open-source infrastructure like PyTorch and Llama.38:04–40:40 · Guest disagreement 0/10 Model Alignment and Preference Optimization: From RLHF to KTO and APO Matt references specific preference alignment techniques like KTO and APO. Douwe breaks down why moving away from paired RLHF preferences allows enterprises to optimize models from direct feedback on as few as 100 samples.40:40–44:34 · Guest disagreement 0/10 Building Specialized RAG Agents for Enterprise Applications Matt presents a thoughtful mental model of retrieval as a tool within broader agentic workflows. Douwe validates this framing, adding that virtually all high-value enterprise agents are fundamentally specialized RAG agents.44:34–49:35 · Guest disagreement 0/10 The Role of Synthetic Data in End-to-End RAG Optimization Matt asks about practical enterprise adoption barriers and customer expectations regarding AI accuracy. Douwe explains that 100 percent LLM accuracy is impossible, forcing enterprise deployments to manage non-deterministic risks using audit trails and claim verification.1:45–7:22 · Matt pushing back 3/10 Assessing Frontier Model Releases, AI Benchmarks, and Test-Time Compute Matt demonstrates solid familiarity with recent frontier model releases, scaling law discussions, and test-time compute. Douwe gently reframes the dichotomy around unsupervised versus reasoning models, explaining that chain-of-thought capabilities mean models like GPT-4o are already reasoning models.7:22–13:52 · Matt pushing back 2/10 DeepSeek Impact, Training Cost Realities, and Geopolitical Narratives Douwe rejects dominant media narratives around DeepSeek, explaining that the reported six million dollar training cost figure misleadingly omits the 100x prior R&D exploration compute. He also re-educates on model hallucination, noting it is ill-defined and often a desirable feature for creative writing tasks.13:52–20:41 · Matt pushing back 1/10 Origins and Evolution of Retrieval-Augmented Generation (RAG) Douwe shares the story of inventing RAG at Meta FAIR using early vector search tools like FAISS. Matt acts as an engaging guide, asking about Meta FAIR's open research culture and early reactions to transformer papers.20:41–30:55 · Matt pushing back 3/10 Naive RAG vs. Advanced RAG and the Long-Context Debate Matt brings up industry debates on fine-tuning and long-context windows replacing RAG. Douwe dismantles these false dichotomies, using a vivid Harry Potter analogy to demonstrate the extreme computational inefficiency of processing massive context windows for basic facts.30:56–33:34 · Matt pushing back 1/10 Document Extraction and Layout Understanding in RAG Systems Douwe details how layout segmentation models process visual structures in enterprise PDFs like tables and charts. Matt asks whether document extraction issues stem from human decisions or technical bottlenecks.33:34–35:54 · Matt pushing back 1/10 Enterprise Vector Databases vs. On-Demand API Tool Calling Matt asks whether future architectures will shift from pre-indexed vector databases to live API tool calling. Douwe explains how on-demand tool calling complements vector search by preserving native enterprise access control.35:54–38:04 · Matt pushing back 0/10 Contextual Language Models and Grounding to Eliminate Hallucinations Douwe outlines Contextual Language Models (CLMs) fine-tuned on top of Llama to strictly ground responses and admit when context is missing. Matt and Douwe highlight Meta's massive contributions to open-source infrastructure like PyTorch and Llama.38:04–40:40 · Matt pushing back 0/10 Model Alignment and Preference Optimization: From RLHF to KTO and APO Matt references specific preference alignment techniques like KTO and APO. Douwe breaks down why moving away from paired RLHF preferences allows enterprises to optimize models from direct feedback on as few as 100 samples.40:40–44:34 · Matt pushing back 2/10 Building Specialized RAG Agents for Enterprise Applications Matt presents a thoughtful mental model of retrieval as a tool within broader agentic workflows. Douwe validates this framing, adding that virtually all high-value enterprise agents are fundamentally specialized RAG agents.44:34–49:35 · Matt pushing back 1/10 The Role of Synthetic Data in End-to-End RAG Optimization Matt asks about practical enterprise adoption barriers and customer expectations regarding AI accuracy. Douwe explains that 100 percent LLM accuracy is impossible, forcing enterprise deployments to manage non-deterministic risks using audit trails and claim verification.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 70.5% · guest 29.5%0:00 · Matt 70.5% · guest 29.5%3:00 · Matt 36% · guest 64%3:00 · Matt 36% · guest 64%6:00 · Matt 34.5% · guest 65.5%6:00 · Matt 34.5% · guest 65.5%9:00 · Matt 10.1% · guest 89.9%9:00 · Matt 10.1% · guest 89.9%12:00 · Matt 33.6% · guest 66.4%12:00 · Matt 33.6% · guest 66.4%15:00 · Matt 6.9% · guest 93.1%15:00 · Matt 6.9% · guest 93.1%18:00 · Matt 26.3% · guest 73.7%18:00 · Matt 26.3% · guest 73.7%21:00 · Matt 4% · guest 96%21:00 · Matt 4% · guest 96%24:00 · Matt 41.1% · guest 58.9%24:00 · Matt 41.1% · guest 58.9%27:00 · Matt 5.5% · guest 94.5%27:00 · Matt 5.5% · guest 94.5%30:00 · Matt 21.6% · guest 78.4%30:00 · Matt 21.6% · guest 78.4%33:00 · Matt 18.2% · guest 81.8%33:00 · Matt 18.2% · guest 81.8%36:00 · Matt 22.8% · guest 77.2%36:00 · Matt 22.8% · guest 77.2%39:00 · Matt 24.2% · guest 75.8%39:00 · Matt 24.2% · guest 75.8%42:00 · Matt 27.7% · guest 72.3%42:00 · Matt 27.7% · guest 72.3%45:00 · Matt 18.6% · guest 81.4%45:00 · Matt 18.6% · guest 81.4%48:00 · Matt 24.9% · guest 75.1%48:00 · Matt 24.9% · guest 75.1%
Sharpest disagreement ▶ 9:40 Debunking the DeepSeek $6M Training Cost Narrative

Douwe forcefully rejects the dominant media framing around DeepSeek's efficiency, arguing that citing a six million dollar single training run ignores the 100x prior R&D computation costs required to discover that run.

Hardest push from Matt ▶ 26:50 Host Defends Content Creators and VCs

When Douwe criticizes journalists and venture capitalists for creating false binary tech narratives, Matt jokingly pushes back by noting that VCs and content creators increasingly overlap to deliver valuable industry analysis.

Biggest teaching moment ▶ 27:50 The Harry Potter Analogy for Long-Context Inefficiency

Douwe educates the host on why long-context model windows cannot replace RAG systems, illustrating how computationally wasteful it is to feed all seven Harry Potter books into a context window just to identify the headmaster.

Matt holds his own ▶ 40:40 Framing Retrieval as an Agentic Tool

Matt demonstrates sharp architectural understanding by framing retrieval not merely as a fixed pipeline lookup step, but as a specialized tool dynamically called within multi-step agentic workflows.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Assessing Frontier Model Releases, AI Benchmarks, and Test-Time Compute 4423 Matt demonstrates solid familiarity with recent frontier model releases, scaling law discussions, and test-time compute. Douwe gently reframes the dichotomy around unsupervised versus reasoning models, explaining that chain-of-thought capabilities mean models like GPT-4o are already reasoning models.
DeepSeek Impact, Training Cost Realities, and Geopolitical Narratives 4642 Douwe rejects dominant media narratives around DeepSeek, explaining that the reported six million dollar training cost figure misleadingly omits the 100x prior R&D exploration compute. He also re-educates on model hallucination, noting it is ill-defined and often a desirable feature for creative writing tasks.
Origins and Evolution of Retrieval-Augmented Generation (RAG) 2401 Douwe shares the story of inventing RAG at Meta FAIR using early vector search tools like FAISS. Matt acts as an engaging guide, asking about Meta FAIR's open research culture and early reactions to transformer papers.
Naive RAG vs. Advanced RAG and the Long-Context Debate 4623 Matt brings up industry debates on fine-tuning and long-context windows replacing RAG. Douwe dismantles these false dichotomies, using a vivid Harry Potter analogy to demonstrate the extreme computational inefficiency of processing massive context windows for basic facts.
Document Extraction and Layout Understanding in RAG Systems 3401 Douwe details how layout segmentation models process visual structures in enterprise PDFs like tables and charts. Matt asks whether document extraction issues stem from human decisions or technical bottlenecks.
Enterprise Vector Databases vs. On-Demand API Tool Calling 4401 Matt asks whether future architectures will shift from pre-indexed vector databases to live API tool calling. Douwe explains how on-demand tool calling complements vector search by preserving native enterprise access control.
Contextual Language Models and Grounding to Eliminate Hallucinations 3400 Douwe outlines Contextual Language Models (CLMs) fine-tuned on top of Llama to strictly ground responses and admit when context is missing. Matt and Douwe highlight Meta's massive contributions to open-source infrastructure like PyTorch and Llama.
Model Alignment and Preference Optimization: From RLHF to KTO and APO 4500 Matt references specific preference alignment techniques like KTO and APO. Douwe breaks down why moving away from paired RLHF preferences allows enterprises to optimize models from direct feedback on as few as 100 samples.
Building Specialized RAG Agents for Enterprise Applications 5402 Matt presents a thoughtful mental model of retrieval as a tool within broader agentic workflows. Douwe validates this framing, adding that virtually all high-value enterprise agents are fundamentally specialized RAG agents.
The Role of Synthetic Data in End-to-End RAG Optimization 4501 Matt asks about practical enterprise adoption barriers and customer expectations regarding AI accuracy. Douwe explains that 100 percent LLM accuracy is impossible, forcing enterprise deployments to manage non-deterministic risks using audit trails and claim verification.

Statements from this episode (15)

Insight
Kiela: DeepSeek proved frontier AI models can rely on synthetic data
“We have kind of an existence proof now that it's actually not that hard to do this and so you don't need to invest all that much in, in data, and you can use synthetic data and get a pretty good model out of that”
Douwe Kiela Mar 6, 2025 ▶ 3:12
Opinion
Kiela: GPT-4o is already effectively a reasoning model via chain of thought
“I mean, you could argue that GPT-IV-O is also already a reasoning model. It just hasn't been trained on reasoning specifically, but, ah, it can do chain of thought, right? So if it can do chain of thought, it's basically already a reasoning model. It just hasn…”
Douwe Kiela Mar 6, 2025 ▶ 7:00
Assertion Not checkable as stated
Kiela: DeepSeek's total development cost was at least 100x its $6M training
“So I would guess that they spent at least a hundred X The amount of that, that single training run, right?”
Douwe Kiela Mar 6, 2025 ▶ 9:58
Opinion
Kiela: Core language model development is almost solved and plateauing
“It's not even really about language models anymore. That has almost been solved, right? That's kind of why you see things plateauing off a little bit as well.”
Douwe Kiela Mar 6, 2025 ▶ 11:20
Prediction Not checkable as stated
Kiela: AI is heading toward specialized language models over generalists
“Where we're headed is that we will have more specialized language models.”
Douwe Kiela Mar 6, 2025 ▶ 13:31
Assertion Contradicted
Kiela: FAIR was the first team to build a generative RAG model
“Why RAG became the way you name these things is because it's generative, right? So we were the first ones to have a generative model there.”
Douwe Kiela Mar 6, 2025 ▶ 18:38
Opinion
Kiela: Attention mechanism, not Transformers, was the real AI breakthrough
“So I would say, and maybe I'm biased because one of my best friends is, is on the original attention paper, but that was the real breakthrough. It's just like figuring out that you have this attention mechanism that actually allows you to yeah, to do a much be…”
Douwe Kiela Mar 6, 2025 ▶ 20:19
Assertion Contradicted
Kiela: FAISS was the first vector database
“In the initial paper, we used a vector database or a face. So the words vector database didn't exist at the time. But so face was the first vector database.”
Douwe Kiela Mar 6, 2025 ▶ 21:28
Insight
Kiela: Fine-tuning cannot inject new knowledge into AI models
“One common misconception about fine tuning is a lot of people think that you can inject new knowledge into a model using fine tuning. And that is not true.”
Douwe Kiela Mar 6, 2025 ▶ 28:19
Opinion
Kiela: Long-context LLMs are inherently incredibly wasteful
“Long context models are inherently incredibly wasteful. You're paying for all this compute, and that's maybe why some of the companies that are trying to really sell long context model, long context window models, they will make more money from that, right?”
Douwe Kiela Mar 6, 2025 ▶ 29:23
Insight
Kiela: Enterprise RAG fails if complex document data is not properly extracted
“If you want to have a enterprise grade rag system, you are only as good as the data that goes into that rag system. So if you can't extract the data in the right way, so if you have like a sort of table structure and it has like nested information, you can't g…”
Douwe Kiela Mar 6, 2025 ▶ 31:43
Insight
Kiela: Advanced RAG systems break down when scaling to a million PDFs
“You can build a very awesome demo on a couple of PDFs and things will probably work. But then you have to scale it up to a million PDFs, and then everything breaks down. And the reason for that is that a lot of these kind of advanced RAG systems still actually…”
Douwe Kiela Mar 6, 2025 ▶ 35:14
Assertion Not checkable as stated
Kiela: Aligning language models needs only 100 examples via Anchored Preference Optimization
“So you can train on this when you only have like a hundred examples, you can really make a meaningful, meaningful difference.”
Douwe Kiela Mar 6, 2025 ▶ 40:20
Prediction Not checkable as stated
Kiela: AI systems will probably never reach 100 percent accuracy
“When are we getting to a hundred percent accuracy? And I had to give them the bad news that probably never.”
Douwe Kiela Mar 6, 2025 ▶ 48:21
Insight
Kiela: Retrieval is the only way AI agents can handle proprietary data
“Really focused on retrieval because that's really the only way you get these agents to work on your data and your problems.”
Douwe Kiela Mar 6, 2025 ▶ 50:02
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.