Jun 6, 2024 · 36m · no-priors

No Priors Ep. 67 | With Voyage AI Co-Founder and CEO

Tengyu Ma · 25m spoken Sarah Guo · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Sarah Guo interviews Stanford professor and Voyage AI CEO Tengyu Ma about the economics, technical design, and algorithmic optimization of Retrieval-Augmented Generation (RAG) systems. Ma breaks down the RAG versus long-context debate, introduces training efficiency breakthroughs like the Sophia optimizer, and shares insights on bridging academic research with AI entrepreneurship.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23% of the talking time here. How this is scored →

The hosts as informed peer 5.9 Guest teaching 3.8 Guest disagreement 1.2 The hosts pushing back 1.3
05100:0010:0020:0030:000:29–4:27 · The hosts as informed peer 6/10 Tengyu Ma's Research Agenda and Focus Areas Sarah shows strong familiarity with Tengyu's background across theory, RL, and optimizers, interjecting with industry context about Adam's age. Tengyu explains his academic evolution from matrix completion to contrastive learning and the Sophia optimizer.4:28–6:51 · The hosts as informed peer 5/10 Founding Voyage AI and the Evolution of Applied ML Sarah connects Tengyu's transition into entrepreneurship to Conviction's investment thesis about foundation models commoditizing earlier pipeline steps. Tengyu describes the historical shift from 7-step applied ML workflows to prompt and RAG architectures.6:51–9:36 · The hosts as informed peer 5/10 Understanding RAG Systems and Core Retrieval Components Sarah prompts an architectural breakdown of RAG systems and chimes in on vectorizing diverse modalities like code. Tengyu outlines the core mechanics of embedding vectorization, retrieval grounding, and hallucination reduction.9:36–15:40 · The hosts as informed peer 7/10 Real-World Applications and Enterprise Use Cases for RAG Sarah poses the industry counterarguments against RAG, citing agent chaining and infinite-context LLMs. Tengyu addresses the debate systematically, using hardware caching analogies to argue that hierarchical retrieval remains vastly more cost-efficient than long context.15:40–18:02 · The hosts as informed peer 7/10 Token Limits and Scale in Enterprise Contexts Sarah contextualizes Gemini 1.5 Pro's 1M token window in practical terms (code lines, book lengths) and articulates why enterprise scale still demands retrieval. Tengyu reinforces the math with enterprise cost multiples.18:02–20:20 · The hosts as informed peer 6/10 Agent Chaining and Iterative Retrieval Approaches Sarah probes agent chaining as an alternative data management approach. Tengyu reframes agent chaining as orthogonal, explaining that agents still require embedding models and iterative retrieval for efficiency.20:20–23:20 · The hosts as informed peer 5/10 Optimizing Retrieval Performance and Pipeline Simplicity Sarah asks how builders should optimize RAG pipelines beyond the core LLM. Tengyu explains his vision where smarter foundation models eliminate brittle software heuristics like document chunking.23:20–27:08 · The hosts as informed peer 7/10 Domain-Specific Fine-Tuning and Latency Budgets Sarah explains inference-time latency mechanics in search pipelines to clarify why parameter count matters. Tengyu provides benchmarks on domain-specific fine-tuning across code and legal corpora.27:09–30:51 · The hosts as informed peer 6/10 Customization for Enterprises and Practical Builder Advice Sarah prompts practical recommendations for developers and asks for forward-looking predictions on model evolution. Tengyu details profiling strategies and predicts simplified 3-4 component architectures.30:52–35:59 · The hosts as informed peer 5/10 Lessons Learned Transitioning from Academia to Founder Sarah questions the role of academic labs amid massive industrial scaling laws. Tengyu outlines why universities must target 3-to-5-year breakthrough horizons like fundamental optimizers and reasoning conjectures rather than short-term scale.0:29–4:27 · Guest teaching 4/10 Tengyu Ma's Research Agenda and Focus Areas Sarah shows strong familiarity with Tengyu's background across theory, RL, and optimizers, interjecting with industry context about Adam's age. Tengyu explains his academic evolution from matrix completion to contrastive learning and the Sophia optimizer.4:28–6:51 · Guest teaching 3/10 Founding Voyage AI and the Evolution of Applied ML Sarah connects Tengyu's transition into entrepreneurship to Conviction's investment thesis about foundation models commoditizing earlier pipeline steps. Tengyu describes the historical shift from 7-step applied ML workflows to prompt and RAG architectures.6:51–9:36 · Guest teaching 4/10 Understanding RAG Systems and Core Retrieval Components Sarah prompts an architectural breakdown of RAG systems and chimes in on vectorizing diverse modalities like code. Tengyu outlines the core mechanics of embedding vectorization, retrieval grounding, and hallucination reduction.9:36–15:40 · Guest teaching 5/10 Real-World Applications and Enterprise Use Cases for RAG Sarah poses the industry counterarguments against RAG, citing agent chaining and infinite-context LLMs. Tengyu addresses the debate systematically, using hardware caching analogies to argue that hierarchical retrieval remains vastly more cost-efficient than long context.15:40–18:02 · Guest teaching 3/10 Token Limits and Scale in Enterprise Contexts Sarah contextualizes Gemini 1.5 Pro's 1M token window in practical terms (code lines, book lengths) and articulates why enterprise scale still demands retrieval. Tengyu reinforces the math with enterprise cost multiples.18:02–20:20 · Guest teaching 4/10 Agent Chaining and Iterative Retrieval Approaches Sarah probes agent chaining as an alternative data management approach. Tengyu reframes agent chaining as orthogonal, explaining that agents still require embedding models and iterative retrieval for efficiency.20:20–23:20 · Guest teaching 4/10 Optimizing Retrieval Performance and Pipeline Simplicity Sarah asks how builders should optimize RAG pipelines beyond the core LLM. Tengyu explains his vision where smarter foundation models eliminate brittle software heuristics like document chunking.23:20–27:08 · Guest teaching 4/10 Domain-Specific Fine-Tuning and Latency Budgets Sarah explains inference-time latency mechanics in search pipelines to clarify why parameter count matters. Tengyu provides benchmarks on domain-specific fine-tuning across code and legal corpora.27:09–30:51 · Guest teaching 3/10 Customization for Enterprises and Practical Builder Advice Sarah prompts practical recommendations for developers and asks for forward-looking predictions on model evolution. Tengyu details profiling strategies and predicts simplified 3-4 component architectures.30:52–35:59 · Guest teaching 4/10 Lessons Learned Transitioning from Academia to Founder Sarah questions the role of academic labs amid massive industrial scaling laws. Tengyu outlines why universities must target 3-to-5-year breakthrough horizons like fundamental optimizers and reasoning conjectures rather than short-term scale.0:29–4:27 · Guest disagreement 1/10 Tengyu Ma's Research Agenda and Focus Areas Sarah shows strong familiarity with Tengyu's background across theory, RL, and optimizers, interjecting with industry context about Adam's age. Tengyu explains his academic evolution from matrix completion to contrastive learning and the Sophia optimizer.4:28–6:51 · Guest disagreement 1/10 Founding Voyage AI and the Evolution of Applied ML Sarah connects Tengyu's transition into entrepreneurship to Conviction's investment thesis about foundation models commoditizing earlier pipeline steps. Tengyu describes the historical shift from 7-step applied ML workflows to prompt and RAG architectures.6:51–9:36 · Guest disagreement 1/10 Understanding RAG Systems and Core Retrieval Components Sarah prompts an architectural breakdown of RAG systems and chimes in on vectorizing diverse modalities like code. Tengyu outlines the core mechanics of embedding vectorization, retrieval grounding, and hallucination reduction.9:36–15:40 · Guest disagreement 2/10 Real-World Applications and Enterprise Use Cases for RAG Sarah poses the industry counterarguments against RAG, citing agent chaining and infinite-context LLMs. Tengyu addresses the debate systematically, using hardware caching analogies to argue that hierarchical retrieval remains vastly more cost-efficient than long context.15:40–18:02 · Guest disagreement 1/10 Token Limits and Scale in Enterprise Contexts Sarah contextualizes Gemini 1.5 Pro's 1M token window in practical terms (code lines, book lengths) and articulates why enterprise scale still demands retrieval. Tengyu reinforces the math with enterprise cost multiples.18:02–20:20 · Guest disagreement 2/10 Agent Chaining and Iterative Retrieval Approaches Sarah probes agent chaining as an alternative data management approach. Tengyu reframes agent chaining as orthogonal, explaining that agents still require embedding models and iterative retrieval for efficiency.20:20–23:20 · Guest disagreement 1/10 Optimizing Retrieval Performance and Pipeline Simplicity Sarah asks how builders should optimize RAG pipelines beyond the core LLM. Tengyu explains his vision where smarter foundation models eliminate brittle software heuristics like document chunking.23:20–27:08 · Guest disagreement 1/10 Domain-Specific Fine-Tuning and Latency Budgets Sarah explains inference-time latency mechanics in search pipelines to clarify why parameter count matters. Tengyu provides benchmarks on domain-specific fine-tuning across code and legal corpora.27:09–30:51 · Guest disagreement 1/10 Customization for Enterprises and Practical Builder Advice Sarah prompts practical recommendations for developers and asks for forward-looking predictions on model evolution. Tengyu details profiling strategies and predicts simplified 3-4 component architectures.30:52–35:59 · Guest disagreement 1/10 Lessons Learned Transitioning from Academia to Founder Sarah questions the role of academic labs amid massive industrial scaling laws. Tengyu outlines why universities must target 3-to-5-year breakthrough horizons like fundamental optimizers and reasoning conjectures rather than short-term scale.0:29–4:27 · The hosts pushing back 1/10 Tengyu Ma's Research Agenda and Focus Areas Sarah shows strong familiarity with Tengyu's background across theory, RL, and optimizers, interjecting with industry context about Adam's age. Tengyu explains his academic evolution from matrix completion to contrastive learning and the Sophia optimizer.4:28–6:51 · The hosts pushing back 1/10 Founding Voyage AI and the Evolution of Applied ML Sarah connects Tengyu's transition into entrepreneurship to Conviction's investment thesis about foundation models commoditizing earlier pipeline steps. Tengyu describes the historical shift from 7-step applied ML workflows to prompt and RAG architectures.6:51–9:36 · The hosts pushing back 1/10 Understanding RAG Systems and Core Retrieval Components Sarah prompts an architectural breakdown of RAG systems and chimes in on vectorizing diverse modalities like code. Tengyu outlines the core mechanics of embedding vectorization, retrieval grounding, and hallucination reduction.9:36–15:40 · The hosts pushing back 3/10 Real-World Applications and Enterprise Use Cases for RAG Sarah poses the industry counterarguments against RAG, citing agent chaining and infinite-context LLMs. Tengyu addresses the debate systematically, using hardware caching analogies to argue that hierarchical retrieval remains vastly more cost-efficient than long context.15:40–18:02 · The hosts pushing back 1/10 Token Limits and Scale in Enterprise Contexts Sarah contextualizes Gemini 1.5 Pro's 1M token window in practical terms (code lines, book lengths) and articulates why enterprise scale still demands retrieval. Tengyu reinforces the math with enterprise cost multiples.18:02–20:20 · The hosts pushing back 2/10 Agent Chaining and Iterative Retrieval Approaches Sarah probes agent chaining as an alternative data management approach. Tengyu reframes agent chaining as orthogonal, explaining that agents still require embedding models and iterative retrieval for efficiency.20:20–23:20 · The hosts pushing back 1/10 Optimizing Retrieval Performance and Pipeline Simplicity Sarah asks how builders should optimize RAG pipelines beyond the core LLM. Tengyu explains his vision where smarter foundation models eliminate brittle software heuristics like document chunking.23:20–27:08 · The hosts pushing back 1/10 Domain-Specific Fine-Tuning and Latency Budgets Sarah explains inference-time latency mechanics in search pipelines to clarify why parameter count matters. Tengyu provides benchmarks on domain-specific fine-tuning across code and legal corpora.27:09–30:51 · The hosts pushing back 1/10 Customization for Enterprises and Practical Builder Advice Sarah prompts practical recommendations for developers and asks for forward-looking predictions on model evolution. Tengyu details profiling strategies and predicts simplified 3-4 component architectures.30:52–35:59 · The hosts pushing back 1/10 Lessons Learned Transitioning from Academia to Founder Sarah questions the role of academic labs amid massive industrial scaling laws. Tengyu outlines why universities must target 3-to-5-year breakthrough horizons like fundamental optimizers and reasoning conjectures rather than short-term scale.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 26.2% · guest 73.8%0:00 · the hosts 26.2% · guest 73.8%3:00 · the hosts 11.9% · guest 88.1%3:00 · the hosts 11.9% · guest 88.1%6:00 · the hosts 19% · guest 81%6:00 · the hosts 19% · guest 81%9:00 · the hosts 47.4% · guest 52.6%9:00 · the hosts 47.4% · guest 52.6%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 51.8% · guest 48.2%15:00 · the hosts 51.8% · guest 48.2%18:00 · the hosts 23.6% · guest 76.4%18:00 · the hosts 23.6% · guest 76.4%21:00 · the hosts 5% · guest 95%21:00 · the hosts 5% · guest 95%24:00 · the hosts 25.2% · guest 74.8%24:00 · the hosts 25.2% · guest 74.8%27:00 · the hosts 25.5% · guest 74.5%27:00 · the hosts 25.5% · guest 74.5%30:00 · the hosts 23.1% · guest 76.9%30:00 · the hosts 23.1% · guest 76.9%33:00 · the hosts 11.1% · guest 88.9%33:00 · the hosts 11.1% · guest 88.9%36:00 · the hosts 100% · guest 0%36:00 · the hosts 100% · guest 0%
Sharpest disagreement ▶ 18:09 Reframing agent chaining as dependent on retrieval

Tengyu rejects the premise that agent chaining displaces retrieval systems, arguing instead that multi-step agents fundamentally rely on embeddings and small models to remain computationally viable.

Hardest push from the hosts ▶ 10:42 Challenging RAG with infinite context and agent architectures

Sarah directly frames the leading counterarguments from top frontier labs questioning whether RAG architectures will be rendered obsolete by infinite context windows and agentic chaining.

Biggest teaching moment ▶ 13:00 First-principles memory hierarchy comparison

Tengyu educates listeners on the theoretical cost and memory bottlenecks of full-context transformers, using computer architecture caching levels to prove why retrieval remains essential.

The host holds their own ▶ 16:12 Translating token metrics to concrete engineering limits

Sarah demonstrates deep technical and operational domain knowledge by calculating the real-world scale limits of 1M token windows against enterprise codebases and media requirements.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Tengyu Ma's Research Agenda and Focus Areas 6411 Sarah shows strong familiarity with Tengyu's background across theory, RL, and optimizers, interjecting with industry context about Adam's age. Tengyu explains his academic evolution from matrix completion to contrastive learning and the Sophia optimizer.
Founding Voyage AI and the Evolution of Applied ML 5311 Sarah connects Tengyu's transition into entrepreneurship to Conviction's investment thesis about foundation models commoditizing earlier pipeline steps. Tengyu describes the historical shift from 7-step applied ML workflows to prompt and RAG architectures.
Understanding RAG Systems and Core Retrieval Components 5411 Sarah prompts an architectural breakdown of RAG systems and chimes in on vectorizing diverse modalities like code. Tengyu outlines the core mechanics of embedding vectorization, retrieval grounding, and hallucination reduction.
Real-World Applications and Enterprise Use Cases for RAG 7523 Sarah poses the industry counterarguments against RAG, citing agent chaining and infinite-context LLMs. Tengyu addresses the debate systematically, using hardware caching analogies to argue that hierarchical retrieval remains vastly more cost-efficient than long context.
Token Limits and Scale in Enterprise Contexts 7311 Sarah contextualizes Gemini 1.5 Pro's 1M token window in practical terms (code lines, book lengths) and articulates why enterprise scale still demands retrieval. Tengyu reinforces the math with enterprise cost multiples.
Agent Chaining and Iterative Retrieval Approaches 6422 Sarah probes agent chaining as an alternative data management approach. Tengyu reframes agent chaining as orthogonal, explaining that agents still require embedding models and iterative retrieval for efficiency.
Optimizing Retrieval Performance and Pipeline Simplicity 5411 Sarah asks how builders should optimize RAG pipelines beyond the core LLM. Tengyu explains his vision where smarter foundation models eliminate brittle software heuristics like document chunking.
Domain-Specific Fine-Tuning and Latency Budgets 7411 Sarah explains inference-time latency mechanics in search pipelines to clarify why parameter count matters. Tengyu provides benchmarks on domain-specific fine-tuning across code and legal corpora.
Customization for Enterprises and Practical Builder Advice 6311 Sarah prompts practical recommendations for developers and asks for forward-looking predictions on model evolution. Tengyu details profiling strategies and predicts simplified 3-4 component architectures.
Lessons Learned Transitioning from Academia to Founder 5411 Sarah questions the role of academic labs amid massive industrial scaling laws. Tengyu outlines why universities must target 3-to-5-year breakthrough horizons like fundamental optimizers and reasoning conjectures rather than short-term scale.

Statements from this episode (19)

Prediction Not checkable as stated
Ma: AI Is Running Out of Compute and Training Data
“My vision is that in the future the efficiency is very important because we are running out of data and compute. So we have to either use the data much better and use the compute much better.”
Tengyu Ma Jun 6, 2024 ▶ 1:23
Assertion Supported
Ma: Sophia Optimizer Improves LLM Pre-Training Efficiency by 2x
“One of the paper we wrote last year was Sophia which we found, where we found that we have a nutrient optimizer, which can improve the training efficiency by two X for pre-training.”
Tengyu Ma Jun 6, 2024 ▶ 2:59
Assertion Not checkable as stated
Ma: Meta Achieved 1.6x Training Efficiency Gain Using Sophia Optimizer
“And recently I think one of the Facebook friends actually used this in their large scale multimodal training. And they found that in on that scale, I don't know exactly how many parameters there are, but I think, I assume it's kind of more than a hundred billi…”
Tengyu Ma Jun 6, 2024 ▶ 3:54
Insight
Ma: RAG Response Quality Is Bottlenecked by Retrieval Quality
“Right now for implementing Rack the bottleneck seems to be that, you know, it's not very hard to implement it, right? You can just connect the components and have your Rack system ready very quickly. But the bottleneck seems to be the quality. Of the response …”
Tengyu Ma Jun 6, 2024 ▶ 7:03
Insight
Ma: Fine-Tuning Often Fails Due to Data Demands and Hallucinations
“Fine tuning in many cases doesn't work because you need a lot of data to see the results and there are still hallucinations even after fine tuning.”
Tengyu Ma Jun 6, 2024 ▶ 12:06
Prediction Held up
Ma: RAG Will Remain Much Cheaper Than Long-Context Windows
“My prediction is that reg will be much cheaper than long contacts going forward.”
Tengyu Ma Jun 6, 2024 ▶ 14:11
Opinion
Ma: 100M-Token Enterprises Cannot Afford Long-Context Inference Costs
“So right now, if they have a hundred million tokens, I don't think they can use long context transformers at all because it's way too expensive.”
Tengyu Ma Jun 6, 2024 ▶ 17:43
Insight
Ma: Agent Chaining Architectures Still Rely on Embedding Models
“On the first level bit I would say is that I think it's kind of orthogonal to embedding models and re-rankers to some degree, because even when you have agent chaining, right, you still probably use embedding models as part of the chain, right?”
Tengyu Ma Jun 6, 2024 ▶ 18:18
Prediction Not checkable as stated
Ma: Iterative Retrieval Will Diminish as Embedding Models Improve
“However, in the long run, my suspicion is that iterative retrieval will be useful, but it will be a bit less useful as the If the embedding models becomes more and more clever, right? So once the embedding models are more clever, then maybe one run or two runs…”
Tengyu Ma Jun 6, 2024 ▶ 20:01
Disclosure
Ma: Voyage AI Trains Embedding Models on Trillions of Tokens
“We train our network on trillions of tokens at least and we fine tune them for special use cases.”
Tengyu Ma Jun 6, 2024 ▶ 22:05
Prediction Not checkable as stated
Ma: RAG Software Heuristics Will Vanish as Embedding Models Improve
“And my long term vision here is that some of the software engineering layers on top of the networks will be less and less needed when the networks are more and more clever.”
Tengyu Ma Jun 6, 2024 ▶ 22:20
Prediction Not checkable as stated
Ma: Multimodal Embeddings Will Eliminate Image-to-Text Preprocessing
“So maybe in the future, you don't have to turn your images into description of images and then give it to the text embedding model. That's what people are doing right now. Like everything is turned into text and then use a text embedding model. But when the em…”
Tengyu Ma Jun 6, 2024 ▶ 23:02
Disclosure
Ma: Voyage Fine-Tuned Models on 2T Code and 1T Legal Tokens
“We fine tune on two trillions of code snippets, tokens, and then we get a code embedding model and we do the fine tuning on one trillion legal tokens.”
Tengyu Ma Jun 6, 2024 ▶ 23:54
Insight
Ma: Latency Limits Embedding Models to 10 Billion Parameters
“Basically it's impossible to use more than ten billion parameters. For embedding models.”
Tengyu Ma Jun 6, 2024 ▶ 24:36
Assertion Supported
Ma: Domain-Specific Fine-Tuning Yields 15-20% Retrieval Gains in Code
“And we have seen like five to 20% of improvements by this domain specific Fine-tuning depending on the particular domains. For code, we have seen 15 to 20% of improvement, partly because we have a lot of data there.”
Tengyu Ma Jun 6, 2024 ▶ 25:03
Assertion Supported
Ma: Voyage AI Produces Embeddings 3-4x Smaller Than Competitors
“So we produce embeddings that is like a three X, you know, four X smaller dimension than some of the competitors.”
Tengyu Ma Jun 6, 2024 ▶ 26:39
Assertion Not checkable as stated
Ma: Proprietary Data Fine-Tuning Adds 10-20% Retrieval Accuracy
“So we fine tune on the proprietary data of a particular company, and we can see 10 to 20% improvement on top of the domain specific in fine tuning as well.”
Tengyu Ma Jun 6, 2024 ▶ 27:19
Insight
Ma: Academic AI Labs Should Avoid Competing on Model Scaling
“My view is that I think academia probably should work on some different questions from what industry is good at, right? So if we are just only working on how to scale up the system, then obviously the incentive is not right. The, we don't have enough capital t…”
Tengyu Ma Jun 6, 2024 ▶ 33:19
Opinion
Ma: Scaling Web Data Unlikely to Solve Complex Math Conjectures
“It's very unclear whether you can really the scaling law is really enough to get Get you to prove Riemann hypothesis or any of the math conjectures. So you know, and also you have to be superhuman performance in some sense, right? So if you turn on just the co…”
Tengyu Ma Jun 6, 2024 ▶ 35:16
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.