May 4, 2023 · 37m · no-priors

No Priors Ep. 15 | With Kelvin Guu, Staff Research Scientist, Google Brain

Kelvin Guu · 25m spoken Elad Gil · 4m spoken Sarah Guo · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Google Brain Staff Research Scientist Kelvin Guu joins Sarah Guo and Elad Gil on No Priors to discuss retrieval-augmented language modeling (REALM), modular model adaptation, direct knowledge editing, and the architectural requirements for autonomous agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.5% of the talking time here. How this is scored →

The hosts as informed peer 5.4 Guest teaching 5.7 Guest disagreement 0.1 The hosts pushing back 0.6
05100:0010:0020:0030:000:05–5:36 · The hosts as informed peer 4/10 Background in Mathematics, NLP, and Early Work at Google Sarah and Elad prompt Kelvin on his transition from math and NLP to Google Brain and the origins of REALM. Kelvin delivers a clear technical breakdown of dense retrieval, vector spaces, and cross-attention masked language modeling.5:36–8:43 · The hosts as informed peer 6/10 Retrieval vs. Scale and Comparison with Mixture of Experts Elad synthesizes how scaling and repetition build parametric memory versus smaller models needing retrieval. Kelvin expands with a comparison between granular retrieval augmentation and coarse-grained Mixture of Experts (Branch-Train-Merge).8:43–11:53 · The hosts as informed peer 4/10 Modularity, Model Adaptation Trends, and Instruction Following in FLAN Sarah asks about modularity and instruction following. Kelvin details Google's FLAN research on multitask zero-shot generalization, while explaining why simple prompting still cannot resolve issues like hallucinations.11:54–16:19 · The hosts as informed peer 5/10 Continuous Learning, Prompt Tuning, and Attribution via Simfluence The hosts inquire about continuous real-time weight updates and training data explainability. Kelvin unpacks the mechanics of prompt tuning and the counterfactual simulation framework behind his Simfluence paper.16:20–27:20 · The hosts as informed peer 6/10 Model Surgery and Direct Knowledge Editing with ROME Kelvin describes the ROME model surgery paper and the limitations of current prompt-based autonomous agents. Elad and Sarah contribute relevant neurobiology parallels and failure modes in multi-step reasoning.27:20–32:01 · The hosts as informed peer 6/10 Knowledge Representation Trade-offs and Accessible Model Customization Sarah and Kelvin explore the trade-off between canonical knowledge base centralization and dense model coverage. Elad connects Kelvin's framing of personal customization to Constitutional AI.32:02–37:09 · The hosts as informed peer 7/10 Research Advice and the Future of Human Cognitive Skills Kelvin suggests problem formulation and validation will remain durable human skills. Elad pushes back by reviewing the rapid progression of game-playing AI from Go and Poker to Diplomacy, questioning whether human cognitive advantages will persist.0:05–5:36 · Guest teaching 6/10 Background in Mathematics, NLP, and Early Work at Google Sarah and Elad prompt Kelvin on his transition from math and NLP to Google Brain and the origins of REALM. Kelvin delivers a clear technical breakdown of dense retrieval, vector spaces, and cross-attention masked language modeling.5:36–8:43 · Guest teaching 6/10 Retrieval vs. Scale and Comparison with Mixture of Experts Elad synthesizes how scaling and repetition build parametric memory versus smaller models needing retrieval. Kelvin expands with a comparison between granular retrieval augmentation and coarse-grained Mixture of Experts (Branch-Train-Merge).8:43–11:53 · Guest teaching 6/10 Modularity, Model Adaptation Trends, and Instruction Following in FLAN Sarah asks about modularity and instruction following. Kelvin details Google's FLAN research on multitask zero-shot generalization, while explaining why simple prompting still cannot resolve issues like hallucinations.11:54–16:19 · Guest teaching 6/10 Continuous Learning, Prompt Tuning, and Attribution via Simfluence The hosts inquire about continuous real-time weight updates and training data explainability. Kelvin unpacks the mechanics of prompt tuning and the counterfactual simulation framework behind his Simfluence paper.16:20–27:20 · Guest teaching 7/10 Model Surgery and Direct Knowledge Editing with ROME Kelvin describes the ROME model surgery paper and the limitations of current prompt-based autonomous agents. Elad and Sarah contribute relevant neurobiology parallels and failure modes in multi-step reasoning.27:20–32:01 · Guest teaching 5/10 Knowledge Representation Trade-offs and Accessible Model Customization Sarah and Kelvin explore the trade-off between canonical knowledge base centralization and dense model coverage. Elad connects Kelvin's framing of personal customization to Constitutional AI.32:02–37:09 · Guest teaching 4/10 Research Advice and the Future of Human Cognitive Skills Kelvin suggests problem formulation and validation will remain durable human skills. Elad pushes back by reviewing the rapid progression of game-playing AI from Go and Poker to Diplomacy, questioning whether human cognitive advantages will persist.0:05–5:36 · Guest disagreement 0/10 Background in Mathematics, NLP, and Early Work at Google Sarah and Elad prompt Kelvin on his transition from math and NLP to Google Brain and the origins of REALM. Kelvin delivers a clear technical breakdown of dense retrieval, vector spaces, and cross-attention masked language modeling.5:36–8:43 · Guest disagreement 0/10 Retrieval vs. Scale and Comparison with Mixture of Experts Elad synthesizes how scaling and repetition build parametric memory versus smaller models needing retrieval. Kelvin expands with a comparison between granular retrieval augmentation and coarse-grained Mixture of Experts (Branch-Train-Merge).8:43–11:53 · Guest disagreement 0/10 Modularity, Model Adaptation Trends, and Instruction Following in FLAN Sarah asks about modularity and instruction following. Kelvin details Google's FLAN research on multitask zero-shot generalization, while explaining why simple prompting still cannot resolve issues like hallucinations.11:54–16:19 · Guest disagreement 0/10 Continuous Learning, Prompt Tuning, and Attribution via Simfluence The hosts inquire about continuous real-time weight updates and training data explainability. Kelvin unpacks the mechanics of prompt tuning and the counterfactual simulation framework behind his Simfluence paper.16:20–27:20 · Guest disagreement 0/10 Model Surgery and Direct Knowledge Editing with ROME Kelvin describes the ROME model surgery paper and the limitations of current prompt-based autonomous agents. Elad and Sarah contribute relevant neurobiology parallels and failure modes in multi-step reasoning.27:20–32:01 · Guest disagreement 0/10 Knowledge Representation Trade-offs and Accessible Model Customization Sarah and Kelvin explore the trade-off between canonical knowledge base centralization and dense model coverage. Elad connects Kelvin's framing of personal customization to Constitutional AI.32:02–37:09 · Guest disagreement 1/10 Research Advice and the Future of Human Cognitive Skills Kelvin suggests problem formulation and validation will remain durable human skills. Elad pushes back by reviewing the rapid progression of game-playing AI from Go and Poker to Diplomacy, questioning whether human cognitive advantages will persist.0:05–5:36 · The hosts pushing back 0/10 Background in Mathematics, NLP, and Early Work at Google Sarah and Elad prompt Kelvin on his transition from math and NLP to Google Brain and the origins of REALM. Kelvin delivers a clear technical breakdown of dense retrieval, vector spaces, and cross-attention masked language modeling.5:36–8:43 · The hosts pushing back 0/10 Retrieval vs. Scale and Comparison with Mixture of Experts Elad synthesizes how scaling and repetition build parametric memory versus smaller models needing retrieval. Kelvin expands with a comparison between granular retrieval augmentation and coarse-grained Mixture of Experts (Branch-Train-Merge).8:43–11:53 · The hosts pushing back 0/10 Modularity, Model Adaptation Trends, and Instruction Following in FLAN Sarah asks about modularity and instruction following. Kelvin details Google's FLAN research on multitask zero-shot generalization, while explaining why simple prompting still cannot resolve issues like hallucinations.11:54–16:19 · The hosts pushing back 0/10 Continuous Learning, Prompt Tuning, and Attribution via Simfluence The hosts inquire about continuous real-time weight updates and training data explainability. Kelvin unpacks the mechanics of prompt tuning and the counterfactual simulation framework behind his Simfluence paper.16:20–27:20 · The hosts pushing back 1/10 Model Surgery and Direct Knowledge Editing with ROME Kelvin describes the ROME model surgery paper and the limitations of current prompt-based autonomous agents. Elad and Sarah contribute relevant neurobiology parallels and failure modes in multi-step reasoning.27:20–32:01 · The hosts pushing back 0/10 Knowledge Representation Trade-offs and Accessible Model Customization Sarah and Kelvin explore the trade-off between canonical knowledge base centralization and dense model coverage. Elad connects Kelvin's framing of personal customization to Constitutional AI.32:02–37:09 · The hosts pushing back 3/10 Research Advice and the Future of Human Cognitive Skills Kelvin suggests problem formulation and validation will remain durable human skills. Elad pushes back by reviewing the rapid progression of game-playing AI from Go and Poker to Diplomacy, questioning whether human cognitive advantages will persist.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 17.6% · guest 82.4%0:00 · the hosts 17.6% · guest 82.4%3:00 · the hosts 17.3% · guest 82.7%3:00 · the hosts 17.3% · guest 82.7%6:00 · the hosts 23.9% · guest 76.1%6:00 · the hosts 23.9% · guest 76.1%9:00 · the hosts 6.4% · guest 93.6%9:00 · the hosts 6.4% · guest 93.6%12:00 · the hosts 14.5% · guest 85.5%12:00 · the hosts 14.5% · guest 85.5%15:00 · the hosts 10.3% · guest 89.7%15:00 · the hosts 10.3% · guest 89.7%18:00 · the hosts 10.5% · guest 89.5%18:00 · the hosts 10.5% · guest 89.5%21:00 · the hosts 46.2% · guest 53.8%21:00 · the hosts 46.2% · guest 53.8%24:00 · the hosts 32.2% · guest 67.8%24:00 · the hosts 32.2% · guest 67.8%27:00 · the hosts 23.1% · guest 76.9%27:00 · the hosts 23.1% · guest 76.9%30:00 · the hosts 27% · guest 73%30:00 · the hosts 27% · guest 73%33:00 · the hosts 23.5% · guest 76.5%33:00 · the hosts 23.5% · guest 76.5%36:00 · the hosts 100% · guest 0%36:00 · the hosts 100% · guest 0%
Sharpest disagreement ▶ 21:01 Dismissing agent workflow hacking as temporary

Kelvin firmly dismisses the current wave of open-source agent workflow engineering, arguing it will likely be discarded just like manual feature engineering was in earlier ML paradigms.

Hardest push from the hosts ▶ 35:47 Challenging the assumption of enduring human problem formulation

Elad challenges Kelvin's optimism about human creativity remaining distinct from AI by tracing how skepticism in game benchmarks collapsed across Chess, Go, Poker, and Diplomacy.

Biggest teaching moment ▶ 16:36 Explaining ROME parameter editing as lookup table surgery

Kelvin explains how weight matrices operate as key-value lookup tables in transformers and how localized editing propagates factual updates across interconnected knowledge queries.

The host holds their own ▶ 35:47 Citing Noam Brown's Diplomacy work on capability slopes

Elad demonstrates command of AI research history by systematically citing milestones in AI game literature to argue that complex human strategic domains inevitably get surpassed.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Background in Mathematics, NLP, and Early Work at Google 4600 Sarah and Elad prompt Kelvin on his transition from math and NLP to Google Brain and the origins of REALM. Kelvin delivers a clear technical breakdown of dense retrieval, vector spaces, and cross-attention masked language modeling.
Retrieval vs. Scale and Comparison with Mixture of Experts 6600 Elad synthesizes how scaling and repetition build parametric memory versus smaller models needing retrieval. Kelvin expands with a comparison between granular retrieval augmentation and coarse-grained Mixture of Experts (Branch-Train-Merge).
Modularity, Model Adaptation Trends, and Instruction Following in FLAN 4600 Sarah asks about modularity and instruction following. Kelvin details Google's FLAN research on multitask zero-shot generalization, while explaining why simple prompting still cannot resolve issues like hallucinations.
Continuous Learning, Prompt Tuning, and Attribution via Simfluence 5600 The hosts inquire about continuous real-time weight updates and training data explainability. Kelvin unpacks the mechanics of prompt tuning and the counterfactual simulation framework behind his Simfluence paper.
Model Surgery and Direct Knowledge Editing with ROME 6701 Kelvin describes the ROME model surgery paper and the limitations of current prompt-based autonomous agents. Elad and Sarah contribute relevant neurobiology parallels and failure modes in multi-step reasoning.
Knowledge Representation Trade-offs and Accessible Model Customization 6500 Sarah and Kelvin explore the trade-off between canonical knowledge base centralization and dense model coverage. Elad connects Kelvin's framing of personal customization to Constitutional AI.
Research Advice and the Future of Human Cognitive Skills 7413 Kelvin suggests problem formulation and validation will remain durable human skills. Elad pushes back by reviewing the rapid progression of game-playing AI from Go and Poker to Diplomacy, questioning whether human cognitive advantages will persist.

Statements from this episode (16)

Assertion Supported
Guu: BERT Demonstrated Vast Unprogrammed World Knowledge Purely From Pre-Training
“I think one of the things that became very apparent early on when playing with BERT was unlike all the prior generations of models, it had a large amount of world knowledge that we didn't deliberately encode into it. It wasn't in, you know, the fine tuning dat…”
Kelvin Guu May 4, 2023 ▶ 1:23
Opinion
Guu: Tool-Using AI Models Still Struggle to Fulfill Original Retrieval Promises
“And since then we've encountered many interesting challenges on top of that idea that I would say even in, in the sort of systems you see today, these tool using models that issue Google searches or provide citations, they still face some challenges in terms o…”
Kelvin Guu May 4, 2023 ▶ 3:06
Assertion Supported
Guu: REALM Rewarded Document Retrieval That Improved Masked Token Predictions
“And the way Realm was being trained in this method was we would say, okay, you can go and you can retrieve some documents. And if those documents help you fill the blank better, learn that those are useful. And if they don't help you fill the blank better, the…”
Kelvin Guu May 4, 2023 ▶ 5:19
Insight
Guu: Retrieval Models Best Serve Enterprise Needs Through Modularity and Privacy
“So it's actually been a case where I think more and more the benefits of a retrieval augmented model are for modularity. For personal information that you might not want in the main model, for adapting to, say, like an enterprise customer that has special info…”
Kelvin Guu May 4, 2023 ▶ 6:10
Assertion Supported
Guu: LLM Fact Memorization Requires Hitting Minimum Training Data Frequencies
“There are some papers that I can point to later in the show notes that kind of show how memorization scales with the number of times something shows up in the corpus, and you have to hit a certain frequency before these models can kind of accurately remember t…”
Kelvin Guu May 4, 2023 ▶ 6:28
Insight
Guu: Adapting LLMs to Code Requires Specialized Training Over Pure Retrieval
“So if you just want to get very precise factoid information, like what is my Wi-Fi password, retrieval augmentation is going to be very good. But if, for example, you're trying to adapt a language model to a new enterprise, and they have some kind of a special…”
Kelvin Guu May 4, 2023 ▶ 8:14
Assertion Supported
Guu: FLAN Demonstrated Zero-Shot Generalization After Multi-Task Training on 100 Tasks
“What we were able to show was that if you train on a hundred tasks, if we then show the model a new 101 task, it will be able to adapt to that without having seen that particular task before.”
Kelvin Guu May 4, 2023 ▶ 10:39
Insight
Guu: Continuous Real-Time ML Weight Updating Is Rare Due to Validation Costs
“Those I initially thought would have been more popular in kind of production grade settings, but they come with a maintenance cost, which is if you have something updating live, you don't have the opportunity to validate and check that everything is going well…”
Kelvin Guu May 4, 2023 ▶ 12:11
Insight
Guu: Without Attribution, Distinguishing LLM Generalization From Memorization Is Impossible
“Without being able to track that down, we'll never know if these models are generalizing or just kind of cleverly patching together what they know.”
Kelvin Guu May 4, 2023 ▶ 15:53
Opinion
Guu: Model Surgery Research Enables Unique Modularity Over Retrieval and MoE
“I feel that This is a very exciting area for research because it provides a different kind of modularity from retrieval augmented models or mixture of experts, one that actually allows a kind of generalization that's very interesting.”
Kelvin Guu May 4, 2023 ▶ 17:54
Prediction Not checkable as stated
Guu: AI Agents Relying Purely on Explicit Reasoning May Fail to Scale
“That's something that seems to be missing from these autonomous agents right now. Everything is explicit reasoning, and at some point that might not scale or that might become brittle.”
Kelvin Guu May 4, 2023 ▶ 20:39
Prediction Not checkable as stated
Guu: Fundamental AI Advances May Make Agent Workflow Engineering Obsolete
“But it's quite possible that fundamental advances could come along and this kind of workflow hacking or workflow engineering could go the same way as feature engineering or specialized architectures.”
Kelvin Guu May 4, 2023 ▶ 21:28
Insight
Guu: AI Researchers Should Prioritize Making Dense Representations More Controllable
“If I were kind of advising a student or something on looking into knowledge representations, at the moment I would say there's a lot of momentum on, on dense models continuing to capture more and more of the different applications, and so if there is a way tha…”
Kelvin Guu May 4, 2023 ▶ 28:39
Prediction Not checkable as stated
Guu: LLM Providers Will Maximize Prompting Capabilities to Ensure Ease of Use
“So I think there's a strong incentive to make that happen. So the folks who are providing large language models, they want to make their approaches as easy to use as a possible. And so anything that can go into prompting, it seems to me that people will try to…”
Kelvin Guu May 4, 2023 ▶ 31:00
Insight
Guu: LLMs Shift Computing Value From Technical Proficiency to Problem Formulation
“I feel that there's maybe been a shift in what kind of skill is valuable. So at a certain earlier point in time, having technical proficiency was a huge differentiating factor. And if you didn't have that, you just couldn't pursue certain ideas. Whereas now mo…”
Kelvin Guu May 4, 2023 ▶ 34:36
Insight
Guu: Human Technical Skill Remains Necessary to Validate Automated LLM Outputs
“Another thing that has come up in some discussions is that even if the large language model is autopiloting a lot of work for you, someone still needs to validate that. And that may still require a great deal of technical skill, unless you prompt a large langu…”
Kelvin Guu May 4, 2023 ▶ 35:23
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.