Apr 11, 2024 · 1h 5m · latent-space

Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit

Jungwon Byun · 27m spoken Andreas Stuhlmüller · 21m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Elicit co-founders Andreas Stuhlmüller and Jungwon Byun join the Latent Space Podcast to discuss building an AI research assistant for scientific literature, detailing their journey from alignment research to developing reproducible notebook workflows, faithful RAG architectures, and scalable reasoning systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.8 Guest teaching 5.0 Guest disagreement 1.3 The hosts pushing back 2.2
05100:0015:0030:0045:001:00:001:20–4:10 · The hosts as informed peer 2/10 Andreas Stuhlmüller's Journey to AI and Founding Ought Swix prompts Andreas on his background, prompting an autobiographical overview of his path from early AI exploration to founding Ought. The dynamic is exploratory and narrative with minimal friction.4:10–9:03 · The hosts as informed peer 1/10 Jungwon Byun's Background and Ought's Early Research Agenda Jungwon delivers a detailed monologue explaining how Ought conducted human-in-the-loop experiments to simulate superintelligent systems before foundation models took off. The hosts listen as she explains process supervision.9:03–13:46 · The hosts as informed peer 4/10 Co-Founder Matching and Values Alignment Swix and Alessio probe the co-founder matching process and the counterintuitive transition from research alignment to product execution. The guests explain their 50-page evaluation document and mutual alignment.13:46–18:17 · The hosts as informed peer 4/10 Defining Elicit: The AI Research Assistant for Systematic Reviews Alessio asks for the precise definition of Elicit and questions its scope beyond literature reviews. Jungwon explains how automating systematic reviews and meta-analyses provides the foundation for general scientific reasoning.18:17–24:04 · The hosts as informed peer 5/10 Generalist AI Research Platforms vs. Domain-Specific Tools Alessio brings up BrightWave to challenge whether domain-specific assistants or a single generalist platform will win. Andreas counters that high-level scientific reasoning is inherently cross-domain and likely a winner-take-all problem.24:04–27:13 · The hosts as informed peer 5/10 Market Opportunity, VC Skepticism, and GPT Version Shifts Swix pushes on market sizing and investor skepticism that researchers lack money. Andreas and Jungwon defend their thesis by pointing to the massive total economic value of global R&D and foundational truth-seeking.27:13–29:16 · The hosts as informed peer 5/10 Tabular Workflows, Concept Grouping, and Building Defensible Moats Swix explores Elicit's tabular pivot features and defensibility against new foundational models. Jungwon dismisses generic moat worries by emphasizing that user workflow problem spaces remain vastly larger than current model coverage.29:16–32:11 · The hosts as informed peer 5/10 Implementing Constitutional AI for Faithful Abstract Summaries Swix inquires about cost management and early Constitutional AI work done with an intern. Andreas explains how they created constitutions and fine-tuned open-source models using RLHF for faithful abstract summarization.32:11–36:36 · The hosts as informed peer 6/10 Model Evaluation, Monitoring, and Selecting Model Architectures Swix presses on monitoring production hallucination rates and evaluating vision models like Claude Haiku and GPT-4V. Andreas differentiates basic latency DevOps from offline training-time evaluation suites.36:36–41:09 · The hosts as informed peer 6/10 Developing Computational Notebooks for Iterative Research Workflows Alessio and Andreas discuss transitioning from conversational chatbots to computational notebooks. Andreas explains how notebooks allow defining, debugging, and executing reproducible multi-step data pipelines across thousands of papers.41:09–45:02 · The hosts as informed peer 6/10 Human-in-the-Loop Evaluation and Uncertainty Calibration Swix expresses skepticism regarding model self-reported uncertainty, stating models often hallucinate calibration. Andreas explains using separate evaluation models (e.g., Llama generating and GPT-4 judging) to achieve reliable uncertainty estimates.45:02–48:05 · The hosts as informed peer 5/10 Credit-Based Pricing and Converting Compute to Answer Quality The discussion covers credit-based pricing models and variable compute scaling. Andreas explains the philosophy of allowing users to invest arbitrary compute for deep multi-step verification and rigorous answers.48:05–52:32 · The hosts as informed peer 7/10 RAG Pipelines vs. Massive Context Windows Alessio and Swix challenge the necessity of RAG versus ultra-long context windows. Andreas argues that structured RAG provides essential debuggability and causal attribution that monolithic context windows obscure.52:32–54:34 · The hosts as informed peer 7/10 Hard Grounding and Balancing Context with Model Knowledge Alessio shares a specific example with hardware specs to illustrate how hard grounding can suppress parametric model knowledge. Andreas distinguishes between strict background-synthesis tasks and truth-seeking agent loops.54:34–58:23 · The hosts as informed peer 4/10 Custom Columns and Unexpected Diagnostic Medical Use Cases Alessio invites discussion on underrated features, leading Jungwon to showcase customizable extraction prompts and describe an unexpected diagnostic test interpretation workflow in rural clinical medicine.58:23–1:01:32 · The hosts as informed peer 5/10 Systematizing Scientific Discovery and Developing World Models Alessio brings up scientific resource allocation and automated discovery. Andreas explains why current models lack the structured world models needed to make surprising, cross-domain scientific discoveries.1:20–4:10 · Guest teaching 4/10 Andreas Stuhlmüller's Journey to AI and Founding Ought Swix prompts Andreas on his background, prompting an autobiographical overview of his path from early AI exploration to founding Ought. The dynamic is exploratory and narrative with minimal friction.4:10–9:03 · Guest teaching 6/10 Jungwon Byun's Background and Ought's Early Research Agenda Jungwon delivers a detailed monologue explaining how Ought conducted human-in-the-loop experiments to simulate superintelligent systems before foundation models took off. The hosts listen as she explains process supervision.9:03–13:46 · Guest teaching 4/10 Co-Founder Matching and Values Alignment Swix and Alessio probe the co-founder matching process and the counterintuitive transition from research alignment to product execution. The guests explain their 50-page evaluation document and mutual alignment.13:46–18:17 · Guest teaching 5/10 Defining Elicit: The AI Research Assistant for Systematic Reviews Alessio asks for the precise definition of Elicit and questions its scope beyond literature reviews. Jungwon explains how automating systematic reviews and meta-analyses provides the foundation for general scientific reasoning.18:17–24:04 · Guest teaching 5/10 Generalist AI Research Platforms vs. Domain-Specific Tools Alessio brings up BrightWave to challenge whether domain-specific assistants or a single generalist platform will win. Andreas counters that high-level scientific reasoning is inherently cross-domain and likely a winner-take-all problem.24:04–27:13 · Guest teaching 6/10 Market Opportunity, VC Skepticism, and GPT Version Shifts Swix pushes on market sizing and investor skepticism that researchers lack money. Andreas and Jungwon defend their thesis by pointing to the massive total economic value of global R&D and foundational truth-seeking.27:13–29:16 · Guest teaching 4/10 Tabular Workflows, Concept Grouping, and Building Defensible Moats Swix explores Elicit's tabular pivot features and defensibility against new foundational models. Jungwon dismisses generic moat worries by emphasizing that user workflow problem spaces remain vastly larger than current model coverage.29:16–32:11 · Guest teaching 5/10 Implementing Constitutional AI for Faithful Abstract Summaries Swix inquires about cost management and early Constitutional AI work done with an intern. Andreas explains how they created constitutions and fine-tuned open-source models using RLHF for faithful abstract summarization.32:11–36:36 · Guest teaching 5/10 Model Evaluation, Monitoring, and Selecting Model Architectures Swix presses on monitoring production hallucination rates and evaluating vision models like Claude Haiku and GPT-4V. Andreas differentiates basic latency DevOps from offline training-time evaluation suites.36:36–41:09 · Guest teaching 5/10 Developing Computational Notebooks for Iterative Research Workflows Alessio and Andreas discuss transitioning from conversational chatbots to computational notebooks. Andreas explains how notebooks allow defining, debugging, and executing reproducible multi-step data pipelines across thousands of papers.41:09–45:02 · Guest teaching 5/10 Human-in-the-Loop Evaluation and Uncertainty Calibration Swix expresses skepticism regarding model self-reported uncertainty, stating models often hallucinate calibration. Andreas explains using separate evaluation models (e.g., Llama generating and GPT-4 judging) to achieve reliable uncertainty estimates.45:02–48:05 · Guest teaching 4/10 Credit-Based Pricing and Converting Compute to Answer Quality The discussion covers credit-based pricing models and variable compute scaling. Andreas explains the philosophy of allowing users to invest arbitrary compute for deep multi-step verification and rigorous answers.48:05–52:32 · Guest teaching 5/10 RAG Pipelines vs. Massive Context Windows Alessio and Swix challenge the necessity of RAG versus ultra-long context windows. Andreas argues that structured RAG provides essential debuggability and causal attribution that monolithic context windows obscure.52:32–54:34 · Guest teaching 5/10 Hard Grounding and Balancing Context with Model Knowledge Alessio shares a specific example with hardware specs to illustrate how hard grounding can suppress parametric model knowledge. Andreas distinguishes between strict background-synthesis tasks and truth-seeking agent loops.54:34–58:23 · Guest teaching 6/10 Custom Columns and Unexpected Diagnostic Medical Use Cases Alessio invites discussion on underrated features, leading Jungwon to showcase customizable extraction prompts and describe an unexpected diagnostic test interpretation workflow in rural clinical medicine.58:23–1:01:32 · Guest teaching 6/10 Systematizing Scientific Discovery and Developing World Models Alessio brings up scientific resource allocation and automated discovery. Andreas explains why current models lack the structured world models needed to make surprising, cross-domain scientific discoveries.1:20–4:10 · Guest disagreement 1/10 Andreas Stuhlmüller's Journey to AI and Founding Ought Swix prompts Andreas on his background, prompting an autobiographical overview of his path from early AI exploration to founding Ought. The dynamic is exploratory and narrative with minimal friction.4:10–9:03 · Guest disagreement 1/10 Jungwon Byun's Background and Ought's Early Research Agenda Jungwon delivers a detailed monologue explaining how Ought conducted human-in-the-loop experiments to simulate superintelligent systems before foundation models took off. The hosts listen as she explains process supervision.9:03–13:46 · Guest disagreement 1/10 Co-Founder Matching and Values Alignment Swix and Alessio probe the co-founder matching process and the counterintuitive transition from research alignment to product execution. The guests explain their 50-page evaluation document and mutual alignment.13:46–18:17 · Guest disagreement 1/10 Defining Elicit: The AI Research Assistant for Systematic Reviews Alessio asks for the precise definition of Elicit and questions its scope beyond literature reviews. Jungwon explains how automating systematic reviews and meta-analyses provides the foundation for general scientific reasoning.18:17–24:04 · Guest disagreement 2/10 Generalist AI Research Platforms vs. Domain-Specific Tools Alessio brings up BrightWave to challenge whether domain-specific assistants or a single generalist platform will win. Andreas counters that high-level scientific reasoning is inherently cross-domain and likely a winner-take-all problem.24:04–27:13 · Guest disagreement 2/10 Market Opportunity, VC Skepticism, and GPT Version Shifts Swix pushes on market sizing and investor skepticism that researchers lack money. Andreas and Jungwon defend their thesis by pointing to the massive total economic value of global R&D and foundational truth-seeking.27:13–29:16 · Guest disagreement 1/10 Tabular Workflows, Concept Grouping, and Building Defensible Moats Swix explores Elicit's tabular pivot features and defensibility against new foundational models. Jungwon dismisses generic moat worries by emphasizing that user workflow problem spaces remain vastly larger than current model coverage.29:16–32:11 · Guest disagreement 1/10 Implementing Constitutional AI for Faithful Abstract Summaries Swix inquires about cost management and early Constitutional AI work done with an intern. Andreas explains how they created constitutions and fine-tuned open-source models using RLHF for faithful abstract summarization.32:11–36:36 · Guest disagreement 2/10 Model Evaluation, Monitoring, and Selecting Model Architectures Swix presses on monitoring production hallucination rates and evaluating vision models like Claude Haiku and GPT-4V. Andreas differentiates basic latency DevOps from offline training-time evaluation suites.36:36–41:09 · Guest disagreement 1/10 Developing Computational Notebooks for Iterative Research Workflows Alessio and Andreas discuss transitioning from conversational chatbots to computational notebooks. Andreas explains how notebooks allow defining, debugging, and executing reproducible multi-step data pipelines across thousands of papers.41:09–45:02 · Guest disagreement 2/10 Human-in-the-Loop Evaluation and Uncertainty Calibration Swix expresses skepticism regarding model self-reported uncertainty, stating models often hallucinate calibration. Andreas explains using separate evaluation models (e.g., Llama generating and GPT-4 judging) to achieve reliable uncertainty estimates.45:02–48:05 · Guest disagreement 1/10 Credit-Based Pricing and Converting Compute to Answer Quality The discussion covers credit-based pricing models and variable compute scaling. Andreas explains the philosophy of allowing users to invest arbitrary compute for deep multi-step verification and rigorous answers.48:05–52:32 · Guest disagreement 2/10 RAG Pipelines vs. Massive Context Windows Alessio and Swix challenge the necessity of RAG versus ultra-long context windows. Andreas argues that structured RAG provides essential debuggability and causal attribution that monolithic context windows obscure.52:32–54:34 · Guest disagreement 2/10 Hard Grounding and Balancing Context with Model Knowledge Alessio shares a specific example with hardware specs to illustrate how hard grounding can suppress parametric model knowledge. Andreas distinguishes between strict background-synthesis tasks and truth-seeking agent loops.54:34–58:23 · Guest disagreement 0/10 Custom Columns and Unexpected Diagnostic Medical Use Cases Alessio invites discussion on underrated features, leading Jungwon to showcase customizable extraction prompts and describe an unexpected diagnostic test interpretation workflow in rural clinical medicine.58:23–1:01:32 · Guest disagreement 1/10 Systematizing Scientific Discovery and Developing World Models Alessio brings up scientific resource allocation and automated discovery. Andreas explains why current models lack the structured world models needed to make surprising, cross-domain scientific discoveries.1:20–4:10 · The hosts pushing back 1/10 Andreas Stuhlmüller's Journey to AI and Founding Ought Swix prompts Andreas on his background, prompting an autobiographical overview of his path from early AI exploration to founding Ought. The dynamic is exploratory and narrative with minimal friction.4:10–9:03 · The hosts pushing back 0/10 Jungwon Byun's Background and Ought's Early Research Agenda Jungwon delivers a detailed monologue explaining how Ought conducted human-in-the-loop experiments to simulate superintelligent systems before foundation models took off. The hosts listen as she explains process supervision.9:03–13:46 · The hosts pushing back 2/10 Co-Founder Matching and Values Alignment Swix and Alessio probe the co-founder matching process and the counterintuitive transition from research alignment to product execution. The guests explain their 50-page evaluation document and mutual alignment.13:46–18:17 · The hosts pushing back 2/10 Defining Elicit: The AI Research Assistant for Systematic Reviews Alessio asks for the precise definition of Elicit and questions its scope beyond literature reviews. Jungwon explains how automating systematic reviews and meta-analyses provides the foundation for general scientific reasoning.18:17–24:04 · The hosts pushing back 3/10 Generalist AI Research Platforms vs. Domain-Specific Tools Alessio brings up BrightWave to challenge whether domain-specific assistants or a single generalist platform will win. Andreas counters that high-level scientific reasoning is inherently cross-domain and likely a winner-take-all problem.24:04–27:13 · The hosts pushing back 4/10 Market Opportunity, VC Skepticism, and GPT Version Shifts Swix pushes on market sizing and investor skepticism that researchers lack money. Andreas and Jungwon defend their thesis by pointing to the massive total economic value of global R&D and foundational truth-seeking.27:13–29:16 · The hosts pushing back 2/10 Tabular Workflows, Concept Grouping, and Building Defensible Moats Swix explores Elicit's tabular pivot features and defensibility against new foundational models. Jungwon dismisses generic moat worries by emphasizing that user workflow problem spaces remain vastly larger than current model coverage.29:16–32:11 · The hosts pushing back 1/10 Implementing Constitutional AI for Faithful Abstract Summaries Swix inquires about cost management and early Constitutional AI work done with an intern. Andreas explains how they created constitutions and fine-tuned open-source models using RLHF for faithful abstract summarization.32:11–36:36 · The hosts pushing back 3/10 Model Evaluation, Monitoring, and Selecting Model Architectures Swix presses on monitoring production hallucination rates and evaluating vision models like Claude Haiku and GPT-4V. Andreas differentiates basic latency DevOps from offline training-time evaluation suites.36:36–41:09 · The hosts pushing back 2/10 Developing Computational Notebooks for Iterative Research Workflows Alessio and Andreas discuss transitioning from conversational chatbots to computational notebooks. Andreas explains how notebooks allow defining, debugging, and executing reproducible multi-step data pipelines across thousands of papers.41:09–45:02 · The hosts pushing back 4/10 Human-in-the-Loop Evaluation and Uncertainty Calibration Swix expresses skepticism regarding model self-reported uncertainty, stating models often hallucinate calibration. Andreas explains using separate evaluation models (e.g., Llama generating and GPT-4 judging) to achieve reliable uncertainty estimates.45:02–48:05 · The hosts pushing back 1/10 Credit-Based Pricing and Converting Compute to Answer Quality The discussion covers credit-based pricing models and variable compute scaling. Andreas explains the philosophy of allowing users to invest arbitrary compute for deep multi-step verification and rigorous answers.48:05–52:32 · The hosts pushing back 4/10 RAG Pipelines vs. Massive Context Windows Alessio and Swix challenge the necessity of RAG versus ultra-long context windows. Andreas argues that structured RAG provides essential debuggability and causal attribution that monolithic context windows obscure.52:32–54:34 · The hosts pushing back 3/10 Hard Grounding and Balancing Context with Model Knowledge Alessio shares a specific example with hardware specs to illustrate how hard grounding can suppress parametric model knowledge. Andreas distinguishes between strict background-synthesis tasks and truth-seeking agent loops.54:34–58:23 · The hosts pushing back 1/10 Custom Columns and Unexpected Diagnostic Medical Use Cases Alessio invites discussion on underrated features, leading Jungwon to showcase customizable extraction prompts and describe an unexpected diagnostic test interpretation workflow in rural clinical medicine.58:23–1:01:32 · The hosts pushing back 2/10 Systematizing Scientific Discovery and Developing World Models Alessio brings up scientific resource allocation and automated discovery. Andreas explains why current models lack the structured world models needed to make surprising, cross-domain scientific discoveries.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 25:15 Dismissing short-sighted VC skepticism on researcher market size

Andreas directly challenges the narrow VC view that selling to researchers is unviable, arguing that improving broad R&D efficiency unlocks massive economic value.

Hardest push from the hosts ▶ 44:08 Swix challenges model-generated uncertainty calibration

Swix refuses to accept that LLMs can self-report confidence accurately, citing his own empirical experience where models simply hallucinate their calibration.

Biggest teaching moment ▶ 7:25 Jungwon breaks down process supervision for evaluating superhuman systems

Jungwon provides a deep explanation of decomposing cognitive tasks into verifiable sub-steps to supervise processes rather than blindly trusting model outputs.

The host holds their own ▶ 52:45 Alessio tests context grounding boundaries with hardware benchmark examples

Alessio demonstrates technical depth by citing an empirical AMD MI300 versus NVIDIA NVLink prompt experiment where document grounding suppressed accurate model knowledge.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Andreas Stuhlmüller's Journey to AI and Founding Ought 2411 Swix prompts Andreas on his background, prompting an autobiographical overview of his path from early AI exploration to founding Ought. The dynamic is exploratory and narrative with minimal friction.
Jungwon Byun's Background and Ought's Early Research Agenda 1610 Jungwon delivers a detailed monologue explaining how Ought conducted human-in-the-loop experiments to simulate superintelligent systems before foundation models took off. The hosts listen as she explains process supervision.
Co-Founder Matching and Values Alignment 4412 Swix and Alessio probe the co-founder matching process and the counterintuitive transition from research alignment to product execution. The guests explain their 50-page evaluation document and mutual alignment.
Defining Elicit: The AI Research Assistant for Systematic Reviews 4512 Alessio asks for the precise definition of Elicit and questions its scope beyond literature reviews. Jungwon explains how automating systematic reviews and meta-analyses provides the foundation for general scientific reasoning.
Generalist AI Research Platforms vs. Domain-Specific Tools 5523 Alessio brings up BrightWave to challenge whether domain-specific assistants or a single generalist platform will win. Andreas counters that high-level scientific reasoning is inherently cross-domain and likely a winner-take-all problem.
Market Opportunity, VC Skepticism, and GPT Version Shifts 5624 Swix pushes on market sizing and investor skepticism that researchers lack money. Andreas and Jungwon defend their thesis by pointing to the massive total economic value of global R&D and foundational truth-seeking.
Tabular Workflows, Concept Grouping, and Building Defensible Moats 5412 Swix explores Elicit's tabular pivot features and defensibility against new foundational models. Jungwon dismisses generic moat worries by emphasizing that user workflow problem spaces remain vastly larger than current model coverage.
Implementing Constitutional AI for Faithful Abstract Summaries 5511 Swix inquires about cost management and early Constitutional AI work done with an intern. Andreas explains how they created constitutions and fine-tuned open-source models using RLHF for faithful abstract summarization.
Model Evaluation, Monitoring, and Selecting Model Architectures 6523 Swix presses on monitoring production hallucination rates and evaluating vision models like Claude Haiku and GPT-4V. Andreas differentiates basic latency DevOps from offline training-time evaluation suites.
Developing Computational Notebooks for Iterative Research Workflows 6512 Alessio and Andreas discuss transitioning from conversational chatbots to computational notebooks. Andreas explains how notebooks allow defining, debugging, and executing reproducible multi-step data pipelines across thousands of papers.
Human-in-the-Loop Evaluation and Uncertainty Calibration 6524 Swix expresses skepticism regarding model self-reported uncertainty, stating models often hallucinate calibration. Andreas explains using separate evaluation models (e.g., Llama generating and GPT-4 judging) to achieve reliable uncertainty estimates.
Credit-Based Pricing and Converting Compute to Answer Quality 5411 The discussion covers credit-based pricing models and variable compute scaling. Andreas explains the philosophy of allowing users to invest arbitrary compute for deep multi-step verification and rigorous answers.
RAG Pipelines vs. Massive Context Windows 7524 Alessio and Swix challenge the necessity of RAG versus ultra-long context windows. Andreas argues that structured RAG provides essential debuggability and causal attribution that monolithic context windows obscure.
Hard Grounding and Balancing Context with Model Knowledge 7523 Alessio shares a specific example with hardware specs to illustrate how hard grounding can suppress parametric model knowledge. Andreas distinguishes between strict background-synthesis tasks and truth-seeking agent loops.
Custom Columns and Unexpected Diagnostic Medical Use Cases 4601 Alessio invites discussion on underrated features, leading Jungwon to showcase customizable extraction prompts and describe an unexpected diagnostic test interpretation workflow in rural clinical medicine.
Systematizing Scientific Discovery and Developing World Models 5612 Alessio brings up scientific resource allocation and automated discovery. Andreas explains why current models lack the structured world models needed to make surprising, cross-domain scientific discoveries.

Statements from this episode (27)

Insight
Stuhlmüller: Academia cannot build great software tools due to paper timelines
“It's really hard to actually build interesting tools as an academic. You can't really hire great engineers. Everything is kind of on a paper to paper timeline.”
Andreas Stuhlmüller Apr 11, 2024 ▶ 3:12
Insight
Byun: Supervising step-by-step AI reasoning makes models far easier to evaluate
“The importance of supervising the process of AI systems, not just the outcomes. And so a big part of how, then, like, how Elicit is built is, We're very intentional about not just throwing a ton of data into a model and training it and then saying, cool, here'…”
Jungwon Byun Apr 11, 2024 ▶ 8:12
Assertion Not checkable as stated
Stuhlmüller: Elicit co-founders wrote a 50-page mutual evaluation document before starting
“We also did a pretty lengthy mutual evaluation process where we had a Google Doc where we had all kinds of questions for each other, and I think it ended up being around 50 pages or so of, like, various, like, questions and back and forth.”
Andreas Stuhlmüller Apr 11, 2024 ▶ 10:04
Disclosure
Stuhlmüller: Elicit builds scaffolding rather than training foundation models
“The way we are building Elicit is not let's train a foundation model to do more stuff. It's like let's build a scaffolding such that we can deploy powerful models to good ends.”
Andreas Stuhlmüller Apr 11, 2024 ▶ 11:20
Assertion Not checkable as stated
Stuhlmüller: Elicit continues to use T5-based models
“We do also use, like, T-Five-based models, even, even now but started, yeah, started with GPT-II.”
Andreas Stuhlmüller Apr 11, 2024 ▶ 13:18
Assertion Supported
Byun: Scientific meta-analysis typically takes five people over a year
“Lisset was very much inspired by this workflow in literature called systematic reviews or meta-analysis, which is basically the human state of the art for summarizing scientific literature. It typically involves like five people working together for over a yea…”
Jungwon Byun Apr 11, 2024 ▶ 17:18
Insight
Byun: Highly structured, reproducible research workflows are uniquely amenable to automation
“Because it's so structured and designed to be reproducible, it's really amenable to automation. So that's kind of the one, the workflow that we want to automate first.”
Jungwon Byun Apr 11, 2024 ▶ 18:00
Prediction Not checkable as stated
Stuhlmüller: Generalist AI research platforms will be a winner-take-all market
“So I think there will be, at least within research, I think there will be, like, one best platform, more or less for this type of generalist research. I think there may still be, like, some particular tools, like, for genomics, like, particular types of module…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 19:06
Insight
Byun: Research, not probability modeling, is the bottleneck in forecasting
“The thing that's blocking people from making interesting predictions about important events in the world is less kind of on the probabilistic side and much more on the research side.”
Jungwon Byun Apr 11, 2024 ▶ 20:46
Disclosure
Stuhlmüller: Seed VCs urged Elicit to build legal AI over research
“We did encounter, I guess talking to VCs for our seed round. A lot of VCs were like, you know, researchers, they don't have any money. Why don't you build a legal assistant?”
Andreas Stuhlmüller Apr 11, 2024 ▶ 25:17
Opinion
Byun: GPT-3 was a qualitative shift, while GPT-4 was an extension
“I think GPT-III was a big change because it kind of said, oh, now is the time to build to you that we can use AI to build these tools. And then GPT-IV was maybe a little bit more of an extension of GPT-III. It felt less like a level, GPT-III over GPT-II was li…”
Jungwon Byun Apr 11, 2024 ▶ 26:47
Insight
Byun: Foundational models will not commoditize Elicit due to deep workflow specialization
“I think about this a lot in the context of moats. People are like, oh, what's your moat? What happens if GPT-V comes out? It's like, if GPT-V comes out, there's still like all of this other space that we can go into. And so I think being really obsessed with t…”
Jungwon Byun Apr 11, 2024 ▶ 28:57
Assertion Not checkable as stated
Byun: Anthropic's Constitutional AI slashed Elicit's query costs tenfold in days
“At the start of twenty-twenty-three, Anthropik kind of launched their constitutional AI paper and within a few days, I think four days, he had basically implemented that in production, and then we had it in-app, like, a week or so after that, and he has since …”
Jungwon Byun Apr 11, 2024 ▶ 30:10
Insight
Byun: Early LLMs prioritized answering questions over faithfulness to source text
“At the time, the models hadn't been trained at all to be faithful to a text. So they were just generating. So then when you ask them a question, they tried too hard to ask, answer the question, and didn't try hard enough to answer the question given the text o…”
Jungwon Byun Apr 11, 2024 ▶ 31:52
Opinion
Stuhlmüller: Claude Haiku offers an optimal balance of cost and accuracy
“Specifically, I think Cloud Haiku is like a good point on the kind of Pareto frontier, so I think it's like, it's not the, it's neither the cheapest model nor is it the most accurate, most high quality model, but it's just like a really good trade-off between …”
Andreas Stuhlmüller Apr 11, 2024 ▶ 34:16
Disclosure
Stuhlmüller: Closed-source models consume most of Elicit's compute budget
“I'd say, like, in terms of number of careers, it's maybe similar. In terms of, like, cost and compute, I think the closed models make, make up more of the budget, since the main cases where you want to use closed models are cases where they're just smarter, wh…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 36:17
Insight
Stuhlmüller: Notebooks enable debugging and scaling workflows far better than chat
“But the important difference in our minds is with notebooks you can define a process. So in, in data science you can go like, here's like my data analysis process that takes in a CSV and then does some Extraction, and then generates a figure at the end, and yo…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 39:20
Insight
Stuhlmüller: AI primitives should be semantic tasks, not granular chain-of-thought
“I think chain of thought is maybe still like kind of one level lower on the abstraction hierarchy than we would think of notebooks. I think we'll probably want to think about more semantic pieces, like a building block is more like A paper search, or an extrac…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 40:41
Assertion Not checkable as stated
Byun: LLM self-reported uncertainty is reasonably well-calibrated in production
“We found it to be pretty calibrated. There varies on the model.”
Jungwon Byun Apr 11, 2024 ▶ 44:29
Insight
Stuhlmüller: Separate evaluator models yield better uncertainty estimates than self-evaluation
“I think in some cases we also use the different models for the uncertainty estimates. Yes, then, for the question answering. So, one model would say, here's my chain of thought, here's my answer, and then a different type of model. Let's say the first model is…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 44:33
Opinion
Stuhlmüller: Trading inference compute for answer accuracy is undervalued in AI
“Being able to invest more or less compute into getting more or less accurate answers is, I think, one of the core things we care about, and that I think is currently undervalued in the AI space.”
Andreas Stuhlmüller Apr 11, 2024 ▶ 46:03
Insight
Stuhlmüller: List-wise re-ranking outperforms per-item scoring in search
“In the past, I think a lot of ranking was kind of per item ranking where you would score each individual item, maybe using increasingly expensive scoring methods, and then rank based on the scores, but I think list-wise re-ranking where you have a model that c…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 48:49
Insight
Stuhlmüller: Pure long-context LLMs are significantly harder to debug than RAG
“In one sense, I think you're right that the throw everything into the context window thing is easier to maintain because you just can swap out a model. In another sense, it's, if things go wrong, it's harder to debug, where, like, if you know, here's the proce…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 51:48
Insight
Stuhlmüller: Agentic search ideally balances parametric memory with retrieved context documents
“I think probably the ideal thing looks a bit more like agent control where the model can issue a query that then is intended to surface documents that substantiate its hunch. So I would, that's maybe a reasonable middle ground between model just telling you an…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 53:50
Prediction Not checkable as stated
Stuhlmüller: In 10-20 years, today's scientific methods will look incredibly unsystematic
“Probably, yeah, I, I'd guess in like, 1020 years, we'll look back and it will be incredible how unsystematic science was back in the day.”
Andreas Stuhlmüller Apr 11, 2024 ▶ 59:39
Insight
Stuhlmüller: AI requires deeper world models to make novel scientific discoveries
“Having deeper models of how, let's see, what are the underlying structures of different domains, how they're related or not related, I think will be an important ingredient for models actually being able to make novel contributions.”
Andreas Stuhlmüller Apr 11, 2024 ▶ 1:01:19
Insight
Stuhlmüller: AI orchestration resembles software engineering far more than ML research
“I think a lot of this looks more like traditional software engineering than it does look like machine learning research, and I think the people who are, like, really good at building good abstractions building applications that can kind of survive even if some…”
Andreas Stuhlmüller Apr 11, 2024 ▶ 1:03:27
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.