Apr 11, 2024 · 1h 5m · latent-space
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Elicit co-founders Andreas Stuhlmüller and Jungwon Byun join the Latent Space Podcast to discuss building an AI research assistant for scientific literature, detailing their journey from alignment research to developing reproducible notebook workflows, faithful RAG architectures, and scalable reasoning systems.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Andreas directly challenges the narrow VC view that selling to researchers is unviable, arguing that improving broad R&D efficiency unlocks massive economic value.
Hardest push from the hosts ▶ 44:08 Swix challenges model-generated uncertainty calibrationSwix refuses to accept that LLMs can self-report confidence accurately, citing his own empirical experience where models simply hallucinate their calibration.
Biggest teaching moment ▶ 7:25 Jungwon breaks down process supervision for evaluating superhuman systemsJungwon provides a deep explanation of decomposing cognitive tasks into verifiable sub-steps to supervise processes rather than blindly trusting model outputs.
The host holds their own ▶ 52:45 Alessio tests context grounding boundaries with hardware benchmark examplesAlessio demonstrates technical depth by citing an empirical AMD MI300 versus NVIDIA NVLink prompt experiment where document grounding suppressed accurate model knowledge.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Andreas Stuhlmüller's Journey to AI and Founding Ought | 2 | 4 | 1 | 1 | Swix prompts Andreas on his background, prompting an autobiographical overview of his path from early AI exploration to founding Ought. The dynamic is exploratory and narrative with minimal friction. | |
| Jungwon Byun's Background and Ought's Early Research Agenda | 1 | 6 | 1 | 0 | Jungwon delivers a detailed monologue explaining how Ought conducted human-in-the-loop experiments to simulate superintelligent systems before foundation models took off. The hosts listen as she explains process supervision. | |
| Co-Founder Matching and Values Alignment | 4 | 4 | 1 | 2 | Swix and Alessio probe the co-founder matching process and the counterintuitive transition from research alignment to product execution. The guests explain their 50-page evaluation document and mutual alignment. | |
| Defining Elicit: The AI Research Assistant for Systematic Reviews | 4 | 5 | 1 | 2 | Alessio asks for the precise definition of Elicit and questions its scope beyond literature reviews. Jungwon explains how automating systematic reviews and meta-analyses provides the foundation for general scientific reasoning. | |
| Generalist AI Research Platforms vs. Domain-Specific Tools | 5 | 5 | 2 | 3 | Alessio brings up BrightWave to challenge whether domain-specific assistants or a single generalist platform will win. Andreas counters that high-level scientific reasoning is inherently cross-domain and likely a winner-take-all problem. | |
| Market Opportunity, VC Skepticism, and GPT Version Shifts | 5 | 6 | 2 | 4 | Swix pushes on market sizing and investor skepticism that researchers lack money. Andreas and Jungwon defend their thesis by pointing to the massive total economic value of global R&D and foundational truth-seeking. | |
| Tabular Workflows, Concept Grouping, and Building Defensible Moats | 5 | 4 | 1 | 2 | Swix explores Elicit's tabular pivot features and defensibility against new foundational models. Jungwon dismisses generic moat worries by emphasizing that user workflow problem spaces remain vastly larger than current model coverage. | |
| Implementing Constitutional AI for Faithful Abstract Summaries | 5 | 5 | 1 | 1 | Swix inquires about cost management and early Constitutional AI work done with an intern. Andreas explains how they created constitutions and fine-tuned open-source models using RLHF for faithful abstract summarization. | |
| Model Evaluation, Monitoring, and Selecting Model Architectures | 6 | 5 | 2 | 3 | Swix presses on monitoring production hallucination rates and evaluating vision models like Claude Haiku and GPT-4V. Andreas differentiates basic latency DevOps from offline training-time evaluation suites. | |
| Developing Computational Notebooks for Iterative Research Workflows | 6 | 5 | 1 | 2 | Alessio and Andreas discuss transitioning from conversational chatbots to computational notebooks. Andreas explains how notebooks allow defining, debugging, and executing reproducible multi-step data pipelines across thousands of papers. | |
| Human-in-the-Loop Evaluation and Uncertainty Calibration | 6 | 5 | 2 | 4 | Swix expresses skepticism regarding model self-reported uncertainty, stating models often hallucinate calibration. Andreas explains using separate evaluation models (e.g., Llama generating and GPT-4 judging) to achieve reliable uncertainty estimates. | |
| Credit-Based Pricing and Converting Compute to Answer Quality | 5 | 4 | 1 | 1 | The discussion covers credit-based pricing models and variable compute scaling. Andreas explains the philosophy of allowing users to invest arbitrary compute for deep multi-step verification and rigorous answers. | |
| RAG Pipelines vs. Massive Context Windows | 7 | 5 | 2 | 4 | Alessio and Swix challenge the necessity of RAG versus ultra-long context windows. Andreas argues that structured RAG provides essential debuggability and causal attribution that monolithic context windows obscure. | |
| Hard Grounding and Balancing Context with Model Knowledge | 7 | 5 | 2 | 3 | Alessio shares a specific example with hardware specs to illustrate how hard grounding can suppress parametric model knowledge. Andreas distinguishes between strict background-synthesis tasks and truth-seeking agent loops. | |
| Custom Columns and Unexpected Diagnostic Medical Use Cases | 4 | 6 | 0 | 1 | Alessio invites discussion on underrated features, leading Jungwon to showcase customizable extraction prompts and describe an unexpected diagnostic test interpretation workflow in rural clinical medicine. | |
| Systematizing Scientific Discovery and Developing World Models | 5 | 6 | 1 | 2 | Alessio brings up scientific resource allocation and automated discovery. Andreas explains why current models lack the structured world models needed to make surprising, cross-domain scientific discoveries. |