Jan 11, 2024 · 1h 35m · latent-space

The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert

Nathan Lambert · 1h 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Podcast, AI researcher Dr. Nathan Lambert delivers a comprehensive masterclass on Reinforcement Learning from Human Feedback (RLHF), exploring its mathematical mechanics, historical lineage, data economics, and emerging alternatives like Direct Preference Optimization (DPO). The discussion provides critical technical insights into how alignment steers frontier language models, the transition toward synthetic supervision, and the maturation of open-source alignment ecosystems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.6 Guest teaching 5.3 Guest disagreement 1.4 The hosts pushing back 0.8
05100:0020:0040:001:00:001:20:000:45–3:43 · The hosts as informed peer 3/10 Ultra-Endurance Sports and Staying Grounded Casual introductory rapport where Swix and Nathan discuss ultra-endurance running, trail access in Berkeley, and Nathan's Substack essays.3:43–8:34 · The hosts as informed peer 4/10 Motivation for the RLHF Presentation Alessio and Nathan discuss how RL transitioned from traditional robotics to LLM alignment, citing the Llama 2 paper's surprise at RL's efficacy.8:34–13:35 · The hosts as informed peer 3/10 The Historical and Philosophical Roots of RLHF Nathan details the philosophical and decision-theoretic assumptions underpinning RLHF, citing Aristotle and the Von Neumann-Morgenstern utility theorem.13:36–17:18 · The hosts as informed peer 4/10 Bridging Social Sciences and AI Engineering Alessio inquires about the intersection of sociology and AI systems engineering, while Nathan explains why RLHF formulation resembles a bandit problem more than traditional RL.17:18–21:04 · The hosts as informed peer 4/10 Chain-of-Thought Reasoning and Process Reward Models Nathan breaks down process reward models and historical milestones from TAMER (2008) to the Christiano et al. (2017) human preference paper.21:05–23:09 · The hosts as informed peer 6/10 Signal Richness and Negative Feedback in Alignment Swix highlights the distinction between next-token prediction and negative feedback in RLHF; Nathan agrees it is a richer, more compute-efficient supervisory signal.23:09–27:32 · The hosts as informed peer 4/10 The Core Assumption: Evaluating vs. Generating Text Nathan reviews the core assumption from OpenAI's 2018–2020 papers that evaluating text is significantly easier than generating high-quality text.27:33–30:05 · The hosts as informed peer 5/10 Instruction Tuning vs. RLHF in Practical Applications Nathan advises startups against building full RLHF pipelines unless critical to their moat, advocating instruction tuning and DPO instead.30:05–34:06 · The hosts as informed peer 5/10 The Power and Limits of Instruction Tuning at Scale The conversation touches on early open-source instruction tuning milestones like Vicuna, contrasting small model limits against 70B parameter scaling.34:08–38:26 · The hosts as informed peer 4/10 Mathematical Formulation of the RLHF Objective Nathan explains the mathematical RLHF objective, detailing the role of the KL-divergence constraint to prevent model mode collapse and reward hacking.38:26–42:30 · The hosts as informed peer 7/10 Challenges in Preference Aggregation and Annotation Guidelines Swix introduces Arrow's Impossibility Theorem regarding preference aggregation; Nathan acknowledges social choice limitations and explains how annotator guidelines try to bypass it.42:31–46:06 · The hosts as informed peer 6/10 Preference Data Categories and Granular Metadata Swix asks whether safety guardrails should be handled in RLHF or via downstream classifier filters, prompting Nathan to outline industrial deployment trade-offs.46:06–51:32 · The hosts as informed peer 6/10 The High Economic Cost of Human Preference Data Swix and Nathan break down the massive economic costs of human preference data versus synthetic data generation via GPT-4.51:33–55:57 · The hosts as informed peer 4/10 Open Preference Datasets and LMSYS Chatbot Arena Nathan describes the Bradley-Terry preference loss function math and the 65-75% inter-annotator agreement ceiling observed across industry datasets.55:58–59:11 · The hosts as informed peer 3/10 Implementing Proximal Policy Optimization (PPO) for LLMs Nathan deconstructs the multi-model architecture needed for PPO (actor, critic, reference, reward models), emphasizing systems engineering overhead.59:14–1:01:38 · The hosts as informed peer 5/10 Impact of RLHF on Benchmark Capabilities vs. Style Nathan points out that RLHF mostly alters model tone and style rather than factual academic capabilities, citing GPT-4 technical report exam charts.1:01:40–1:05:12 · The hosts as informed peer 4/10 Non-PPO Alternatives: Best-of-N and Rejection Sampling Nathan reviews lightweight non-PPO optimization strategies including Best-of-N inference sampling and iterative rejection sampling.1:05:12–1:09:48 · The hosts as informed peer 4/10 Unpacking Anthropic's Constitutional AI Nathan clarifies Anthropic's Constitutional AI, correcting common misconceptions by explaining its dual-phase nature of instruction revision and principle sampling.1:09:49–1:12:31 · The hosts as informed peer 7/10 Superalignment and Weak-to-Strong Generalization When Nathan notes skepticism around OpenAI's weak-to-strong generalization paper, Swix pushes back with a comprehensive explanation connecting it to superalignment.1:12:32–1:17:40 · The hosts as informed peer 5/10 Direct Preference Optimization (DPO) Mechanics & Debate Nathan explains DPO mechanics and implicit reward formulations, playfully critiquing academic researchers who avoid explaining the intuitive mechanics.1:17:40–1:21:19 · The hosts as informed peer 4/10 AI2's Mission and Scaling TÜLU 2 via DPO Nathan details the Allen Institute's mission and the training run behind Tulu 2 70B using JAX TPUs and DPO recipes.1:21:21–1:26:46 · The hosts as informed peer 5/10 Evaluating LLMs: Leaderboards and Real-World Usage Nathan critiques automated leaderboard gaming and explains the significance of LMSYS Chatbot Arena human Elo rankings.1:26:47–1:29:43 · The hosts as informed peer 4/10 Comparing Chat Evaluation Benchmarks: MT-Bench vs. AlpacaEval Nathan contrasts the evaluation methodologies of MT-Bench (multi-turn GPT-4 scoring) and AlpacaEval (pairwise win rate versus DaVinci-003).1:29:46–1:33:20 · The hosts as informed peer 5/10 Open Frontiers and Outer-Loop Optimization in RLHF Nathan outlines outer-loop RLHF optimization (continual retraining on user prompts and fixes) as described by John Schulman, wrapping up with startup landscape analysis.0:45–3:43 · Guest teaching 1/10 Ultra-Endurance Sports and Staying Grounded Casual introductory rapport where Swix and Nathan discuss ultra-endurance running, trail access in Berkeley, and Nathan's Substack essays.3:43–8:34 · Guest teaching 4/10 Motivation for the RLHF Presentation Alessio and Nathan discuss how RL transitioned from traditional robotics to LLM alignment, citing the Llama 2 paper's surprise at RL's efficacy.8:34–13:35 · Guest teaching 6/10 The Historical and Philosophical Roots of RLHF Nathan details the philosophical and decision-theoretic assumptions underpinning RLHF, citing Aristotle and the Von Neumann-Morgenstern utility theorem.13:36–17:18 · Guest teaching 5/10 Bridging Social Sciences and AI Engineering Alessio inquires about the intersection of sociology and AI systems engineering, while Nathan explains why RLHF formulation resembles a bandit problem more than traditional RL.17:18–21:04 · Guest teaching 6/10 Chain-of-Thought Reasoning and Process Reward Models Nathan breaks down process reward models and historical milestones from TAMER (2008) to the Christiano et al. (2017) human preference paper.21:05–23:09 · Guest teaching 4/10 Signal Richness and Negative Feedback in Alignment Swix highlights the distinction between next-token prediction and negative feedback in RLHF; Nathan agrees it is a richer, more compute-efficient supervisory signal.23:09–27:32 · Guest teaching 5/10 The Core Assumption: Evaluating vs. Generating Text Nathan reviews the core assumption from OpenAI's 2018–2020 papers that evaluating text is significantly easier than generating high-quality text.27:33–30:05 · Guest teaching 6/10 Instruction Tuning vs. RLHF in Practical Applications Nathan advises startups against building full RLHF pipelines unless critical to their moat, advocating instruction tuning and DPO instead.30:05–34:06 · Guest teaching 5/10 The Power and Limits of Instruction Tuning at Scale The conversation touches on early open-source instruction tuning milestones like Vicuna, contrasting small model limits against 70B parameter scaling.34:08–38:26 · Guest teaching 7/10 Mathematical Formulation of the RLHF Objective Nathan explains the mathematical RLHF objective, detailing the role of the KL-divergence constraint to prevent model mode collapse and reward hacking.38:26–42:30 · Guest teaching 4/10 Challenges in Preference Aggregation and Annotation Guidelines Swix introduces Arrow's Impossibility Theorem regarding preference aggregation; Nathan acknowledges social choice limitations and explains how annotator guidelines try to bypass it.42:31–46:06 · Guest teaching 5/10 Preference Data Categories and Granular Metadata Swix asks whether safety guardrails should be handled in RLHF or via downstream classifier filters, prompting Nathan to outline industrial deployment trade-offs.46:06–51:32 · Guest teaching 5/10 The High Economic Cost of Human Preference Data Swix and Nathan break down the massive economic costs of human preference data versus synthetic data generation via GPT-4.51:33–55:57 · Guest teaching 6/10 Open Preference Datasets and LMSYS Chatbot Arena Nathan describes the Bradley-Terry preference loss function math and the 65-75% inter-annotator agreement ceiling observed across industry datasets.55:58–59:11 · Guest teaching 7/10 Implementing Proximal Policy Optimization (PPO) for LLMs Nathan deconstructs the multi-model architecture needed for PPO (actor, critic, reference, reward models), emphasizing systems engineering overhead.59:14–1:01:38 · Guest teaching 5/10 Impact of RLHF on Benchmark Capabilities vs. Style Nathan points out that RLHF mostly alters model tone and style rather than factual academic capabilities, citing GPT-4 technical report exam charts.1:01:40–1:05:12 · Guest teaching 6/10 Non-PPO Alternatives: Best-of-N and Rejection Sampling Nathan reviews lightweight non-PPO optimization strategies including Best-of-N inference sampling and iterative rejection sampling.1:05:12–1:09:48 · Guest teaching 8/10 Unpacking Anthropic's Constitutional AI Nathan clarifies Anthropic's Constitutional AI, correcting common misconceptions by explaining its dual-phase nature of instruction revision and principle sampling.1:09:49–1:12:31 · Guest teaching 3/10 Superalignment and Weak-to-Strong Generalization When Nathan notes skepticism around OpenAI's weak-to-strong generalization paper, Swix pushes back with a comprehensive explanation connecting it to superalignment.1:12:32–1:17:40 · Guest teaching 6/10 Direct Preference Optimization (DPO) Mechanics & Debate Nathan explains DPO mechanics and implicit reward formulations, playfully critiquing academic researchers who avoid explaining the intuitive mechanics.1:17:40–1:21:19 · Guest teaching 5/10 AI2's Mission and Scaling TÜLU 2 via DPO Nathan details the Allen Institute's mission and the training run behind Tulu 2 70B using JAX TPUs and DPO recipes.1:21:21–1:26:46 · Guest teaching 5/10 Evaluating LLMs: Leaderboards and Real-World Usage Nathan critiques automated leaderboard gaming and explains the significance of LMSYS Chatbot Arena human Elo rankings.1:26:47–1:29:43 · Guest teaching 6/10 Comparing Chat Evaluation Benchmarks: MT-Bench vs. AlpacaEval Nathan contrasts the evaluation methodologies of MT-Bench (multi-turn GPT-4 scoring) and AlpacaEval (pairwise win rate versus DaVinci-003).1:29:46–1:33:20 · Guest teaching 6/10 Open Frontiers and Outer-Loop Optimization in RLHF Nathan outlines outer-loop RLHF optimization (continual retraining on user prompts and fixes) as described by John Schulman, wrapping up with startup landscape analysis.0:45–3:43 · Guest disagreement 1/10 Ultra-Endurance Sports and Staying Grounded Casual introductory rapport where Swix and Nathan discuss ultra-endurance running, trail access in Berkeley, and Nathan's Substack essays.3:43–8:34 · Guest disagreement 1/10 Motivation for the RLHF Presentation Alessio and Nathan discuss how RL transitioned from traditional robotics to LLM alignment, citing the Llama 2 paper's surprise at RL's efficacy.8:34–13:35 · Guest disagreement 1/10 The Historical and Philosophical Roots of RLHF Nathan details the philosophical and decision-theoretic assumptions underpinning RLHF, citing Aristotle and the Von Neumann-Morgenstern utility theorem.13:36–17:18 · Guest disagreement 2/10 Bridging Social Sciences and AI Engineering Alessio inquires about the intersection of sociology and AI systems engineering, while Nathan explains why RLHF formulation resembles a bandit problem more than traditional RL.17:18–21:04 · Guest disagreement 1/10 Chain-of-Thought Reasoning and Process Reward Models Nathan breaks down process reward models and historical milestones from TAMER (2008) to the Christiano et al. (2017) human preference paper.21:05–23:09 · Guest disagreement 1/10 Signal Richness and Negative Feedback in Alignment Swix highlights the distinction between next-token prediction and negative feedback in RLHF; Nathan agrees it is a richer, more compute-efficient supervisory signal.23:09–27:32 · Guest disagreement 1/10 The Core Assumption: Evaluating vs. Generating Text Nathan reviews the core assumption from OpenAI's 2018–2020 papers that evaluating text is significantly easier than generating high-quality text.27:33–30:05 · Guest disagreement 2/10 Instruction Tuning vs. RLHF in Practical Applications Nathan advises startups against building full RLHF pipelines unless critical to their moat, advocating instruction tuning and DPO instead.30:05–34:06 · Guest disagreement 1/10 The Power and Limits of Instruction Tuning at Scale The conversation touches on early open-source instruction tuning milestones like Vicuna, contrasting small model limits against 70B parameter scaling.34:08–38:26 · Guest disagreement 1/10 Mathematical Formulation of the RLHF Objective Nathan explains the mathematical RLHF objective, detailing the role of the KL-divergence constraint to prevent model mode collapse and reward hacking.38:26–42:30 · Guest disagreement 1/10 Challenges in Preference Aggregation and Annotation Guidelines Swix introduces Arrow's Impossibility Theorem regarding preference aggregation; Nathan acknowledges social choice limitations and explains how annotator guidelines try to bypass it.42:31–46:06 · Guest disagreement 1/10 Preference Data Categories and Granular Metadata Swix asks whether safety guardrails should be handled in RLHF or via downstream classifier filters, prompting Nathan to outline industrial deployment trade-offs.46:06–51:32 · Guest disagreement 1/10 The High Economic Cost of Human Preference Data Swix and Nathan break down the massive economic costs of human preference data versus synthetic data generation via GPT-4.51:33–55:57 · Guest disagreement 1/10 Open Preference Datasets and LMSYS Chatbot Arena Nathan describes the Bradley-Terry preference loss function math and the 65-75% inter-annotator agreement ceiling observed across industry datasets.55:58–59:11 · Guest disagreement 2/10 Implementing Proximal Policy Optimization (PPO) for LLMs Nathan deconstructs the multi-model architecture needed for PPO (actor, critic, reference, reward models), emphasizing systems engineering overhead.59:14–1:01:38 · Guest disagreement 2/10 Impact of RLHF on Benchmark Capabilities vs. Style Nathan points out that RLHF mostly alters model tone and style rather than factual academic capabilities, citing GPT-4 technical report exam charts.1:01:40–1:05:12 · Guest disagreement 1/10 Non-PPO Alternatives: Best-of-N and Rejection Sampling Nathan reviews lightweight non-PPO optimization strategies including Best-of-N inference sampling and iterative rejection sampling.1:05:12–1:09:48 · Guest disagreement 2/10 Unpacking Anthropic's Constitutional AI Nathan clarifies Anthropic's Constitutional AI, correcting common misconceptions by explaining its dual-phase nature of instruction revision and principle sampling.1:09:49–1:12:31 · Guest disagreement 2/10 Superalignment and Weak-to-Strong Generalization When Nathan notes skepticism around OpenAI's weak-to-strong generalization paper, Swix pushes back with a comprehensive explanation connecting it to superalignment.1:12:32–1:17:40 · Guest disagreement 3/10 Direct Preference Optimization (DPO) Mechanics & Debate Nathan explains DPO mechanics and implicit reward formulations, playfully critiquing academic researchers who avoid explaining the intuitive mechanics.1:17:40–1:21:19 · Guest disagreement 1/10 AI2's Mission and Scaling TÜLU 2 via DPO Nathan details the Allen Institute's mission and the training run behind Tulu 2 70B using JAX TPUs and DPO recipes.1:21:21–1:26:46 · Guest disagreement 2/10 Evaluating LLMs: Leaderboards and Real-World Usage Nathan critiques automated leaderboard gaming and explains the significance of LMSYS Chatbot Arena human Elo rankings.1:26:47–1:29:43 · Guest disagreement 1/10 Comparing Chat Evaluation Benchmarks: MT-Bench vs. AlpacaEval Nathan contrasts the evaluation methodologies of MT-Bench (multi-turn GPT-4 scoring) and AlpacaEval (pairwise win rate versus DaVinci-003).1:29:46–1:33:20 · Guest disagreement 1/10 Open Frontiers and Outer-Loop Optimization in RLHF Nathan outlines outer-loop RLHF optimization (continual retraining on user prompts and fixes) as described by John Schulman, wrapping up with startup landscape analysis.0:45–3:43 · The hosts pushing back 0/10 Ultra-Endurance Sports and Staying Grounded Casual introductory rapport where Swix and Nathan discuss ultra-endurance running, trail access in Berkeley, and Nathan's Substack essays.3:43–8:34 · The hosts pushing back 1/10 Motivation for the RLHF Presentation Alessio and Nathan discuss how RL transitioned from traditional robotics to LLM alignment, citing the Llama 2 paper's surprise at RL's efficacy.8:34–13:35 · The hosts pushing back 0/10 The Historical and Philosophical Roots of RLHF Nathan details the philosophical and decision-theoretic assumptions underpinning RLHF, citing Aristotle and the Von Neumann-Morgenstern utility theorem.13:36–17:18 · The hosts pushing back 1/10 Bridging Social Sciences and AI Engineering Alessio inquires about the intersection of sociology and AI systems engineering, while Nathan explains why RLHF formulation resembles a bandit problem more than traditional RL.17:18–21:04 · The hosts pushing back 0/10 Chain-of-Thought Reasoning and Process Reward Models Nathan breaks down process reward models and historical milestones from TAMER (2008) to the Christiano et al. (2017) human preference paper.21:05–23:09 · The hosts pushing back 2/10 Signal Richness and Negative Feedback in Alignment Swix highlights the distinction between next-token prediction and negative feedback in RLHF; Nathan agrees it is a richer, more compute-efficient supervisory signal.23:09–27:32 · The hosts pushing back 0/10 The Core Assumption: Evaluating vs. Generating Text Nathan reviews the core assumption from OpenAI's 2018–2020 papers that evaluating text is significantly easier than generating high-quality text.27:33–30:05 · The hosts pushing back 1/10 Instruction Tuning vs. RLHF in Practical Applications Nathan advises startups against building full RLHF pipelines unless critical to their moat, advocating instruction tuning and DPO instead.30:05–34:06 · The hosts pushing back 1/10 The Power and Limits of Instruction Tuning at Scale The conversation touches on early open-source instruction tuning milestones like Vicuna, contrasting small model limits against 70B parameter scaling.34:08–38:26 · The hosts pushing back 0/10 Mathematical Formulation of the RLHF Objective Nathan explains the mathematical RLHF objective, detailing the role of the KL-divergence constraint to prevent model mode collapse and reward hacking.38:26–42:30 · The hosts pushing back 2/10 Challenges in Preference Aggregation and Annotation Guidelines Swix introduces Arrow's Impossibility Theorem regarding preference aggregation; Nathan acknowledges social choice limitations and explains how annotator guidelines try to bypass it.42:31–46:06 · The hosts pushing back 2/10 Preference Data Categories and Granular Metadata Swix asks whether safety guardrails should be handled in RLHF or via downstream classifier filters, prompting Nathan to outline industrial deployment trade-offs.46:06–51:32 · The hosts pushing back 1/10 The High Economic Cost of Human Preference Data Swix and Nathan break down the massive economic costs of human preference data versus synthetic data generation via GPT-4.51:33–55:57 · The hosts pushing back 0/10 Open Preference Datasets and LMSYS Chatbot Arena Nathan describes the Bradley-Terry preference loss function math and the 65-75% inter-annotator agreement ceiling observed across industry datasets.55:58–59:11 · The hosts pushing back 0/10 Implementing Proximal Policy Optimization (PPO) for LLMs Nathan deconstructs the multi-model architecture needed for PPO (actor, critic, reference, reward models), emphasizing systems engineering overhead.59:14–1:01:38 · The hosts pushing back 1/10 Impact of RLHF on Benchmark Capabilities vs. Style Nathan points out that RLHF mostly alters model tone and style rather than factual academic capabilities, citing GPT-4 technical report exam charts.1:01:40–1:05:12 · The hosts pushing back 0/10 Non-PPO Alternatives: Best-of-N and Rejection Sampling Nathan reviews lightweight non-PPO optimization strategies including Best-of-N inference sampling and iterative rejection sampling.1:05:12–1:09:48 · The hosts pushing back 1/10 Unpacking Anthropic's Constitutional AI Nathan clarifies Anthropic's Constitutional AI, correcting common misconceptions by explaining its dual-phase nature of instruction revision and principle sampling.1:09:49–1:12:31 · The hosts pushing back 4/10 Superalignment and Weak-to-Strong Generalization When Nathan notes skepticism around OpenAI's weak-to-strong generalization paper, Swix pushes back with a comprehensive explanation connecting it to superalignment.1:12:32–1:17:40 · The hosts pushing back 1/10 Direct Preference Optimization (DPO) Mechanics & Debate Nathan explains DPO mechanics and implicit reward formulations, playfully critiquing academic researchers who avoid explaining the intuitive mechanics.1:17:40–1:21:19 · The hosts pushing back 0/10 AI2's Mission and Scaling TÜLU 2 via DPO Nathan details the Allen Institute's mission and the training run behind Tulu 2 70B using JAX TPUs and DPO recipes.1:21:21–1:26:46 · The hosts pushing back 1/10 Evaluating LLMs: Leaderboards and Real-World Usage Nathan critiques automated leaderboard gaming and explains the significance of LMSYS Chatbot Arena human Elo rankings.1:26:47–1:29:43 · The hosts pushing back 0/10 Comparing Chat Evaluation Benchmarks: MT-Bench vs. AlpacaEval Nathan contrasts the evaluation methodologies of MT-Bench (multi-turn GPT-4 scoring) and AlpacaEval (pairwise win rate versus DaVinci-003).1:29:46–1:33:20 · The hosts pushing back 0/10 Open Frontiers and Outer-Loop Optimization in RLHF Nathan outlines outer-loop RLHF optimization (continual retraining on user prompts and fixes) as described by John Schulman, wrapping up with startup landscape analysis.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:16:00 Nathan critiques academic DPO researchers dodging intuitive explanations

Nathan candidly criticizes how academic authors deflect practical questions by telling people to just stare at mathematical equations on poster boards.

Hardest push from the hosts ▶ 1:10:20 Swix challenges dismissive view on weak-to-strong generalization

When Nathan brushes aside OpenAI's weak-to-strong paper as potential safety-washing, Swix firmly intervenes to articulate the exact conceptual lineage connecting Constitutional AI to superalignment.

Biggest teaching moment ▶ 1:06:00 Nathan demystifies how Constitutional AI actually operates

Nathan breaks down how Constitutional AI is widely misunderstood, educating the hosts on its two distinct phases: principle-based instruction revision followed by sampling constitutional prompts.

The host holds their own ▶ 38:46 Swix introduces Arrow's Impossibility Theorem into the RLHF framework

Swix demonstrates advanced interdisciplinary domain knowledge by directly connecting welfare economics' Arrow Impossibility Theorem to the foundational flaws of pairwise preference aggregation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Ultra-Endurance Sports and Staying Grounded 3110 Casual introductory rapport where Swix and Nathan discuss ultra-endurance running, trail access in Berkeley, and Nathan's Substack essays.
Motivation for the RLHF Presentation 4411 Alessio and Nathan discuss how RL transitioned from traditional robotics to LLM alignment, citing the Llama 2 paper's surprise at RL's efficacy.
The Historical and Philosophical Roots of RLHF 3610 Nathan details the philosophical and decision-theoretic assumptions underpinning RLHF, citing Aristotle and the Von Neumann-Morgenstern utility theorem.
Bridging Social Sciences and AI Engineering 4521 Alessio inquires about the intersection of sociology and AI systems engineering, while Nathan explains why RLHF formulation resembles a bandit problem more than traditional RL.
Chain-of-Thought Reasoning and Process Reward Models 4610 Nathan breaks down process reward models and historical milestones from TAMER (2008) to the Christiano et al. (2017) human preference paper.
Signal Richness and Negative Feedback in Alignment 6412 Swix highlights the distinction between next-token prediction and negative feedback in RLHF; Nathan agrees it is a richer, more compute-efficient supervisory signal.
The Core Assumption: Evaluating vs. Generating Text 4510 Nathan reviews the core assumption from OpenAI's 2018–2020 papers that evaluating text is significantly easier than generating high-quality text.
Instruction Tuning vs. RLHF in Practical Applications 5621 Nathan advises startups against building full RLHF pipelines unless critical to their moat, advocating instruction tuning and DPO instead.
The Power and Limits of Instruction Tuning at Scale 5511 The conversation touches on early open-source instruction tuning milestones like Vicuna, contrasting small model limits against 70B parameter scaling.
Mathematical Formulation of the RLHF Objective 4710 Nathan explains the mathematical RLHF objective, detailing the role of the KL-divergence constraint to prevent model mode collapse and reward hacking.
Challenges in Preference Aggregation and Annotation Guidelines 7412 Swix introduces Arrow's Impossibility Theorem regarding preference aggregation; Nathan acknowledges social choice limitations and explains how annotator guidelines try to bypass it.
Preference Data Categories and Granular Metadata 6512 Swix asks whether safety guardrails should be handled in RLHF or via downstream classifier filters, prompting Nathan to outline industrial deployment trade-offs.
The High Economic Cost of Human Preference Data 6511 Swix and Nathan break down the massive economic costs of human preference data versus synthetic data generation via GPT-4.
Open Preference Datasets and LMSYS Chatbot Arena 4610 Nathan describes the Bradley-Terry preference loss function math and the 65-75% inter-annotator agreement ceiling observed across industry datasets.
Implementing Proximal Policy Optimization (PPO) for LLMs 3720 Nathan deconstructs the multi-model architecture needed for PPO (actor, critic, reference, reward models), emphasizing systems engineering overhead.
Impact of RLHF on Benchmark Capabilities vs. Style 5521 Nathan points out that RLHF mostly alters model tone and style rather than factual academic capabilities, citing GPT-4 technical report exam charts.
Non-PPO Alternatives: Best-of-N and Rejection Sampling 4610 Nathan reviews lightweight non-PPO optimization strategies including Best-of-N inference sampling and iterative rejection sampling.
Unpacking Anthropic's Constitutional AI 4821 Nathan clarifies Anthropic's Constitutional AI, correcting common misconceptions by explaining its dual-phase nature of instruction revision and principle sampling.
Superalignment and Weak-to-Strong Generalization 7324 When Nathan notes skepticism around OpenAI's weak-to-strong generalization paper, Swix pushes back with a comprehensive explanation connecting it to superalignment.
Direct Preference Optimization (DPO) Mechanics & Debate 5631 Nathan explains DPO mechanics and implicit reward formulations, playfully critiquing academic researchers who avoid explaining the intuitive mechanics.
AI2's Mission and Scaling TÜLU 2 via DPO 4510 Nathan details the Allen Institute's mission and the training run behind Tulu 2 70B using JAX TPUs and DPO recipes.
Evaluating LLMs: Leaderboards and Real-World Usage 5521 Nathan critiques automated leaderboard gaming and explains the significance of LMSYS Chatbot Arena human Elo rankings.
Comparing Chat Evaluation Benchmarks: MT-Bench vs. AlpacaEval 4610 Nathan contrasts the evaluation methodologies of MT-Bench (multi-turn GPT-4 scoring) and AlpacaEval (pairwise win rate versus DaVinci-003).
Open Frontiers and Outer-Loop Optimization in RLHF 5610 Nathan outlines outer-loop RLHF optimization (continual retraining on user prompts and fixes) as described by John Schulman, wrapping up with startup landscape analysis.

Statements from this episode (57)

Opinion
Lambert: OpenAI's rumored Q* was likely just a moderate benchmark bump
“They probably just got like a moderate bump on one of their benchmarks, and then everyone lost their minds, so it doesn't really matter.”
Nathan Lambert Jan 11, 2024 ▶ 3:34
Insight
Lambert: AI field hasn't learned much more, just articulates ignorance better
“I think it's, I try to do it every six or 12 months is my current, is my estimated cadence, just to refine the ways that I say things, and people will see That we don't know that much more, but we have a bit of better way of saying what we don't know.”
Nathan Lambert Jan 11, 2024 ▶ 3:58
Opinion
Lambert: Major tech companies now realize they need dedicated RLHF teams
“I think any major company that wasn't doing RLHF is now realizing they have to have a team around this.”
Nathan Lambert Jan 11, 2024 ▶ 5:31
Assertion Not checkable as stated
Lambert: Open-source and research communities lack large-scale RLHF capabilities
“At the same time, we don't have a lot of that in the like open and research communities at the same scale.”
Nathan Lambert Jan 11, 2024 ▶ 5:38
Opinion
Lambert: Anthropic is considered the master of RLHF techniques
“Everyone knows Anthropik is kind of the masters of this”
Nathan Lambert Jan 11, 2024 ▶ 5:49
Prediction Not checkable as stated
Nathan Lambert: RL will remain a distinct field from language modeling
“I think in the long run it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about, that it's just viewing the whole problem formulation different than predicting text, really,…”
Nathan Lambert Jan 11, 2024 ▶ 7:40
Opinion
Lambert: Reinforcement learning in language models is contrived and not real RL
“And the view of RL in language models is pretty contrived already. So it's not like we're doing real RL.”
Nathan Lambert Jan 11, 2024 ▶ 7:59
Insight
Lambert: RLHF reward models lack inductive biases for true preferences
“In RLHF calling things a preference model is a little annoying because There's no inductive bias of what a preference is.”
Nathan Lambert Jan 11, 2024 ▶ 12:41
Insight
Lambert: Claude's constitution dictates output priorities, not model beliefs
“If you look at Claude's constitution, like, that doesn't mean the model believes these things. It's just trying Trained and to prioritize these things.”
Nathan Lambert Jan 11, 2024 ▶ 13:06
Assertion Not checkable as stated
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Nathan Lambert Jan 11, 2024 ▶ 14:06
Assertion Not checkable as stated
Lambert: Real-Time Human Feedback in LLM RL Loops Is Far From Feasible
“Setting up the infrastructure to take tens of thousands of prompts and generate them and then show them to a human and collect the human responses and then show that Shove that into your training architecture is very far away from working, so we don't really h…”
Nathan Lambert Jan 11, 2024 ▶ 16:13
Insight
Lambert: RLHF operates as a contextual bandit problem, not sequential RL
“What happens is the state is a prompt, and then you do a completion, and then you throw it away, and you grab a new prompt. Where really in like RL, you, as an RL researcher, you would think of this as being like, you take a state, you complete, Get some compl…”
Nathan Lambert Jan 11, 2024 ▶ 16:40
Prediction Not checkable as stated
Lambert: AI community will clarify if chain-of-thought maps to RL within a year
“I think in the next year that'll probably get kind of made more concrete by the community on, like, if you can easily draw out, like, if chain of thought reasoning is more like RL.”
Nathan Lambert Jan 11, 2024 ▶ 18:00
Insight
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Nathan Lambert Jan 11, 2024 ▶ 22:14
Insight
Lambert: Instruction tuning is more important than RLHF for most practitioners
“I think for most people, instruction tuning is probably still more important in their day-to-day life. I think instruction tuning works very well. You can write samples by hand that make sense. You can get the model to learn from them. You could do this with v…”
Nathan Lambert Jan 11, 2024 ▶ 28:11
Assertion Supported
Lambert: DPO benchmark gains rely largely on the UltraFeedback dataset
“Everyone's using this ultra feedback data set and it boosts AlpacaVal, MTBench, TruthfulQA, and like the qualitative model a bit. We don't really know why.”
Nathan Lambert Jan 11, 2024 ▶ 29:20
Insight
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Nathan Lambert Jan 11, 2024 ▶ 29:37
Opinion
Lambert: Open-source claims of ChatGPT-level performance are overblown
“I think the claims of ChatGPT level are long overblown in most of the things in open source.”
Nathan Lambert Jan 11, 2024 ▶ 30:44
Insight
Lambert: Scaling from 7B to 70B parameters fixes nuance and repetition
“I think the things that people see now is like the small models don't really handle nuance as well, and they could be more repetitive if, even if they have really good instruction tuning, but if you take that kind of seven to seventy billion parameter jump, li…”
Nathan Lambert Jan 11, 2024 ▶ 31:24
Opinion
Lambert: The vast majority of instruction tuning data remains simple Q&A
“There's much more, like there's surely kind of more tricky things that people do, but I still think the vast majority of it is question and answer. It's like, please explain this topic to me, generate this thing for me. That hasn't changed that much this year.…”
Nathan Lambert Jan 11, 2024 ▶ 33:49
Insight
Lambert: RLHF fails completely without a KL divergence constraint
“The KL constraint prevents that. There's not that much documented work on that, but there's a lot of people that know if you take that away, it just doesn't work at all.”
Nathan Lambert Jan 11, 2024 ▶ 35:51
Assertion Supported
Lambert: Training LLM reward models on 0-to-10 ratings failed
“People tried that with language models, which is if you have a prompt and a completion and you just have someone rate it from zero to 10, could you then train a reward model on all of these completions and zero to 10 ratings and see if you could actually chang…”
Nathan Lambert Jan 11, 2024 ▶ 37:43
Insight
Lambert: RLHF avoids controversial preferences, focusing on correctness and style
“The reason this really is done on a deep level is that you're not actually trying to model any, like, contestable preference in this. Like, you're not trying to go into things that are controversial or anything. It's really, like, the notion of preference is t…”
Nathan Lambert Jan 11, 2024 ▶ 39:02
Opinion
Lambert: Annotator disagreement data in RLHF is not used for anything
“I don't really think this disagreement data is used for anything, but it's good to know, like, what the distribution of prompts is, who's doing it, how many samples you have, controlling the workforce.”
Nathan Lambert Jan 11, 2024 ▶ 42:20
Opinion
Lambert: OpenAI trains its RLHF models primarily on good-vs-good answer comparisons
“I think open AIs of the world are all in good answer, and have learned to eliminate everything else.”
Nathan Lambert Jan 11, 2024 ▶ 43:19
Assertion Supported
Lambert: Anthropic, ChatGPT, and Bard Use Post-Generation Moderation Classifiers
“Anthropic and ChatGPT and Bard almost surely have a classifier after, which is like, is this text good? Is this text bad?”
Nathan Lambert Jan 11, 2024 ▶ 45:13
Insight
Lambert: A 100x smaller language model filters output better than RLHF
“You could use like a hundred times smaller language model and do much better at filtering than RLHF”
Nathan Lambert Jan 11, 2024 ▶ 45:20
Assertion Not checkable as stated
Lambert: Meta spent roughly $6M to $8M on Llama 2 preference data
“So I would say, still say, like, six to eight million is safe to say that they're spending, if not more, they're probably also buying other types of data and or throwing out data that they don't like.”
Nathan Lambert Jan 11, 2024 ▶ 46:51
Insight
Lambert: Published AI training compute costs understate total experimentation budgets
“The compute costs listed in the paper always are way lower, because all they have to say is, like, what is one run cost, but they're running tens or hundreds of runs, so it's like, okay, like, it's kind of a meaningless number, yeah.”
Nathan Lambert Jan 11, 2024 ▶ 47:04
Prediction Not checkable as stated
Lambert: Open source will learn to train models on arbitrary preference data
“I really think people in open source and academics are going to figure out how to use any preference data on any model just because they're scrappy.”
Nathan Lambert Jan 11, 2024 ▶ 47:48
Opinion
Lambert: Only 20% to 40% of Meta's Llama RLHF data is useful
“I do think that if we had all the llama data, we wouldn't know what to do with all of it. Like, probably, like, 20 to 40% would be pretty useful for people, but not the whole data set. Like, a lot of it's probably kind of gibberish, because they had a lot of d…”
Nathan Lambert Jan 11, 2024 ▶ 48:20
Assertion Supported
Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans
“Essentially, people also think that synthetic data is, like, GPT-IV is more accurate than humans at labeling preferences, so if you look at these diagrams, like, humans are about 60 to 70% agreement, or, like, that's what the models get to, and if humans are a…”
Nathan Lambert Jan 11, 2024 ▶ 48:53
Prediction Not checkable as stated
Lambert: OpenAI Will Not Aggressively Ban Synthetic Training Scraping
“I don't expect OpenAI to go too crazy on this, because they're just gonna, there's gonna be so much backlash against them.”
Nathan Lambert Jan 11, 2024 ▶ 50:31
Prediction Held up
Lambert: Practitioners Will Adopt Constitutional AI for Preferences in 2024
“I think in twenty-twenty-four at some point people will start doing things like constitutional AI for preferences.”
Nathan Lambert Jan 11, 2024 ▶ 51:25
Assertion Not checkable as stated
Lambert: Zephyr was the first open model to succeed with public RLHF
“I think Zephyr was the first model that showed success with RLHF in the public, but that's a long time from everyone knowing that it was something that people are interested in to having any, like, check mark.”
Nathan Lambert Jan 11, 2024 ▶ 51:42
Opinion
Lambert: Chatbot Arena is the best available evaluation benchmark for LLMs
“I have, if we make it to evaluation, I'd pretty much say that Chat Arena is the best limited evaluation that people have to learn how to use language models, and like, It's very valuable data”
Nathan Lambert Jan 11, 2024 ▶ 52:38
Assertion Supported
Lambert: Anthropic and OpenAI reward model loss functions are mathematically identical
“Fun fact is that these loss functions Look different and anthropic in opening eyes papers, but they're just literally just log transform. So if you start like expantiating both sides and taking the log of both sides, you'll like converge on one of the two, the…”
Nathan Lambert Jan 11, 2024 ▶ 54:41
Assertion Supported
Lambert: RLHF reward models achieve only 65% to 75% validation agreement
“If you look at a test set, you'll have a chosen and rejected, and you can take the reward model you're training, pass in those completions, And you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward, and this …”
Nathan Lambert Jan 11, 2024 ▶ 54:59
Insight
Lambert: RLHF Performance Depends on Data and Systems Over RL Details
“It really ends up being kind of, like, gibberish that I think is less important now, because it's more about data and infrastructure than RL details, than, like, value functions and everything. A lot of the papers have different terms in the equations. I think…”
Nathan Lambert Jan 11, 2024 ▶ 57:58
Insight
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Nathan Lambert Jan 11, 2024 ▶ 58:36
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53
Opinion
Lambert: Scaling PPO in RLHF is a Nightmare
“PPO is kind of a nightmare to scale.”
Nathan Lambert Jan 11, 2024 ▶ 1:01:41
Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”
Nathan Lambert Jan 11, 2024 ▶ 1:03:01
Prediction Not checkable as stated
Lambert: Offline RL could take off due to simpler training pipelines
“There's a few papers that people have published Not a lot of traction. I think it could take off. Some people that I know in the RLHF area really think a lot of people are doing this in industry just because it makes the kind of training process simpler and th…”
Nathan Lambert Jan 11, 2024 ▶ 1:04:08
Prediction Not checkable as stated
Lambert: RL feedback mechanisms will specialize across distinct task domains
“It seems very likely that different feedback will be used for different domains. Chain of thought reasoning is great. For math, and that's where these process reward models are being designed. Probably not great for things like poetry, but as any tool gets bet…”
Nathan Lambert Jan 11, 2024 ▶ 1:04:51
Assertion Supported
Lambert: Most open-source RLHF training runs only last a few epochs
“Most RLHF is only a few epochs, at least in the open models”
Nathan Lambert Jan 11, 2024 ▶ 1:09:04
Insight
Lambert: Anthropic Constitutional AI and OpenAI Superalignment share intellectual roots
“The constitutional AI and the super alignment is, like, very conceptually linked. It's like a group of people that has, like, a very similar intellectual upbringing, and they work together for a long time, like, coming to the same conclusions in different ways…”
Nathan Lambert Jan 11, 2024 ▶ 1:11:37
Insight
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think DPO is closer to RLHF than RLHF is to RL.”
Nathan Lambert Jan 11, 2024 ▶ 1:12:56
Prediction Held up
Lambert: More DPO models will emerge than any other method
“I expect to see more DPO models than anything else in the next six months.”
Nathan Lambert Jan 11, 2024 ▶ 1:15:07
Assertion Not checkable as stated
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Nathan Lambert Jan 11, 2024 ▶ 1:19:31
Disclosure
AI2 Plans to Release Fully Open Pre-Trained LLMs With Data and Code
“The Allen Institute is training, pre-training language models, or pre-training, like, open language models, where we'll be able to share, like, data, code, everything, the kind of horn that everyone likes to get annoyed about these days, it's like, well, I'm n…”
Nathan Lambert Jan 11, 2024 ▶ 1:20:08
Assertion Supported
Lambert: DPO Has Become the Standard Release Expectation for Open-Source LLMs
“I think DPO releases are kind of becoming expected because Mistral released a DPO model as well. I think the slide after this is just like, there's a ton. It's like Intel releases DPO models, Stability releases DPO models. At some point, you just have to accep…”
Nathan Lambert Jan 11, 2024 ▶ 1:20:43
Assertion Supported
Lambert: GPT-4 Turbo Showed a Noticeable Jump on LMSYS Chatbot Arena
“GPT-IV Turbo is also notably ahead of the other GPT-IVs, which it kind of showed up immediately once they added it to the leaderboard, or to the arena, and I was like, all the GPT-IV memes aside, it seems like this is effectively a bump in the model.”
Nathan Lambert Jan 11, 2024 ▶ 1:25:01
Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Nathan Lambert Jan 11, 2024 ▶ 1:26:04
Opinion
Lambert: Frontier Labs Lack Visibility into Cross-Model RLHF Sensitivity
“I think big labs are so over-indexed, are indexed on their own base models, so they don't know, like, what's swapping between CloudBase or GPT-IV-Base, how that would change any notion of preference or what you do with RLHF.”
Nathan Lambert Jan 11, 2024 ▶ 1:30:15
Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27
Opinion
Lambert: Scale AI has historically struggled to retain technical ML talent
“I think they've historically had trouble keeping, like, technical ML talent, but they've started a new research lab, so that should help.”
Nathan Lambert Jan 11, 2024 ▶ 1:34:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.