Oct 9, 2025 · 36m · no-priors

No Priors Ep. 135 | With Humans& Founder Eric Zelikman

Eric Zelikman · 25m spoken Sarah Guo · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

AI researcher and Humans& founder Eric Zelikman discusses the evolution of reasoning models from STaR to frontier systems, advocating for collaborative, memory-rich AI that amplifies human agency rather than pursuing purely autonomous task automation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.9% of the talking time here. How this is scored →

The hosts as informed peer 5.8 Guest teaching 3.8 Guest disagreement 2.0 The hosts pushing back 3.0
05100:0010:0020:0030:002:06–7:55 · The hosts as informed peer 5/10 Bootstrapping Reasoning Models with STaR and Quiet-STaR Sarah prompts Eric to explain the technical intuition behind his seminal STaR and Quiet-STaR papers. Eric walks through the iterative reasoning and policy gradient mechanics in an accessible, collaborative tutorial style without tension.7:55–14:48 · The hosts as informed peer 6/10 Evaluating Modern Model Capabilities and Practical Task Verifiability Sarah pushes Eric to quantify 'reasonably smart' with human benchmarks and probes why models fail at verifiable coding tasks. Eric notes responsiveness trade-offs and domain shifts between RL datasets and real-world codebases.14:50–22:07 · The hosts as informed peer 6/10 Autonomous AI Limits and the Case for Collaboration Sarah challenges Eric by arguing that major labs intentionally design humans out of the loop for scaling efficiency. Eric pushes back against the autonomous agent paradigm, arguing that long-horizon models operating without human feedback lead to loss of agency and unmaintainable code.22:08–34:20 · The hosts as informed peer 6/10 Founding Humans&: Advancing EQ, Multi-Turn Interaction, and Memory Sarah raises counterarguments against modeling individual users, citing human self-inconsistency and distribution shift. Eric explains why current lab structures prioritize single-turn verifiable benchmarks over multi-turn memory and collaboration.2:06–7:55 · Guest teaching 4/10 Bootstrapping Reasoning Models with STaR and Quiet-STaR Sarah prompts Eric to explain the technical intuition behind his seminal STaR and Quiet-STaR papers. Eric walks through the iterative reasoning and policy gradient mechanics in an accessible, collaborative tutorial style without tension.7:55–14:48 · Guest teaching 4/10 Evaluating Modern Model Capabilities and Practical Task Verifiability Sarah pushes Eric to quantify 'reasonably smart' with human benchmarks and probes why models fail at verifiable coding tasks. Eric notes responsiveness trade-offs and domain shifts between RL datasets and real-world codebases.14:50–22:07 · Guest teaching 3/10 Autonomous AI Limits and the Case for Collaboration Sarah challenges Eric by arguing that major labs intentionally design humans out of the loop for scaling efficiency. Eric pushes back against the autonomous agent paradigm, arguing that long-horizon models operating without human feedback lead to loss of agency and unmaintainable code.22:08–34:20 · Guest teaching 4/10 Founding Humans&: Advancing EQ, Multi-Turn Interaction, and Memory Sarah raises counterarguments against modeling individual users, citing human self-inconsistency and distribution shift. Eric explains why current lab structures prioritize single-turn verifiable benchmarks over multi-turn memory and collaboration.2:06–7:55 · Guest disagreement 1/10 Bootstrapping Reasoning Models with STaR and Quiet-STaR Sarah prompts Eric to explain the technical intuition behind his seminal STaR and Quiet-STaR papers. Eric walks through the iterative reasoning and policy gradient mechanics in an accessible, collaborative tutorial style without tension.7:55–14:48 · Guest disagreement 2/10 Evaluating Modern Model Capabilities and Practical Task Verifiability Sarah pushes Eric to quantify 'reasonably smart' with human benchmarks and probes why models fail at verifiable coding tasks. Eric notes responsiveness trade-offs and domain shifts between RL datasets and real-world codebases.14:50–22:07 · Guest disagreement 3/10 Autonomous AI Limits and the Case for Collaboration Sarah challenges Eric by arguing that major labs intentionally design humans out of the loop for scaling efficiency. Eric pushes back against the autonomous agent paradigm, arguing that long-horizon models operating without human feedback lead to loss of agency and unmaintainable code.22:08–34:20 · Guest disagreement 2/10 Founding Humans&: Advancing EQ, Multi-Turn Interaction, and Memory Sarah raises counterarguments against modeling individual users, citing human self-inconsistency and distribution shift. Eric explains why current lab structures prioritize single-turn verifiable benchmarks over multi-turn memory and collaboration.2:06–7:55 · The hosts pushing back 1/10 Bootstrapping Reasoning Models with STaR and Quiet-STaR Sarah prompts Eric to explain the technical intuition behind his seminal STaR and Quiet-STaR papers. Eric walks through the iterative reasoning and policy gradient mechanics in an accessible, collaborative tutorial style without tension.7:55–14:48 · The hosts pushing back 3/10 Evaluating Modern Model Capabilities and Practical Task Verifiability Sarah pushes Eric to quantify 'reasonably smart' with human benchmarks and probes why models fail at verifiable coding tasks. Eric notes responsiveness trade-offs and domain shifts between RL datasets and real-world codebases.14:50–22:07 · The hosts pushing back 4/10 Autonomous AI Limits and the Case for Collaboration Sarah challenges Eric by arguing that major labs intentionally design humans out of the loop for scaling efficiency. Eric pushes back against the autonomous agent paradigm, arguing that long-horizon models operating without human feedback lead to loss of agency and unmaintainable code.22:08–34:20 · The hosts pushing back 4/10 Founding Humans&: Advancing EQ, Multi-Turn Interaction, and Memory Sarah raises counterarguments against modeling individual users, citing human self-inconsistency and distribution shift. Eric explains why current lab structures prioritize single-turn verifiable benchmarks over multi-turn memory and collaboration.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 23.3% · guest 76.7%0:00 · the hosts 23.3% · guest 76.7%3:00 · the hosts 14.4% · guest 85.6%3:00 · the hosts 14.4% · guest 85.6%6:00 · the hosts 22.6% · guest 77.4%6:00 · the hosts 22.6% · guest 77.4%9:00 · the hosts 23.5% · guest 76.5%9:00 · the hosts 23.5% · guest 76.5%12:00 · the hosts 32.7% · guest 67.3%12:00 · the hosts 32.7% · guest 67.3%15:00 · the hosts 23.3% · guest 76.7%15:00 · the hosts 23.3% · guest 76.7%18:00 · the hosts 16.6% · guest 83.4%18:00 · the hosts 16.6% · guest 83.4%21:00 · the hosts 23% · guest 77%21:00 · the hosts 23% · guest 77%24:00 · the hosts 11.1% · guest 88.9%24:00 · the hosts 11.1% · guest 88.9%27:00 · the hosts 32% · guest 68%27:00 · the hosts 32% · guest 68%30:00 · the hosts 19.4% · guest 80.6%30:00 · the hosts 19.4% · guest 80.6%33:00 · the hosts 42.1% · guest 57.9%33:00 · the hosts 42.1% · guest 57.9%36:00 · the hosts 33.5% · guest 66.5%36:00 · the hosts 33.5% · guest 66.5%
Sharpest disagreement ▶ 20:41 Eric rejects the autonomous solo-AI paradigm

Eric explicitly disagrees with mainstream AI researchers, rejecting the prevailing vision of autonomous models solving fundamental problems in isolation over multi-hour runs.

Hardest push from the hosts ▶ 32:49 Sarah challenges feasibility with human self-inconsistency objection

Sarah directly confronts Eric's thesis on modeling users by arguing that humans are unpredictable, inconsistent 'unique snowflakes' that cannot easily be brought into distribution.

Biggest teaching moment ▶ 24:25 Eric explains how benchmark credit assignment distorts research

Eric provides an insider perspective from lab researchers explaining why the industry remains fixated on single-turn tasks due to internal compute and resource allocation metrics.

The host holds their own ▶ 13:41 Sarah identifies unreleased RL distribution gaps

Sarah articulates why specialized coding agents fail in production, detailing how closed proprietary RL data leaves developers stranded when codebases depart from internet distribution.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Bootstrapping Reasoning Models with STaR and Quiet-STaR 5411 Sarah prompts Eric to explain the technical intuition behind his seminal STaR and Quiet-STaR papers. Eric walks through the iterative reasoning and policy gradient mechanics in an accessible, collaborative tutorial style without tension.
Evaluating Modern Model Capabilities and Practical Task Verifiability 6423 Sarah pushes Eric to quantify 'reasonably smart' with human benchmarks and probes why models fail at verifiable coding tasks. Eric notes responsiveness trade-offs and domain shifts between RL datasets and real-world codebases.
Autonomous AI Limits and the Case for Collaboration 6334 Sarah challenges Eric by arguing that major labs intentionally design humans out of the loop for scaling efficiency. Eric pushes back against the autonomous agent paradigm, arguing that long-horizon models operating without human feedback lead to loss of agency and unmaintainable code.
Founding Humans&: Advancing EQ, Multi-Turn Interaction, and Memory 6424 Sarah raises counterarguments against modeling individual users, citing human self-inconsistency and distribution shift. Eric explains why current lab structures prioritize single-turn verifiable benchmarks over multi-turn memory and collaboration.

Statements from this episode (7)

Insight
Zelikman: Training reasoning models on just positive examples causes a plateau
“So if you only train on like the positive examples, then you end up in this kind of like potential minimum where there's just no more data that it can actually solve.”
Eric Zelikman Oct 9, 2025 ▶ 5:44
Assertion Supported
Zelikman: Frontier AI models solve questions that stump actual PhD researchers
“Some of the HLE questions that these models are able to solve are genuinely things that are, like, non-trivial for, like, actual, like, PhD researchers.”
Eric Zelikman Oct 9, 2025 ▶ 8:53
Opinion
Zelikman: Current frontier AI models fundamentally lack emotional intelligence
“One of the core things is that they're not smart Like, emotionally, or, like, they're not smart on the level of, like, actually understanding kind of what people care about, or kind of, like, how to actually, like, help people accomplish the things that they c…”
Eric Zelikman Oct 9, 2025 ▶ 10:04
Insight
Zelikman: AI performance gap persists between verifiable and non-verifiable tasks
“There's still a gap between how well these models perform on verifiable tasks versus not verifiable tasks.”
Eric Zelikman Oct 9, 2025 ▶ 14:40
Assertion Supported
Zelikman: Language models can be trained to simulate students for test design
“Like, even back in my PhD, I think one of my, I guess, less well-known works was actually about, we showed that you can train language models to simulate different kinds of students. For tests. Yeah, yeah. And by simulating students, you can actually design be…”
Eric Zelikman Oct 9, 2025 ▶ 22:54
Opinion
Zelikman: Most AI labs treat humans as mere intermediates to full automation
“And maybe this is a strong statement, but I say for most labs, like the human is kind of, you know, the intermediate until you have like this fully automated, like, you know, system. And so spending a lot of time optimizing things for being really good at unde…”
Eric Zelikman Oct 9, 2025 ▶ 29:26
Insight
Zelikman: AI field underinvests in memory due to task-centric training regimes
“I would say that memory is definitely like a feature that has been under, under-invested in by the field. But I would say that it is kind of difficult to invest in memory in this very, like, task-centric regime. Because if you have, like, A bunch of these, lik…”
Eric Zelikman Oct 9, 2025 ▶ 32:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.