Jul 29, 2025 · 40m · y-combinator

Scaling and the Road to Human-Level AI | Anthropic Co-founder Jared Kaplan · Y Combinator

Jared Kaplan · 32m spoken Diana Hu · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At Y Combinator's Startup School, Anthropic co-founder Jared Kaplan explores how empirical scaling laws, expanding autonomous task horizons, and theoretical physics heuristics drive the continuous development of modern artificial intelligence toward human-level performance.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The partners hold 10.1% of the talking time here. How this is scored →

The partners as informed peer 1.6 Guest teaching 0.9 Guest disagreement 0.3 The partners pushing back 0.5
05100:0015:0030:002:21–6:09 · The partners as informed peer 0/10 Core Mechanics of Pretraining and Reinforcement Learning in AI Solo presentation by Jared Kaplan explaining the core foundations of pre-training and reinforcement learning scaling laws. The host is not on stage.6:09–8:19 · The partners as informed peer 0/10 Scaling Laws in Reinforcement Learning and Performance Metrics Kaplan continues his solo keynote, breaking down RL scaling laws and board game benchmarks like Hex and AlphaGo. There is no host involvement.8:19–11:09 · The partners as informed peer 0/10 AI Capability Axes: Modalities and Task Time Horizon Lengths Monologue segment detailing the two capability axes: modality breadth and task horizon lengths doubling every seven months.11:09–15:33 · The partners as informed peer 0/10 Key Frontiers for Human-Level AI and Practical Startup Recommendations Kaplan concludes his presentation with requirements for human-level AI and advice for startup founders. Host is not present.15:33–18:39 · The partners as informed peer 2/10 Y Combinator Batch Application Announcement with Garry Tan Opens with a brief YC ad voiceover before Diana Hu joins the stage and asks Kaplan about Claude 4's release and its compounding capabilities over the next 12 months.18:39–24:36 · The partners as informed peer 3/10 Human-AI Collaboration Dynamics, End-to-End Workflows, and Scientific Research Diana connects Kaplan's ideas to real YC batch trends moving from co-pilots to full workflow automation and references Dario Amodei's essay 'Machines of Loving Grace'. Kaplan elaborates collaboratively on intelligence breadth versus depth.24:36–30:06 · The partners as informed peer 4/10 Applying Physics Heuristics, Large Matrix Limits, and Interpretability Diana draws upon physics terminology like renormalization and symmetry to ask how theoretical physics heuristics informed AI research. Kaplan clarifies that the main physics transfer was big-matrix approximations and asking simple, naive questions rather than complex formalisms.30:06–40:47 · The partners as informed peer 4/10 Scaling Limits, Compute Efficiency, and Audience Questions on Horizon Lengths Diana asks contrarian questions about scaling breakdowns, compute limits (FP4/ternary representations), and invokes Jevons paradox. Audience members then challenge Kaplan on exponential task horizon jumps versus linear log scaling.2:21–6:09 · Guest teaching 0/10 Core Mechanics of Pretraining and Reinforcement Learning in AI Solo presentation by Jared Kaplan explaining the core foundations of pre-training and reinforcement learning scaling laws. The host is not on stage.6:09–8:19 · Guest teaching 0/10 Scaling Laws in Reinforcement Learning and Performance Metrics Kaplan continues his solo keynote, breaking down RL scaling laws and board game benchmarks like Hex and AlphaGo. There is no host involvement.8:19–11:09 · Guest teaching 0/10 AI Capability Axes: Modalities and Task Time Horizon Lengths Monologue segment detailing the two capability axes: modality breadth and task horizon lengths doubling every seven months.11:09–15:33 · Guest teaching 0/10 Key Frontiers for Human-Level AI and Practical Startup Recommendations Kaplan concludes his presentation with requirements for human-level AI and advice for startup founders. Host is not present.15:33–18:39 · Guest teaching 1/10 Y Combinator Batch Application Announcement with Garry Tan Opens with a brief YC ad voiceover before Diana Hu joins the stage and asks Kaplan about Claude 4's release and its compounding capabilities over the next 12 months.18:39–24:36 · Guest teaching 2/10 Human-AI Collaboration Dynamics, End-to-End Workflows, and Scientific Research Diana connects Kaplan's ideas to real YC batch trends moving from co-pilots to full workflow automation and references Dario Amodei's essay 'Machines of Loving Grace'. Kaplan elaborates collaboratively on intelligence breadth versus depth.24:36–30:06 · Guest teaching 2/10 Applying Physics Heuristics, Large Matrix Limits, and Interpretability Diana draws upon physics terminology like renormalization and symmetry to ask how theoretical physics heuristics informed AI research. Kaplan clarifies that the main physics transfer was big-matrix approximations and asking simple, naive questions rather than complex formalisms.30:06–40:47 · Guest teaching 2/10 Scaling Limits, Compute Efficiency, and Audience Questions on Horizon Lengths Diana asks contrarian questions about scaling breakdowns, compute limits (FP4/ternary representations), and invokes Jevons paradox. Audience members then challenge Kaplan on exponential task horizon jumps versus linear log scaling.2:21–6:09 · Guest disagreement 0/10 Core Mechanics of Pretraining and Reinforcement Learning in AI Solo presentation by Jared Kaplan explaining the core foundations of pre-training and reinforcement learning scaling laws. The host is not on stage.6:09–8:19 · Guest disagreement 0/10 Scaling Laws in Reinforcement Learning and Performance Metrics Kaplan continues his solo keynote, breaking down RL scaling laws and board game benchmarks like Hex and AlphaGo. There is no host involvement.8:19–11:09 · Guest disagreement 0/10 AI Capability Axes: Modalities and Task Time Horizon Lengths Monologue segment detailing the two capability axes: modality breadth and task horizon lengths doubling every seven months.11:09–15:33 · Guest disagreement 0/10 Key Frontiers for Human-Level AI and Practical Startup Recommendations Kaplan concludes his presentation with requirements for human-level AI and advice for startup founders. Host is not present.15:33–18:39 · Guest disagreement 0/10 Y Combinator Batch Application Announcement with Garry Tan Opens with a brief YC ad voiceover before Diana Hu joins the stage and asks Kaplan about Claude 4's release and its compounding capabilities over the next 12 months.18:39–24:36 · Guest disagreement 0/10 Human-AI Collaboration Dynamics, End-to-End Workflows, and Scientific Research Diana connects Kaplan's ideas to real YC batch trends moving from co-pilots to full workflow automation and references Dario Amodei's essay 'Machines of Loving Grace'. Kaplan elaborates collaboratively on intelligence breadth versus depth.24:36–30:06 · Guest disagreement 1/10 Applying Physics Heuristics, Large Matrix Limits, and Interpretability Diana draws upon physics terminology like renormalization and symmetry to ask how theoretical physics heuristics informed AI research. Kaplan clarifies that the main physics transfer was big-matrix approximations and asking simple, naive questions rather than complex formalisms.30:06–40:47 · Guest disagreement 1/10 Scaling Limits, Compute Efficiency, and Audience Questions on Horizon Lengths Diana asks contrarian questions about scaling breakdowns, compute limits (FP4/ternary representations), and invokes Jevons paradox. Audience members then challenge Kaplan on exponential task horizon jumps versus linear log scaling.2:21–6:09 · The partners pushing back 0/10 Core Mechanics of Pretraining and Reinforcement Learning in AI Solo presentation by Jared Kaplan explaining the core foundations of pre-training and reinforcement learning scaling laws. The host is not on stage.6:09–8:19 · The partners pushing back 0/10 Scaling Laws in Reinforcement Learning and Performance Metrics Kaplan continues his solo keynote, breaking down RL scaling laws and board game benchmarks like Hex and AlphaGo. There is no host involvement.8:19–11:09 · The partners pushing back 0/10 AI Capability Axes: Modalities and Task Time Horizon Lengths Monologue segment detailing the two capability axes: modality breadth and task horizon lengths doubling every seven months.11:09–15:33 · The partners pushing back 0/10 Key Frontiers for Human-Level AI and Practical Startup Recommendations Kaplan concludes his presentation with requirements for human-level AI and advice for startup founders. Host is not present.15:33–18:39 · The partners pushing back 0/10 Y Combinator Batch Application Announcement with Garry Tan Opens with a brief YC ad voiceover before Diana Hu joins the stage and asks Kaplan about Claude 4's release and its compounding capabilities over the next 12 months.18:39–24:36 · The partners pushing back 1/10 Human-AI Collaboration Dynamics, End-to-End Workflows, and Scientific Research Diana connects Kaplan's ideas to real YC batch trends moving from co-pilots to full workflow automation and references Dario Amodei's essay 'Machines of Loving Grace'. Kaplan elaborates collaboratively on intelligence breadth versus depth.24:36–30:06 · The partners pushing back 1/10 Applying Physics Heuristics, Large Matrix Limits, and Interpretability Diana draws upon physics terminology like renormalization and symmetry to ask how theoretical physics heuristics informed AI research. Kaplan clarifies that the main physics transfer was big-matrix approximations and asking simple, naive questions rather than complex formalisms.30:06–40:47 · The partners pushing back 2/10 Scaling Limits, Compute Efficiency, and Audience Questions on Horizon Lengths Diana asks contrarian questions about scaling breakdowns, compute limits (FP4/ternary representations), and invokes Jevons paradox. Audience members then challenge Kaplan on exponential task horizon jumps versus linear log scaling.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 0% · guest 100%0:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%15:00 · the partners 14.3% · guest 85.7%15:00 · the partners 14.3% · guest 85.7%18:00 · the partners 35% · guest 65%18:00 · the partners 35% · guest 65%21:00 · the partners 10.7% · guest 89.3%21:00 · the partners 10.7% · guest 89.3%24:00 · the partners 25.5% · guest 74.5%24:00 · the partners 25.5% · guest 74.5%27:00 · the partners 13.4% · guest 86.6%27:00 · the partners 13.4% · guest 86.6%30:00 · the partners 18.8% · guest 81.2%30:00 · the partners 18.8% · guest 81.2%33:00 · the partners 17.9% · guest 82.1%33:00 · the partners 17.9% · guest 82.1%36:00 · the partners 0% · guest 100%36:00 · the partners 0% · guest 100%39:00 · the partners 5.3% · guest 94.7%39:00 · the partners 5.3% · guest 94.7%
Sharpest disagreement ▶ 31:09 Kaplan rejecting the premise that scaling laws are broken

When asked what empirical proof would show scaling is breaking down, Kaplan pushes back against the premise, stating that whenever scaling seemed broken in the past five years, it was merely due to human implementation error in training.

Hardest push from the partners ▶ 29:50 Diana pushing for empirical failure conditions of scaling laws

Diana directly challenges the assumption of perpetual scaling by asking Kaplan what empirical sign would prove that the curve is finally flattening or failing.

Biggest teaching moment ▶ 26:45 Kaplan on questioning researchers about exponential convergence

Kaplan explains how he educated experienced AI researchers by challenging their imprecise claims of exponential convergence and replacing them with rigorous power law models.

The partners hold their own ▶ 33:15 Diana applying Jevons paradox to AI compute efficiency

Diana demonstrates domain fluency by immediately linking Kaplan's explanation of falling inference costs and soaring demand directly to Jevons paradox.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Core Mechanics of Pretraining and Reinforcement Learning in AI 0000 Solo presentation by Jared Kaplan explaining the core foundations of pre-training and reinforcement learning scaling laws. The host is not on stage.
Scaling Laws in Reinforcement Learning and Performance Metrics 0000 Kaplan continues his solo keynote, breaking down RL scaling laws and board game benchmarks like Hex and AlphaGo. There is no host involvement.
AI Capability Axes: Modalities and Task Time Horizon Lengths 0000 Monologue segment detailing the two capability axes: modality breadth and task horizon lengths doubling every seven months.
Key Frontiers for Human-Level AI and Practical Startup Recommendations 0000 Kaplan concludes his presentation with requirements for human-level AI and advice for startup founders. Host is not present.
Y Combinator Batch Application Announcement with Garry Tan 2100 Opens with a brief YC ad voiceover before Diana Hu joins the stage and asks Kaplan about Claude 4's release and its compounding capabilities over the next 12 months.
Human-AI Collaboration Dynamics, End-to-End Workflows, and Scientific Research 3201 Diana connects Kaplan's ideas to real YC batch trends moving from co-pilots to full workflow automation and references Dario Amodei's essay 'Machines of Loving Grace'. Kaplan elaborates collaboratively on intelligence breadth versus depth.
Applying Physics Heuristics, Large Matrix Limits, and Interpretability 4211 Diana draws upon physics terminology like renormalization and symmetry to ask how theoretical physics heuristics informed AI research. Kaplan clarifies that the main physics transfer was big-matrix approximations and asking simple, naive questions rather than complex formalisms.
Scaling Limits, Compute Efficiency, and Audience Questions on Horizon Lengths 4212 Diana asks contrarian questions about scaling breakdowns, compute limits (FP4/ternary representations), and invokes Jevons paradox. Audience members then challenge Kaplan on exponential task horizon jumps versus linear log scaling.

Statements from this episode (19)

Insight
Kaplan: Training frontier AI requires only next-word prediction and reinforcement learning
“So really all there is to training these models is learning to predict the next word and then doing reinforcement learning to learn to do useful tasks.”
Jared Kaplan Jul 29, 2025 ▶ 4:09
Insight
Kaplan: AI scaling trends are as precise as laws of physics
“We found that there's actually something very, very, very precise and surprising underlying AI training. This really blew us away that there are these nice trends that are as precise as anything that you see in physics or astronomy.”
Jared Kaplan Jul 29, 2025 ▶ 5:15
Assertion Supported
Kaplan: Scaling laws apply to reinforcement learning in AI training
“You can see scaling laws in the reinforcement learning phase of AI training.”
Jared Kaplan Jul 29, 2025 ▶ 6:15
Insight
Kaplan: Compute scaling drives AI progress more than researcher cleverness
“Basically you can Scale up the compute in both pre-training and RL and get better and better performance. And I think that's sort of the fundamental thing that is driving AI progress. It's not that AI researchers are really smart or they suddenly got smart. It…”
Jared Kaplan Jul 29, 2025 ▶ 7:54
Assertion Supported
Kaplan: METR found AI task duration doubles roughly every seven months
“And an organization meter studied this very systematically and found yet another scaling trend. They found that If you look at the length of tasks that AI models can do, it's doubling roughly every seven months.”
Jared Kaplan Jul 29, 2025 ▶ 9:41
Prediction Not checkable as stated
Kaplan: AI may execute multi-month tasks within the next few years
“And this kind of picture suggests that over the next few years, we may reach a point where AI models can do tasks that don't just take us minutes or hours, but days, weeks, months, years, et cetera.”
Jared Kaplan Jul 29, 2025 ▶ 10:23
Prediction Not checkable as stated
Kaplan: AI swarms will eventually replicate entire scientific communities
“Eventually we imagine AI models or millions of AI models perhaps working together will be able to do the work That whole human organizations can do. They'll be able to do the kind of work that the entire scientific community currently does.”
Jared Kaplan Jul 29, 2025 ▶ 10:37
Insight
Kaplan: Founders should build products that current AI cannot quite support
“One is I think it's really a good idea to build things that don't quite work yet. This is probably always a good idea. We always want to have ambition, but I think specifically AI models right now are getting better very, very quickly. And I think that's going…”
Jared Kaplan Jul 29, 2025 ▶ 13:48
Assertion Not checkable as stated
Kaplan: Claude 3.7 Sonnet took coding shortcuts just to pass tests
“With Claude III. VII's Sonnet it was already really exciting to use 3.7 for coding, but I think something that everyone noticed was that 3.7 was a little bit too eager. Sometimes it just really wanted to make your tests pass and it would do things that, that y…”
Jared Kaplan Jul 29, 2025 ▶ 16:19
Assertion Supported
Kaplan: Claude 4 will store and retrieve memories across context windows
“Claude IV can blow through its context window with a very complex task, but can also, ah, store memories as files or records, retrieve them in order to sort of keep doing work across many, many, many context windows.”
Jared Kaplan Jul 29, 2025 ▶ 17:19
Prediction Not checkable as stated
Kaplan: AI scaling curves point smoothly toward human-level AGI
“I think that scaling Really suggests a kind of smooth curve towards what I expect is kind of human level AI or AGI.”
Jared Kaplan Jul 29, 2025 ▶ 17:31
Insight
Kaplan: AI's generative capability and judgment are much closer than in humans
“I think one of the sort of basic features of AI that's different about the shape of AI intelligence compared to human intelligence is that there are a lot of things that I can't do, but I can at least judge whether they were done correctly. I think for AI, the…”
Jared Kaplan Jul 29, 2025 ▶ 19:09
Insight
Kaplan: AI holds a major capability overhang in multidisciplinary knowledge synthesis
“I think that we're making a lot of progress on making AI better at deeper tasks like hard coding problems, hard math problems. But I suspect that there's a particular overhang in areas where putting together knowledge that maybe no one human expert would have,…”
Jared Kaplan Jul 29, 2025 ▶ 23:20
Insight
Kaplan: The holy grail of scaling laws is improving the slope
“I think with scaling laws, the holy grail is finding a better slope to the scaling law, because that means that as you put in more compute, you're going to get a bigger and bigger advantage over other AI developers.”
Jared Kaplan Jul 29, 2025 ▶ 27:13
Insight
Kaplan: Interpretability is like neuroscience but with complete observability
“I would say that interpretability is a lot more like biology. It's a lot more like neuroscience. So I think those are kind of the tools. There is some more, more, more mathematics there, but I think it's more like trying to understand the features of the brain…”
Jared Kaplan Jul 29, 2025 ▶ 29:16
Insight
Kaplan: Broken scaling laws usually signal flawed training setups, not fundamental limits
“If scaling laws are failing, it's because we've screwed up AI training in some way. Maybe we got, ah, we got the architecture of the neural network wrong, or there's some bottleneck in training that we don't see, or there's some problem with Precision and the …”
Jared Kaplan Jul 29, 2025 ▶ 30:29
Prediction Held up
Kaplan: AI training and inference will see 3x-10x yearly efficiency gains
“I think that over time, as AI becomes more and more widespread, I think that we're going to really drive down the cost of inference and training dramatically from where we are right now. Sort of, three X to 10 X gains algorithmically, and in sort of scaling up…”
Jared Kaplan Jul 29, 2025 ▶ 32:02
Prediction Not checkable as stated
Kaplan: Most AI value will come from end-to-end frontier models
“I think that you can do a lot of very simple bite-sized tasks, but I think it's just much more convenient to be able to use an AI model that can do a very complex task end-to-end, rather than requiring us as humans to sort of orchestrate a much dumber model to…”
Jared Kaplan Jul 29, 2025 ▶ 34:13
Insight
Kaplan: Modest self-correction capabilities can double autonomous AI task horizons
“A lot of what determines the horizon length of what models can accomplish is their ability to notice that they're doing something wrong and correct it. And I think that's not sort of like a lot of bits of information. It doesn't necessarily require a huge chan…”
Jared Kaplan Jul 29, 2025 ▶ 36:24
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.