Feb 27, 2026 · 1h 5m · latent-space

Measuring Exponential Trends Rising (in AI) — Joel Becker, METR

Joel Becker · 36m spoken Shawn Wang · 14m spoken Alessio Fanelli · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space podcast, METR's Joel Becker discusses empirical AI safety evaluations, the methodology behind the Model Time Horizon benchmark, the dynamics of automated AI R&D, and his personal strategies in prediction markets.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 39.2% of the talking time here. How this is scored →

The hosts as informed peer 5.6 Guest teaching 5.0 Guest disagreement 1.9 The hosts pushing back 3.1
05100:0015:0030:0045:001:00:001:46–4:03 · The hosts as informed peer 4/10 Distinguishing Model Evaluation from Threat Research Swyx and Alessio ask clarifying questions regarding METR's balance between model evaluation and threat research, as well as how threat models have evolved. Joel explains the shift from autonomous replication toward automated R&D acceleration.4:03–8:40 · The hosts as informed peer 5/10 The Origin and Methodology of the Time Horizon Benchmark Alessio probes the origins of the time horizon metric and how METR selects tasks without introducing arbitrary bias. Joel details the internal history, straight trendline emergence, and task constraints like low-context solvability and autograding.8:41–13:27 · The hosts as informed peer 6/10 Task Distributions: From SWAs to RE-Bench The hosts and guest discuss the differences between SWA, HCAST, and RE-Bench tasks, and why task difficulty is measured in human hours rather than agent wall-clock execution time. Joel and Swyx agree on dismissing sensational, unscientific agentic claims.13:27–16:43 · The hosts as informed peer 6/10 Opus 4.5 and the Trend Line Trajectory Swyx suggests Opus 4.5 broke METR's trendline and references Joel's forecasting background. Joel directly rejects both premises, clarifying he is not formally a superforecaster and arguing model progress remains remarkably continuous across compute scales.16:43–24:36 · The hosts as informed peer 6/10 Evaluating Developer Productivity and RCT Methodologies Alessio and Swyx discuss developer productivity RCTs and empirical workflows. Joel educates on methodological hazards in new studies, including developer selection bias, concurrent task multitasking, and marginal value distortion.24:36–30:13 · The hosts as informed peer 6/10 Capability Explosions, Automated R&D, and Physical Loops Swyx and Alessio question whether capability explosions will follow continuous curves or sudden phase transitions like boiling water. Joel analyzes the conditions for recursive self-improvement and notes that hardware bottlenecks prevent immediate software explosion.30:13–34:15 · The hosts as informed peer 6/10 The Long Tail of AI R&D Automation and Metric Nuance Swyx critiques the reduction of AI capability tracking to single summary numbers and advocates for multidimensional capability lists. Joel counters by highlighting the vast real-world physical tail required to automate R&D and challenges Swyx to enumerate an exhaustive list.34:15–42:10 · The hosts as informed peer 7/10 Compute Bottlenecks and Algorithmic Progress Alessio and Joel analyze how compute slowdowns could directly halve algorithmic progress. Swyx demonstrates strong industry domain knowledge, citing cluster timelines, unobserved hyperscaler capex, and SemiAnalysis insights.42:10–49:47 · The hosts as informed peer 6/10 Prediction Markets, Manifold Alpha, and Market Ethics Alessio and Swyx ask Joel about ranking #1 on Manifold Markets. Joel explains his arbitrage maneuver on charity matching markets and questions the broader societal value of prediction markets given gambling dynamics, while Swyx defends price discovery.49:48–56:05 · The hosts as informed peer 6/10 Frontier Evals: AI Village and Real-World Transcripts Joel outlines cutting-edge evaluation approaches such as AI Village and raw transcript analysis. Swyx links these ideas to Noam Brown's multi-agent cooperation research and DeepMind's open-endedness teams.56:05–59:43 · The hosts as informed peer 7/10 Scaffolding, Harness Optimization, and Skill Depreciation Alessio brings up Terminal Bench harness variations and argues in favor of aggressive task overfitting for immediate developer utility. Swyx and Joel debate whether custom scaffolding creates lasting value or gets washed away by base model advancements.59:44–1:03:05 · The hosts as informed peer 5/10 METR's Roadmap and Hiring Priorities Swyx asks about METR's roadmap toward 2026/2030 and what specific qualities METR seeks when rejecting average candidates. Joel outlines requirements around rigorous data intuition, transparent scientific writing, and scrappy execution.1:03:06–1:05:09 · The hosts as informed peer 3/10 Live Band Karaoke and Parting Reflections The conversation closes with lighthearted personal discussion about live band karaoke, acapella music history, and the enduring human element of live performance compared to synthetic AI generation.1:46–4:03 · Guest teaching 4/10 Distinguishing Model Evaluation from Threat Research Swyx and Alessio ask clarifying questions regarding METR's balance between model evaluation and threat research, as well as how threat models have evolved. Joel explains the shift from autonomous replication toward automated R&D acceleration.4:03–8:40 · Guest teaching 5/10 The Origin and Methodology of the Time Horizon Benchmark Alessio probes the origins of the time horizon metric and how METR selects tasks without introducing arbitrary bias. Joel details the internal history, straight trendline emergence, and task constraints like low-context solvability and autograding.8:41–13:27 · Guest teaching 5/10 Task Distributions: From SWAs to RE-Bench The hosts and guest discuss the differences between SWA, HCAST, and RE-Bench tasks, and why task difficulty is measured in human hours rather than agent wall-clock execution time. Joel and Swyx agree on dismissing sensational, unscientific agentic claims.13:27–16:43 · Guest teaching 6/10 Opus 4.5 and the Trend Line Trajectory Swyx suggests Opus 4.5 broke METR's trendline and references Joel's forecasting background. Joel directly rejects both premises, clarifying he is not formally a superforecaster and arguing model progress remains remarkably continuous across compute scales.16:43–24:36 · Guest teaching 6/10 Evaluating Developer Productivity and RCT Methodologies Alessio and Swyx discuss developer productivity RCTs and empirical workflows. Joel educates on methodological hazards in new studies, including developer selection bias, concurrent task multitasking, and marginal value distortion.24:36–30:13 · Guest teaching 5/10 Capability Explosions, Automated R&D, and Physical Loops Swyx and Alessio question whether capability explosions will follow continuous curves or sudden phase transitions like boiling water. Joel analyzes the conditions for recursive self-improvement and notes that hardware bottlenecks prevent immediate software explosion.30:13–34:15 · Guest teaching 6/10 The Long Tail of AI R&D Automation and Metric Nuance Swyx critiques the reduction of AI capability tracking to single summary numbers and advocates for multidimensional capability lists. Joel counters by highlighting the vast real-world physical tail required to automate R&D and challenges Swyx to enumerate an exhaustive list.34:15–42:10 · Guest teaching 5/10 Compute Bottlenecks and Algorithmic Progress Alessio and Joel analyze how compute slowdowns could directly halve algorithmic progress. Swyx demonstrates strong industry domain knowledge, citing cluster timelines, unobserved hyperscaler capex, and SemiAnalysis insights.42:10–49:47 · Guest teaching 6/10 Prediction Markets, Manifold Alpha, and Market Ethics Alessio and Swyx ask Joel about ranking #1 on Manifold Markets. Joel explains his arbitrage maneuver on charity matching markets and questions the broader societal value of prediction markets given gambling dynamics, while Swyx defends price discovery.49:48–56:05 · Guest teaching 6/10 Frontier Evals: AI Village and Real-World Transcripts Joel outlines cutting-edge evaluation approaches such as AI Village and raw transcript analysis. Swyx links these ideas to Noam Brown's multi-agent cooperation research and DeepMind's open-endedness teams.56:05–59:43 · Guest teaching 5/10 Scaffolding, Harness Optimization, and Skill Depreciation Alessio brings up Terminal Bench harness variations and argues in favor of aggressive task overfitting for immediate developer utility. Swyx and Joel debate whether custom scaffolding creates lasting value or gets washed away by base model advancements.59:44–1:03:05 · Guest teaching 4/10 METR's Roadmap and Hiring Priorities Swyx asks about METR's roadmap toward 2026/2030 and what specific qualities METR seeks when rejecting average candidates. Joel outlines requirements around rigorous data intuition, transparent scientific writing, and scrappy execution.1:03:06–1:05:09 · Guest teaching 2/10 Live Band Karaoke and Parting Reflections The conversation closes with lighthearted personal discussion about live band karaoke, acapella music history, and the enduring human element of live performance compared to synthetic AI generation.1:46–4:03 · Guest disagreement 1/10 Distinguishing Model Evaluation from Threat Research Swyx and Alessio ask clarifying questions regarding METR's balance between model evaluation and threat research, as well as how threat models have evolved. Joel explains the shift from autonomous replication toward automated R&D acceleration.4:03–8:40 · Guest disagreement 1/10 The Origin and Methodology of the Time Horizon Benchmark Alessio probes the origins of the time horizon metric and how METR selects tasks without introducing arbitrary bias. Joel details the internal history, straight trendline emergence, and task constraints like low-context solvability and autograding.8:41–13:27 · Guest disagreement 2/10 Task Distributions: From SWAs to RE-Bench The hosts and guest discuss the differences between SWA, HCAST, and RE-Bench tasks, and why task difficulty is measured in human hours rather than agent wall-clock execution time. Joel and Swyx agree on dismissing sensational, unscientific agentic claims.13:27–16:43 · Guest disagreement 4/10 Opus 4.5 and the Trend Line Trajectory Swyx suggests Opus 4.5 broke METR's trendline and references Joel's forecasting background. Joel directly rejects both premises, clarifying he is not formally a superforecaster and arguing model progress remains remarkably continuous across compute scales.16:43–24:36 · Guest disagreement 2/10 Evaluating Developer Productivity and RCT Methodologies Alessio and Swyx discuss developer productivity RCTs and empirical workflows. Joel educates on methodological hazards in new studies, including developer selection bias, concurrent task multitasking, and marginal value distortion.24:36–30:13 · Guest disagreement 2/10 Capability Explosions, Automated R&D, and Physical Loops Swyx and Alessio question whether capability explosions will follow continuous curves or sudden phase transitions like boiling water. Joel analyzes the conditions for recursive self-improvement and notes that hardware bottlenecks prevent immediate software explosion.30:13–34:15 · Guest disagreement 3/10 The Long Tail of AI R&D Automation and Metric Nuance Swyx critiques the reduction of AI capability tracking to single summary numbers and advocates for multidimensional capability lists. Joel counters by highlighting the vast real-world physical tail required to automate R&D and challenges Swyx to enumerate an exhaustive list.34:15–42:10 · Guest disagreement 1/10 Compute Bottlenecks and Algorithmic Progress Alessio and Joel analyze how compute slowdowns could directly halve algorithmic progress. Swyx demonstrates strong industry domain knowledge, citing cluster timelines, unobserved hyperscaler capex, and SemiAnalysis insights.42:10–49:47 · Guest disagreement 3/10 Prediction Markets, Manifold Alpha, and Market Ethics Alessio and Swyx ask Joel about ranking #1 on Manifold Markets. Joel explains his arbitrage maneuver on charity matching markets and questions the broader societal value of prediction markets given gambling dynamics, while Swyx defends price discovery.49:48–56:05 · Guest disagreement 1/10 Frontier Evals: AI Village and Real-World Transcripts Joel outlines cutting-edge evaluation approaches such as AI Village and raw transcript analysis. Swyx links these ideas to Noam Brown's multi-agent cooperation research and DeepMind's open-endedness teams.56:05–59:43 · Guest disagreement 2/10 Scaffolding, Harness Optimization, and Skill Depreciation Alessio brings up Terminal Bench harness variations and argues in favor of aggressive task overfitting for immediate developer utility. Swyx and Joel debate whether custom scaffolding creates lasting value or gets washed away by base model advancements.59:44–1:03:05 · Guest disagreement 1/10 METR's Roadmap and Hiring Priorities Swyx asks about METR's roadmap toward 2026/2030 and what specific qualities METR seeks when rejecting average candidates. Joel outlines requirements around rigorous data intuition, transparent scientific writing, and scrappy execution.1:03:06–1:05:09 · Guest disagreement 1/10 Live Band Karaoke and Parting Reflections The conversation closes with lighthearted personal discussion about live band karaoke, acapella music history, and the enduring human element of live performance compared to synthetic AI generation.1:46–4:03 · The hosts pushing back 2/10 Distinguishing Model Evaluation from Threat Research Swyx and Alessio ask clarifying questions regarding METR's balance between model evaluation and threat research, as well as how threat models have evolved. Joel explains the shift from autonomous replication toward automated R&D acceleration.4:03–8:40 · The hosts pushing back 2/10 The Origin and Methodology of the Time Horizon Benchmark Alessio probes the origins of the time horizon metric and how METR selects tasks without introducing arbitrary bias. Joel details the internal history, straight trendline emergence, and task constraints like low-context solvability and autograding.8:41–13:27 · The hosts pushing back 3/10 Task Distributions: From SWAs to RE-Bench The hosts and guest discuss the differences between SWA, HCAST, and RE-Bench tasks, and why task difficulty is measured in human hours rather than agent wall-clock execution time. Joel and Swyx agree on dismissing sensational, unscientific agentic claims.13:27–16:43 · The hosts pushing back 4/10 Opus 4.5 and the Trend Line Trajectory Swyx suggests Opus 4.5 broke METR's trendline and references Joel's forecasting background. Joel directly rejects both premises, clarifying he is not formally a superforecaster and arguing model progress remains remarkably continuous across compute scales.16:43–24:36 · The hosts pushing back 4/10 Evaluating Developer Productivity and RCT Methodologies Alessio and Swyx discuss developer productivity RCTs and empirical workflows. Joel educates on methodological hazards in new studies, including developer selection bias, concurrent task multitasking, and marginal value distortion.24:36–30:13 · The hosts pushing back 4/10 Capability Explosions, Automated R&D, and Physical Loops Swyx and Alessio question whether capability explosions will follow continuous curves or sudden phase transitions like boiling water. Joel analyzes the conditions for recursive self-improvement and notes that hardware bottlenecks prevent immediate software explosion.30:13–34:15 · The hosts pushing back 4/10 The Long Tail of AI R&D Automation and Metric Nuance Swyx critiques the reduction of AI capability tracking to single summary numbers and advocates for multidimensional capability lists. Joel counters by highlighting the vast real-world physical tail required to automate R&D and challenges Swyx to enumerate an exhaustive list.34:15–42:10 · The hosts pushing back 3/10 Compute Bottlenecks and Algorithmic Progress Alessio and Joel analyze how compute slowdowns could directly halve algorithmic progress. Swyx demonstrates strong industry domain knowledge, citing cluster timelines, unobserved hyperscaler capex, and SemiAnalysis insights.42:10–49:47 · The hosts pushing back 4/10 Prediction Markets, Manifold Alpha, and Market Ethics Alessio and Swyx ask Joel about ranking #1 on Manifold Markets. Joel explains his arbitrage maneuver on charity matching markets and questions the broader societal value of prediction markets given gambling dynamics, while Swyx defends price discovery.49:48–56:05 · The hosts pushing back 3/10 Frontier Evals: AI Village and Real-World Transcripts Joel outlines cutting-edge evaluation approaches such as AI Village and raw transcript analysis. Swyx links these ideas to Noam Brown's multi-agent cooperation research and DeepMind's open-endedness teams.56:05–59:43 · The hosts pushing back 5/10 Scaffolding, Harness Optimization, and Skill Depreciation Alessio brings up Terminal Bench harness variations and argues in favor of aggressive task overfitting for immediate developer utility. Swyx and Joel debate whether custom scaffolding creates lasting value or gets washed away by base model advancements.59:44–1:03:05 · The hosts pushing back 2/10 METR's Roadmap and Hiring Priorities Swyx asks about METR's roadmap toward 2026/2030 and what specific qualities METR seeks when rejecting average candidates. Joel outlines requirements around rigorous data intuition, transparent scientific writing, and scrappy execution.1:03:06–1:05:09 · The hosts pushing back 1/10 Live Band Karaoke and Parting Reflections The conversation closes with lighthearted personal discussion about live band karaoke, acapella music history, and the enduring human element of live performance compared to synthetic AI generation.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 22.1% · guest 77.9%0:00 · the hosts 22.1% · guest 77.9%3:00 · the hosts 26.3% · guest 73.7%3:00 · the hosts 26.3% · guest 73.7%6:00 · the hosts 13.4% · guest 86.6%6:00 · the hosts 13.4% · guest 86.6%9:00 · the hosts 23% · guest 77%9:00 · the hosts 23% · guest 77%12:00 · the hosts 41.4% · guest 58.6%12:00 · the hosts 41.4% · guest 58.6%15:00 · the hosts 31.1% · guest 68.9%15:00 · the hosts 31.1% · guest 68.9%18:00 · the hosts 37.4% · guest 62.6%18:00 · the hosts 37.4% · guest 62.6%21:00 · the hosts 79.4% · guest 20.6%21:00 · the hosts 79.4% · guest 20.6%24:00 · the hosts 61.6% · guest 38.4%24:00 · the hosts 61.6% · guest 38.4%27:00 · the hosts 41.9% · guest 58.1%27:00 · the hosts 41.9% · guest 58.1%30:00 · the hosts 47.6% · guest 52.4%30:00 · the hosts 47.6% · guest 52.4%33:00 · the hosts 37.4% · guest 62.6%33:00 · the hosts 37.4% · guest 62.6%36:00 · the hosts 22.8% · guest 77.2%36:00 · the hosts 22.8% · guest 77.2%39:00 · the hosts 89.4% · guest 10.6%39:00 · the hosts 89.4% · guest 10.6%42:00 · the hosts 33.7% · guest 66.3%42:00 · the hosts 33.7% · guest 66.3%45:00 · the hosts 31.4% · guest 68.6%45:00 · the hosts 31.4% · guest 68.6%48:00 · the hosts 53.5% · guest 46.5%48:00 · the hosts 53.5% · guest 46.5%51:00 · the hosts 29.6% · guest 70.4%51:00 · the hosts 29.6% · guest 70.4%54:00 · the hosts 20.5% · guest 79.5%54:00 · the hosts 20.5% · guest 79.5%57:00 · the hosts 46.4% · guest 53.6%57:00 · the hosts 46.4% · guest 53.6%1:00:00 · the hosts 17.4% · guest 82.6%1:00:00 · the hosts 17.4% · guest 82.6%1:03:00 · the hosts 67.3% · guest 32.7%1:03:00 · the hosts 67.3% · guest 32.7%
Sharpest disagreement ▶ 13:59 Direct rejection of forecasting credentials and discontinuous progress

Joel explicitly challenges Swyx's framing by stating he wants to attack two claims, denying he was ever a formal superforecaster and disputing that Opus 4.5 represented an unprecedented break in the trendline.

Hardest push from the hosts ▶ 58:37 Alessio and Swyx defend pragmatic scaffolding over pure wait-and-see

Alessio and Swyx challenge the fatalistic view that all scaffolding is useless because future models will replace it, insisting developers must build custom harness optimizations for immediate production needs.

Biggest teaching moment ▶ 19:17 Joel explains systemic biases in developer productivity metrics

Joel methodically breaks down why anecdotal and self-reported 10x speedups are economically misleading, explaining how concurrency, task selection bias, and low marginal value distort uplift calculations.

The host holds their own ▶ 39:29 Swyx details cluster timelines and unobserved hardware capex

Swyx demonstrates deep hardware and industry supply-chain knowledge, noting how GPU deployment schedules at major labs mathematically dictate model release windows beyond public financial disclosures.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Distinguishing Model Evaluation from Threat Research 4412 Swyx and Alessio ask clarifying questions regarding METR's balance between model evaluation and threat research, as well as how threat models have evolved. Joel explains the shift from autonomous replication toward automated R&D acceleration.
The Origin and Methodology of the Time Horizon Benchmark 5512 Alessio probes the origins of the time horizon metric and how METR selects tasks without introducing arbitrary bias. Joel details the internal history, straight trendline emergence, and task constraints like low-context solvability and autograding.
Task Distributions: From SWAs to RE-Bench 6523 The hosts and guest discuss the differences between SWA, HCAST, and RE-Bench tasks, and why task difficulty is measured in human hours rather than agent wall-clock execution time. Joel and Swyx agree on dismissing sensational, unscientific agentic claims.
Opus 4.5 and the Trend Line Trajectory 6644 Swyx suggests Opus 4.5 broke METR's trendline and references Joel's forecasting background. Joel directly rejects both premises, clarifying he is not formally a superforecaster and arguing model progress remains remarkably continuous across compute scales.
Evaluating Developer Productivity and RCT Methodologies 6624 Alessio and Swyx discuss developer productivity RCTs and empirical workflows. Joel educates on methodological hazards in new studies, including developer selection bias, concurrent task multitasking, and marginal value distortion.
Capability Explosions, Automated R&D, and Physical Loops 6524 Swyx and Alessio question whether capability explosions will follow continuous curves or sudden phase transitions like boiling water. Joel analyzes the conditions for recursive self-improvement and notes that hardware bottlenecks prevent immediate software explosion.
The Long Tail of AI R&D Automation and Metric Nuance 6634 Swyx critiques the reduction of AI capability tracking to single summary numbers and advocates for multidimensional capability lists. Joel counters by highlighting the vast real-world physical tail required to automate R&D and challenges Swyx to enumerate an exhaustive list.
Compute Bottlenecks and Algorithmic Progress 7513 Alessio and Joel analyze how compute slowdowns could directly halve algorithmic progress. Swyx demonstrates strong industry domain knowledge, citing cluster timelines, unobserved hyperscaler capex, and SemiAnalysis insights.
Prediction Markets, Manifold Alpha, and Market Ethics 6634 Alessio and Swyx ask Joel about ranking #1 on Manifold Markets. Joel explains his arbitrage maneuver on charity matching markets and questions the broader societal value of prediction markets given gambling dynamics, while Swyx defends price discovery.
Frontier Evals: AI Village and Real-World Transcripts 6613 Joel outlines cutting-edge evaluation approaches such as AI Village and raw transcript analysis. Swyx links these ideas to Noam Brown's multi-agent cooperation research and DeepMind's open-endedness teams.
Scaffolding, Harness Optimization, and Skill Depreciation 7525 Alessio brings up Terminal Bench harness variations and argues in favor of aggressive task overfitting for immediate developer utility. Swyx and Joel debate whether custom scaffolding creates lasting value or gets washed away by base model advancements.
METR's Roadmap and Hiring Priorities 5412 Swyx asks about METR's roadmap toward 2026/2030 and what specific qualities METR seeks when rejecting average candidates. Joel outlines requirements around rigorous data intuition, transparent scientific writing, and scrappy execution.
Live Band Karaoke and Parting Reflections 3211 The conversation closes with lighthearted personal discussion about live band karaoke, acapella music history, and the enduring human element of live performance compared to synthetic AI generation.

Statements from this episode (21)

Assertion Supported
Becker: Current Frontier Models Cannot Cause Catastrophic Harm
“We find, we think it's not capable enough, you know, on the basis of some of this capabilities evidence that you've alluded to commit these catastrophic harms.”
Joel Becker Feb 27, 2026 ▶ 2:34
Disclosure
METR Shifts Focus from Autonomous Replication to AI R&D Acceleration
“So, something like the autonomous replication threat model, that is being able to set yourself up and control resources, something like that, has been deprioritized relative to AR and D acceleration. That is, you know, the possibility there could be some capab…”
Joel Becker Feb 27, 2026 ▶ 3:38
Opinion
Becker: AI Models Lag on Time Horizons for Vision Tasks
“Tasks that are requiring of vision capabilities, they're probably to take one example, they're probably much less capable today as measured by time horizon, as for these tasks that are typically not requiring vision, vision capabilities that we give them.”
Joel Becker Feb 27, 2026 ▶ 6:21
Insight
Becker: Low-Context Benchmark Criteria Exclude Real-World Situational Work
“Could a low-context human who was sufficiently skilled at sort of The general skills, but maybe, maybe not the particulars in the background would they be able to achieve success on, on this task? And I think that, that rules out a lot of real work because, yo…”
Joel Becker Feb 27, 2026 ▶ 7:31
Assertion Partly supported
Becker: Claude Opus 4 Solves Atomic Software Tasks 100% Reliably
“Opus-IV. I'm sure can do that task a hundred percent of the time.”
Joel Becker Feb 27, 2026 ▶ 10:01
Assertion Supported
Becker: METR's Hardest Benchmarks Require 20 to 30 Hours of Autonomy
“Then we go up to HCOS tasks, which span from, you know, only a little harder than those small tasks, all the way up to, you know, something like 20:30 hours, which are requiring of more autonomy, more sort of more sort of sequential actions.”
Joel Becker Feb 27, 2026 ▶ 10:05
Insight
Becker: Time Horizon Metric Measures AI Task Difficulty in Human Time
“You know, instead we're just plotting what's the difficulty of tasks they can do over time, and that difficulty is measured in human time.”
Joel Becker Feb 27, 2026 ▶ 11:59
Opinion
Swyx: Public AI Agent Performance Claims Are Unscientific and Marketing-Driven
“The state of people making claims on agent performance is very unscientific and much more anecdotal and sometimes influenced by marketing desires.”
Shawn Wang Feb 27, 2026 ▶ 13:02
Insight
Becker: AI Progress Remains Highly Continuous Across Compute Scales
“You know, in some ways, I think the story of Time Horizon is that progress has been remarkably continuous over, over so many years, so many orders of magnitude of compute and effective compute.”
Joel Becker Feb 27, 2026 ▶ 14:55
Disclosure
Becker: METR Is Rerunning Its Developer Productivity Randomized Controlled Trial
“We have been redoing it in the background.”
Joel Becker Feb 27, 2026 ▶ 17:02
Opinion
Becker: Overly Bullish AI Developer Speedup Estimates Are Inflated
“I do think that very bullish estimates of speed up today are, you know, to some extent inflated by what we document in that original paper, that people's expectations of speed up tend to be too optimistic, it seems. They also tend to be inflated, I think, by n…”
Joel Becker Feb 27, 2026 ▶ 20:08
Prediction Not checkable as stated
Becker: Operational Long Tail Will Delay Full AI R&D Automation
“There's this, Very long tail of things potentially involved in in R&D that would perhaps need to be fully automated in order to lead to capabilities explosion. I expect we're measuring, you know, in some ways only, only a small proportion of, only a small prop…”
Joel Becker Feb 27, 2026 ▶ 31:52
Insight
Becker: Algorithmic Progress in AI Is Strictly a Function of Compute
“The suggestion in this paper is that if you think that algorithmic progress, you know, that, that is coming up with the transformer, coming up with RLHF, you know, MOEs, all of this stuff, better learning rate schedules is, is is itself a function of compute b…”
Joel Becker Feb 27, 2026 ▶ 35:18
Prediction Not checkable as stated
Becker: Halving AI Compute Growth Halves Algorithmic Progress and Milestones
“And both of them both of those components half when compute halves sort of trivially, because compute is halving, and algorithmic progress halves because compute is this important input, and compute halves, then you might expect time horizon growth to half. An…”
Joel Becker Feb 27, 2026 ▶ 36:12
Disclosure
Becker: Donated $5,000 Won on Manifold Markets to Charity
“And then I ended up donating, I can't remember exactly how, how much it was, not, not so much, something like 5000 dollars .”
Joel Becker Feb 27, 2026 ▶ 45:15
Opinion
Becker: Prediction Markets' Social Value May Not Justify Gambling Harms
“I think gambling like behaviors are socially costly and the value of higher quality information is is, is, is real, but, you know, is it worth that disbenefit of people trading away their money, it's not, you know, it's not so clear to me.”
Joel Becker Feb 27, 2026 ▶ 47:05
Disclosure
Swyx: Latent Space Holds Embargoed AI News Without Trading
“We do work with people on embargo and we don't trade.”
Shawn Wang Feb 27, 2026 ▶ 49:39
Opinion
Becker: AI Coding Fails at Merge Readiness Despite High SWE-bench Scores
“Maybe one that I'll call out there is this difference between whether models pass unit tests, whether they succeed by, you know, SWE bench-like scoring kind of meter-like scoring, benchmark-style scoring, versus whether their solution would be merged into main…”
Joel Becker Feb 27, 2026 ▶ 52:26
Insight
Becker: AI Scaffolding Value Does Not Persist Across Model Generations
“Within model generation, it's valuable, and across model generations, it's not so valuable.”
Joel Becker Feb 27, 2026 ▶ 59:09
Disclosure
Becker: Stopped Investing in Personal Software Engineering Skills Due to AI
“Intentionally not investing in engineering skills, because the areas are getting so good, maybe that's the wrong decision.”
Joel Becker Feb 27, 2026 ▶ 59:13
Assertion Supported
Becker: METR Uses Black-Box Methods Over Interpretability for AI Monitoring
“Usually this is black box, not, not white box in, in, in my understanding in, in current work. So, so not using interpretability, but you can imagine in principle doing, doing, doing something more white box.”
Joel Becker Feb 27, 2026 ▶ 1:00:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.