Mar 7, 2025 · 28m · latent-space

Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]

Misha Laskin · 20m spoken Shawn Wang · 3m spoken Alessio Fanelli · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Reflection AI co-founder and CEO Misha Laskin discusses the company's emergence from stealth to build fully autonomous coding agents powered by the convergence of large language models and reinforcement learning. He explains why solving autonomous software engineering is the direct catalyst for general superintelligence and outlines Reflection AI's execution-coupled API and enterprise deployment strategy.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.5% of the talking time here. How this is scored →

The hosts as informed peer 4.8 Guest teaching 4.3 Guest disagreement 1.7 The hosts pushing back 2.2
05100:0010:0020:000:44–4:06 · The hosts as informed peer 3/10 The Core Thesis: RL, LLMs, and Autonomous Coding Swyx opens by inviting Misha to detail Reflection AI's emergence from stealth. Misha delivers an extended monologue outlining the team's thesis on combining RL with LLMs to reach superintelligence through coding.4:07–7:56 · The hosts as informed peer 5/10 The RL Pendulum Shift and Foundation Model Baselines Alessio contextualizes historical RL shifts from OpenAI Dota to language, prompting Misha on computer use versus code. Misha educates the hosts on why mouse-based web browsing lacks internet priors compared to code ergonomics.7:57–10:31 · The hosts as informed peer 6/10 Defining Superintelligence and AI-Native Programmatic Interfaces Swyx pushes back on Misha's leap from deterministic coding agents to broad superintelligence. Misha reframes superintelligence around creative discovery like Move 37 and programmatic API execution.10:33–13:22 · The hosts as informed peer 6/10 Data Mixture: Balancing SFT with Reinforcement Learning Swyx brings up Misha's past commentary regarding ChatGPT imitation learning and references past guest Bret Taylor on human versus AI programming languages. Misha explains the necessity of bootstrapping RL with sensible SFT mixtures.13:23–17:01 · The hosts as informed peer 5/10 Anticipating 'Move 37' Breakthroughs in Code Generation Alessio questions what Move 37 means in practical software engineering and critiques hype around Deep Research. Misha points to DeepSeek's custom attention kernels and warns against closed labs hoarding powerful internal models.17:01–20:22 · The hosts as informed peer 5/10 Product Architecture: Autonomous Backlog Resolution API Swyx compares Reflection's upcoming form factor to Poolside's VS Code extension. Misha differentiates between developer-driven copilots and autonomous execution APIs handling enterprise technical debt and backlog triage.20:23–22:52 · The hosts as informed peer 6/10 Long Context Understanding vs. Agentic Code Localization Swyx raises Magic.dev's 100M context windows to question state-of-the-art code indexing. Misha breaks down the architectural tradeoffs between naive long context attention mechanisms and agentic code localization.22:52–26:00 · The hosts as informed peer 5/10 Evaluating Coding Agents Through Real-World Co-Development Swyx inquires about SWE-bench benchmarks versus competitors like Devin and Magic. Misha argues that autonomous coding cannot be evaluated in a benchmark vacuum and requires direct customer co-development.26:01–27:57 · The hosts as informed peer 2/10 Hiring Across Research, Product, and Company Culture Alessio wraps up by asking about open roles and cultural expectations. Misha outlines their emphasis on high agency, technical craftsmanship, and kindness.0:44–4:06 · Guest teaching 4/10 The Core Thesis: RL, LLMs, and Autonomous Coding Swyx opens by inviting Misha to detail Reflection AI's emergence from stealth. Misha delivers an extended monologue outlining the team's thesis on combining RL with LLMs to reach superintelligence through coding.4:07–7:56 · Guest teaching 6/10 The RL Pendulum Shift and Foundation Model Baselines Alessio contextualizes historical RL shifts from OpenAI Dota to language, prompting Misha on computer use versus code. Misha educates the hosts on why mouse-based web browsing lacks internet priors compared to code ergonomics.7:57–10:31 · Guest teaching 5/10 Defining Superintelligence and AI-Native Programmatic Interfaces Swyx pushes back on Misha's leap from deterministic coding agents to broad superintelligence. Misha reframes superintelligence around creative discovery like Move 37 and programmatic API execution.10:33–13:22 · Guest teaching 4/10 Data Mixture: Balancing SFT with Reinforcement Learning Swyx brings up Misha's past commentary regarding ChatGPT imitation learning and references past guest Bret Taylor on human versus AI programming languages. Misha explains the necessity of bootstrapping RL with sensible SFT mixtures.13:23–17:01 · Guest teaching 4/10 Anticipating 'Move 37' Breakthroughs in Code Generation Alessio questions what Move 37 means in practical software engineering and critiques hype around Deep Research. Misha points to DeepSeek's custom attention kernels and warns against closed labs hoarding powerful internal models.17:01–20:22 · Guest teaching 4/10 Product Architecture: Autonomous Backlog Resolution API Swyx compares Reflection's upcoming form factor to Poolside's VS Code extension. Misha differentiates between developer-driven copilots and autonomous execution APIs handling enterprise technical debt and backlog triage.20:23–22:52 · Guest teaching 5/10 Long Context Understanding vs. Agentic Code Localization Swyx raises Magic.dev's 100M context windows to question state-of-the-art code indexing. Misha breaks down the architectural tradeoffs between naive long context attention mechanisms and agentic code localization.22:52–26:00 · Guest teaching 5/10 Evaluating Coding Agents Through Real-World Co-Development Swyx inquires about SWE-bench benchmarks versus competitors like Devin and Magic. Misha argues that autonomous coding cannot be evaluated in a benchmark vacuum and requires direct customer co-development.26:01–27:57 · Guest teaching 2/10 Hiring Across Research, Product, and Company Culture Alessio wraps up by asking about open roles and cultural expectations. Misha outlines their emphasis on high agency, technical craftsmanship, and kindness.0:44–4:06 · Guest disagreement 1/10 The Core Thesis: RL, LLMs, and Autonomous Coding Swyx opens by inviting Misha to detail Reflection AI's emergence from stealth. Misha delivers an extended monologue outlining the team's thesis on combining RL with LLMs to reach superintelligence through coding.4:07–7:56 · Guest disagreement 2/10 The RL Pendulum Shift and Foundation Model Baselines Alessio contextualizes historical RL shifts from OpenAI Dota to language, prompting Misha on computer use versus code. Misha educates the hosts on why mouse-based web browsing lacks internet priors compared to code ergonomics.7:57–10:31 · Guest disagreement 3/10 Defining Superintelligence and AI-Native Programmatic Interfaces Swyx pushes back on Misha's leap from deterministic coding agents to broad superintelligence. Misha reframes superintelligence around creative discovery like Move 37 and programmatic API execution.10:33–13:22 · Guest disagreement 1/10 Data Mixture: Balancing SFT with Reinforcement Learning Swyx brings up Misha's past commentary regarding ChatGPT imitation learning and references past guest Bret Taylor on human versus AI programming languages. Misha explains the necessity of bootstrapping RL with sensible SFT mixtures.13:23–17:01 · Guest disagreement 2/10 Anticipating 'Move 37' Breakthroughs in Code Generation Alessio questions what Move 37 means in practical software engineering and critiques hype around Deep Research. Misha points to DeepSeek's custom attention kernels and warns against closed labs hoarding powerful internal models.17:01–20:22 · Guest disagreement 1/10 Product Architecture: Autonomous Backlog Resolution API Swyx compares Reflection's upcoming form factor to Poolside's VS Code extension. Misha differentiates between developer-driven copilots and autonomous execution APIs handling enterprise technical debt and backlog triage.20:23–22:52 · Guest disagreement 2/10 Long Context Understanding vs. Agentic Code Localization Swyx raises Magic.dev's 100M context windows to question state-of-the-art code indexing. Misha breaks down the architectural tradeoffs between naive long context attention mechanisms and agentic code localization.22:52–26:00 · Guest disagreement 2/10 Evaluating Coding Agents Through Real-World Co-Development Swyx inquires about SWE-bench benchmarks versus competitors like Devin and Magic. Misha argues that autonomous coding cannot be evaluated in a benchmark vacuum and requires direct customer co-development.26:01–27:57 · Guest disagreement 1/10 Hiring Across Research, Product, and Company Culture Alessio wraps up by asking about open roles and cultural expectations. Misha outlines their emphasis on high agency, technical craftsmanship, and kindness.0:44–4:06 · The hosts pushing back 1/10 The Core Thesis: RL, LLMs, and Autonomous Coding Swyx opens by inviting Misha to detail Reflection AI's emergence from stealth. Misha delivers an extended monologue outlining the team's thesis on combining RL with LLMs to reach superintelligence through coding.4:07–7:56 · The hosts pushing back 2/10 The RL Pendulum Shift and Foundation Model Baselines Alessio contextualizes historical RL shifts from OpenAI Dota to language, prompting Misha on computer use versus code. Misha educates the hosts on why mouse-based web browsing lacks internet priors compared to code ergonomics.7:57–10:31 · The hosts pushing back 5/10 Defining Superintelligence and AI-Native Programmatic Interfaces Swyx pushes back on Misha's leap from deterministic coding agents to broad superintelligence. Misha reframes superintelligence around creative discovery like Move 37 and programmatic API execution.10:33–13:22 · The hosts pushing back 2/10 Data Mixture: Balancing SFT with Reinforcement Learning Swyx brings up Misha's past commentary regarding ChatGPT imitation learning and references past guest Bret Taylor on human versus AI programming languages. Misha explains the necessity of bootstrapping RL with sensible SFT mixtures.13:23–17:01 · The hosts pushing back 3/10 Anticipating 'Move 37' Breakthroughs in Code Generation Alessio questions what Move 37 means in practical software engineering and critiques hype around Deep Research. Misha points to DeepSeek's custom attention kernels and warns against closed labs hoarding powerful internal models.17:01–20:22 · The hosts pushing back 2/10 Product Architecture: Autonomous Backlog Resolution API Swyx compares Reflection's upcoming form factor to Poolside's VS Code extension. Misha differentiates between developer-driven copilots and autonomous execution APIs handling enterprise technical debt and backlog triage.20:23–22:52 · The hosts pushing back 2/10 Long Context Understanding vs. Agentic Code Localization Swyx raises Magic.dev's 100M context windows to question state-of-the-art code indexing. Misha breaks down the architectural tradeoffs between naive long context attention mechanisms and agentic code localization.22:52–26:00 · The hosts pushing back 2/10 Evaluating Coding Agents Through Real-World Co-Development Swyx inquires about SWE-bench benchmarks versus competitors like Devin and Magic. Misha argues that autonomous coding cannot be evaluated in a benchmark vacuum and requires direct customer co-development.26:01–27:57 · The hosts pushing back 1/10 Hiring Across Research, Product, and Company Culture Alessio wraps up by asking about open roles and cultural expectations. Misha outlines their emphasis on high agency, technical craftsmanship, and kindness.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 21.7% · guest 78.3%0:00 · the hosts 21.7% · guest 78.3%3:00 · the hosts 14.2% · guest 85.8%3:00 · the hosts 14.2% · guest 85.8%6:00 · the hosts 32.4% · guest 67.6%6:00 · the hosts 32.4% · guest 67.6%9:00 · the hosts 22.2% · guest 77.8%9:00 · the hosts 22.2% · guest 77.8%12:00 · the hosts 30.7% · guest 69.3%12:00 · the hosts 30.7% · guest 69.3%15:00 · the hosts 32.7% · guest 67.3%15:00 · the hosts 32.7% · guest 67.3%18:00 · the hosts 18.2% · guest 81.8%18:00 · the hosts 18.2% · guest 81.8%21:00 · the hosts 15.2% · guest 84.8%21:00 · the hosts 15.2% · guest 84.8%24:00 · the hosts 10.5% · guest 89.5%24:00 · the hosts 10.5% · guest 89.5%27:00 · the hosts 9.1% · guest 90.9%27:00 · the hosts 9.1% · guest 90.9%
Sharpest disagreement ▶ 23:30 Misha dismisses vacuum benchmarks as valid measures of superintelligence

Misha rejects benchmark-driven evaluations like SWE-bench as insufficient, arguing that without real customer co-development, claims of superintelligence are meaningless.

Hardest push from the hosts ▶ 7:56 Swyx challenges the leap from coding agents to superintelligence

Swyx directly interrupts the pitch to question Misha's assumption that solving code equates to achieving broader superintelligence.

Biggest teaching moment ▶ 6:45 Misha explains why LLMs have code priors but no mouse priors

Misha explains the fundamental architectural reason why GUI and web browsing agents struggle with noisy human data compared to code-native pre-training.

The host holds their own ▶ 10:33 Swyx leverages Misha's past tweets to probe training methodologies

Swyx demonstrates deep domain preparation by citing Misha's specific historical tweets on contractor data and interleaving code with text.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Core Thesis: RL, LLMs, and Autonomous Coding 3411 Swyx opens by inviting Misha to detail Reflection AI's emergence from stealth. Misha delivers an extended monologue outlining the team's thesis on combining RL with LLMs to reach superintelligence through coding.
The RL Pendulum Shift and Foundation Model Baselines 5622 Alessio contextualizes historical RL shifts from OpenAI Dota to language, prompting Misha on computer use versus code. Misha educates the hosts on why mouse-based web browsing lacks internet priors compared to code ergonomics.
Defining Superintelligence and AI-Native Programmatic Interfaces 6535 Swyx pushes back on Misha's leap from deterministic coding agents to broad superintelligence. Misha reframes superintelligence around creative discovery like Move 37 and programmatic API execution.
Data Mixture: Balancing SFT with Reinforcement Learning 6412 Swyx brings up Misha's past commentary regarding ChatGPT imitation learning and references past guest Bret Taylor on human versus AI programming languages. Misha explains the necessity of bootstrapping RL with sensible SFT mixtures.
Anticipating 'Move 37' Breakthroughs in Code Generation 5423 Alessio questions what Move 37 means in practical software engineering and critiques hype around Deep Research. Misha points to DeepSeek's custom attention kernels and warns against closed labs hoarding powerful internal models.
Product Architecture: Autonomous Backlog Resolution API 5412 Swyx compares Reflection's upcoming form factor to Poolside's VS Code extension. Misha differentiates between developer-driven copilots and autonomous execution APIs handling enterprise technical debt and backlog triage.
Long Context Understanding vs. Agentic Code Localization 6522 Swyx raises Magic.dev's 100M context windows to question state-of-the-art code indexing. Misha breaks down the architectural tradeoffs between naive long context attention mechanisms and agentic code localization.
Evaluating Coding Agents Through Real-World Co-Development 5522 Swyx inquires about SWE-bench benchmarks versus competitors like Devin and Magic. Misha argues that autonomous coding cannot be evaluated in a benchmark vacuum and requires direct customer co-development.
Hiring Across Research, Product, and Company Culture 2211 Alessio wraps up by asking about open roles and cultural expectations. Misha outlines their emphasis on high agency, technical craftsmanship, and kindness.

Statements from this episode (16)

Prediction Not checkable as stated
Solving Autonomous Coding Is the Direct Path to AGI
“Our core belief is that if you solve this problem, you solve the autonomous coding problem and build a super intelligent coding agent, that that thing will lead to super intelligence more broadly.”
Misha Laskin Mar 7, 2025 ▶ 3:50
Disclosure
Gemini 1 Proved GPT-4-Level Models Can Bootstrap Reinforcement Learning
“Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were basically at GPT-IV level or above, were capable en…”
Misha Laskin Mar 7, 2025 ▶ 4:34
Prediction Not checkable as stated
Superintelligence Cannot Be Trained Entirely From Scratch
“In the era of language models, I don't think you'll be able to train superintelligence from scratch.”
Misha Laskin Mar 7, 2025 ▶ 5:54
Opinion
Software Engineering Is Ergonomic for LLMs, Making It the Ideal Wedge
“Our belief as a company is that the correct wedge in, the correct starting point to this entire problem is decoding agent, because it's already, you know, software engineering is already what I would call kind of ergonomic for a language model.”
Misha Laskin Mar 7, 2025 ▶ 6:32
Insight
LLMs Have Strong Priors for Code, But None for Mouse Movement
“A language model, for example, has no prior for a mouse movement. It, you know, it really never seen that on the internet, but it has a really strong prior for code. So out of the two categories, say web browsing and coding, coding is the only one that is actu…”
Misha Laskin Mar 7, 2025 ▶ 7:19
Prediction Not checkable as stated
Future UIs Will Be Built as Programmatic Interfaces for AI Models
“Over the coming years, there'll be more kind of AI friendly or language model friendly UIs. And what's friendly to a language model is, is code. So the way a model will be doing work, not just for coding and software engineering, Is by basically making functio…”
Misha Laskin Mar 7, 2025 ▶ 9:37
Insight
Reinforcement Learning Fails Without Initial SFT to Seed Rewardable Behaviors
“It can potentially work otherwise, but practically it only works when the agent has interacted with a reward, right? It's received a positive reward for what it's done. Maybe one out of 10 times, one out of 50 times, but if it's getting zero reward, then you d…”
Misha Laskin Mar 7, 2025 ▶ 11:38
Prediction Not checkable as stated
The Fundamental Language of AI Will Likely Remain Python
“The things that language models understand best tend to be the kind of piece of code that are represented on the internet. You know, maybe I don't think a lot of people would be happy with this, that maybe the kind of fundamental language of AI becomes Python …”
Misha Laskin Mar 7, 2025 ▶ 12:53
Prediction Not checkable as stated
AI Coding Agents Will Discover Unexpected 'Move 37' Breakthrough Solutions
“I think I think there are going to be a lot of move 37”
Misha Laskin Mar 7, 2025 ▶ 13:37
Assertion Not checkable as stated
DeepSeek Revealed Attention Kernel Optimizations Kept Secret by Big Labs
“Deep seeks recent open sourcing of their various code components that they use to train that model, which I think outside of the big labs was not really well known to, right, it was not really well known how to write kind of a kernel that's optimized for this …”
Misha Laskin Mar 7, 2025 ▶ 13:37
Prediction Not checkable as stated
Frontier Labs May Hoard Superintelligent Models and Release Nerfed Versions
“You can imagine you know, the world converging on a few companies have really powerful coding models. They basically release a nerfed version of that to the public at large, and basically have a competitive advantage by having, you know, a super intelligent co…”
Misha Laskin Mar 7, 2025 ▶ 16:17
Opinion
Copilot and Cursor Still Require Human Engineers to Drive Most Work
“Github Copilot or Cursor, which are incredible products. We use them internally. We're very happy with them. But again, this is, these are products where the engineer is driving most of the work. Like, even in agent mode, really the engineer is, like, driving …”
Misha Laskin Mar 7, 2025 ▶ 17:58
Insight
Enterprise Engineers Spend Most of Their Time on Backlog Whack-a-Mole
“As a company gets larger, the, like, engineer goes from spending most of their time on, you know, working on the features that matter, and the kind of more, the more kind of value-driven work at a startup, To a very large company where you have giant code base…”
Misha Laskin Mar 7, 2025 ▶ 19:14
Opinion
Needle-in-a-Haystack Tests Are Crude for Evaluating True Context Understanding
“So it's not just having long context, it's whether your model truly understands what's inside the context, and needle in the haystack tests are pretty crude and not very effective way of testing this kind of capability.”
Misha Laskin Mar 7, 2025 ▶ 21:04
Opinion
Long-Context Attention Will Beat Agentic Localization for Codebase Indexing
“I bet would be on long context understanding and improving the attention mechanism over the long context.”
Misha Laskin Mar 7, 2025 ▶ 22:04
Insight
A 90% SWE-Bench Score Can Still Fall Flat in Customer Environments
“Autonomous coding benchmarks, let's say, like Sweetbench, are useful. I'm not going to discount them. They are useful. But let's say, you know, 90% on Sweetbench could still mean something that just falls over flat within a customer setting.”
Misha Laskin Mar 7, 2025 ▶ 23:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.