Jul 17, 2025 · 1h 2m · no-priors

No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin

Misha Laskin · 46m spoken Sarah Guo · 12m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Reflection AI co-founder and CEO Misha Laskin discusses the launch of Asimov—a specialized code comprehension agent—and outlines how reinforcement learning, vertical integration, and post-training economics will drive the multi-decade journey toward practical artificial superintelligence in enterprise software engineering.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.3% of the talking time here. How this is scored →

The hosts as informed peer 6.0 Guest teaching 5.2 Guest disagreement 1.4 The hosts pushing back 1.1
05100:0015:0030:0045:001:00:000:44–3:25 · The hosts as informed peer 5/10 Co-Designing Product and Research for Superintelligence Sarah sets the stage by asking Misha to distinguish between building superintelligent autonomous systems and pure academic superintelligence. Misha articulates Reflection AI's philosophy of co-designing product and frontier research against ASI-complete domains.3:25–7:36 · The hosts as informed peer 4/10 Misha Laskin's Path from Theoretical Physics to Frontier RL Sarah inquires into Misha's transition from theoretical physics into deep RL at Peter Abbeel's lab and Google DeepMind. Misha explains his conviction that RL scaling atop LLMs is the final paradigm toward superintelligence.7:41–11:52 · The hosts as informed peer 6/10 Introducing Asimov: The Code Comprehension and Research Agent Misha introduces Asimov as a code comprehension agent rather than another code generation tool, arguing 80% of engineer time is spent comprehending systems. Sarah immediately references the recent METR productivity study to validate this distinction.11:52–16:15 · The hosts as informed peer 6/10 Technical Architecture: Long Context, Neural Retrieval, and Tool Reasoning Misha details the technical stack behind Asimov, highlighting post-training small models for long-context neural retrieval and custom tool reasoning. He critiques arbitrary benchmarks like Humanity's Last Exam for lacking enterprise utility.16:16–21:52 · The hosts as informed peer 6/10 Customer-Driven Evaluations and Startup Advantage over Big Labs Misha explains that a startup's sole advantage over frontier labs is tight customer-informed evals versus diffuse academic benchmark optimization. Sarah synthesizes how general labs suffer from disconnected organizational layers between research and product.22:02–25:09 · The hosts as informed peer 5/10 Semantic Queries and Complex Enterprise Debugging Sarah asks for concrete test queries, and Misha explains complex semantic debugging scenarios like silent race conditions and flaky test triage across distributed teams.25:09–28:38 · The hosts as informed peer 7/10 Collaborative Org-Wide Memory and Knowledge Governance Sarah engages deeply on the mechanics of org-wide memory governance, distinguishing static hierarchical role-based access control from dynamic Git pull request review workflows. Misha agrees and compares it to Google code ownership models.28:38–33:39 · The hosts as informed peer 7/10 Frontier Lab Economics: Compute Scaling and RL vs Pre-Training Sarah and Misha analyze the economics of frontier research, noting RL post-training requires two orders of magnitude fewer FLOPs than pre-training, making independent labs viable without cloud provider capture.33:39–37:21 · The hosts as informed peer 6/10 The Reward Modeling Bottleneck and Credit Assignment in RL Misha delivers a masterclass on the reward modeling bottleneck and credit assignment failures in RL, explaining why reasoning models follow meandering garden paths when atomic step verification is missing.37:21–44:27 · The hosts as informed peer 7/10 Generalization Debunked and the Emergence of Jagged Superintelligence Misha drops the contrarian take that generalization is an illusion created by pulling test distribution into training data. Sarah challenges him on whether this resigns the field to brute-force data coverage rather than emergent general reasoning.44:27–48:10 · The hosts as informed peer 6/10 Industry Consolidation: The Windsurf Deal and Frontier Lab Verticalization Misha analyzes the Windsurf acquisition talks and verticalization across labs, warning that UI wrapper startups without in-house intelligence face an existential threat when frontier labs subsidize downstream applications.48:10–55:12 · The hosts as informed peer 7/10 Reinforcement Learning in Robotics versus Language Models Sarah pitches an unconventional data-capture concept for engineering workflows. Misha responds by contrasting the noisiness and hackability of multi-modal robotics rewards with the structured, verifiable nature of language tokens.55:13–57:54 · The hosts as informed peer 5/10 Expanding from Code Reasoning to Broader Enterprise Workflows Misha explains how code comprehension serves as the universal substrate for all enterprise agent workflows because models interface with external software via functional code calls.57:54–1:02:33 · The hosts as informed peer 7/10 Deployment Realities and the Multi-Decade Timeline to ASI Misha tempers short-term AGI hype by predicting a multi-decade enterprise deployment cycle, drawing parallels to DeepMind's specialized research strikes. Sarah articulates her investment thesis on funding distribution capture within this decade.0:44–3:25 · Guest teaching 4/10 Co-Designing Product and Research for Superintelligence Sarah sets the stage by asking Misha to distinguish between building superintelligent autonomous systems and pure academic superintelligence. Misha articulates Reflection AI's philosophy of co-designing product and frontier research against ASI-complete domains.3:25–7:36 · Guest teaching 3/10 Misha Laskin's Path from Theoretical Physics to Frontier RL Sarah inquires into Misha's transition from theoretical physics into deep RL at Peter Abbeel's lab and Google DeepMind. Misha explains his conviction that RL scaling atop LLMs is the final paradigm toward superintelligence.7:41–11:52 · Guest teaching 5/10 Introducing Asimov: The Code Comprehension and Research Agent Misha introduces Asimov as a code comprehension agent rather than another code generation tool, arguing 80% of engineer time is spent comprehending systems. Sarah immediately references the recent METR productivity study to validate this distinction.11:52–16:15 · Guest teaching 6/10 Technical Architecture: Long Context, Neural Retrieval, and Tool Reasoning Misha details the technical stack behind Asimov, highlighting post-training small models for long-context neural retrieval and custom tool reasoning. He critiques arbitrary benchmarks like Humanity's Last Exam for lacking enterprise utility.16:16–21:52 · Guest teaching 5/10 Customer-Driven Evaluations and Startup Advantage over Big Labs Misha explains that a startup's sole advantage over frontier labs is tight customer-informed evals versus diffuse academic benchmark optimization. Sarah synthesizes how general labs suffer from disconnected organizational layers between research and product.22:02–25:09 · Guest teaching 5/10 Semantic Queries and Complex Enterprise Debugging Sarah asks for concrete test queries, and Misha explains complex semantic debugging scenarios like silent race conditions and flaky test triage across distributed teams.25:09–28:38 · Guest teaching 4/10 Collaborative Org-Wide Memory and Knowledge Governance Sarah engages deeply on the mechanics of org-wide memory governance, distinguishing static hierarchical role-based access control from dynamic Git pull request review workflows. Misha agrees and compares it to Google code ownership models.28:38–33:39 · Guest teaching 5/10 Frontier Lab Economics: Compute Scaling and RL vs Pre-Training Sarah and Misha analyze the economics of frontier research, noting RL post-training requires two orders of magnitude fewer FLOPs than pre-training, making independent labs viable without cloud provider capture.33:39–37:21 · Guest teaching 7/10 The Reward Modeling Bottleneck and Credit Assignment in RL Misha delivers a masterclass on the reward modeling bottleneck and credit assignment failures in RL, explaining why reasoning models follow meandering garden paths when atomic step verification is missing.37:21–44:27 · Guest teaching 6/10 Generalization Debunked and the Emergence of Jagged Superintelligence Misha drops the contrarian take that generalization is an illusion created by pulling test distribution into training data. Sarah challenges him on whether this resigns the field to brute-force data coverage rather than emergent general reasoning.44:27–48:10 · Guest teaching 6/10 Industry Consolidation: The Windsurf Deal and Frontier Lab Verticalization Misha analyzes the Windsurf acquisition talks and verticalization across labs, warning that UI wrapper startups without in-house intelligence face an existential threat when frontier labs subsidize downstream applications.48:10–55:12 · Guest teaching 7/10 Reinforcement Learning in Robotics versus Language Models Sarah pitches an unconventional data-capture concept for engineering workflows. Misha responds by contrasting the noisiness and hackability of multi-modal robotics rewards with the structured, verifiable nature of language tokens.55:13–57:54 · Guest teaching 5/10 Expanding from Code Reasoning to Broader Enterprise Workflows Misha explains how code comprehension serves as the universal substrate for all enterprise agent workflows because models interface with external software via functional code calls.57:54–1:02:33 · Guest teaching 5/10 Deployment Realities and the Multi-Decade Timeline to ASI Misha tempers short-term AGI hype by predicting a multi-decade enterprise deployment cycle, drawing parallels to DeepMind's specialized research strikes. Sarah articulates her investment thesis on funding distribution capture within this decade.0:44–3:25 · Guest disagreement 1/10 Co-Designing Product and Research for Superintelligence Sarah sets the stage by asking Misha to distinguish between building superintelligent autonomous systems and pure academic superintelligence. Misha articulates Reflection AI's philosophy of co-designing product and frontier research against ASI-complete domains.3:25–7:36 · Guest disagreement 0/10 Misha Laskin's Path from Theoretical Physics to Frontier RL Sarah inquires into Misha's transition from theoretical physics into deep RL at Peter Abbeel's lab and Google DeepMind. Misha explains his conviction that RL scaling atop LLMs is the final paradigm toward superintelligence.7:41–11:52 · Guest disagreement 1/10 Introducing Asimov: The Code Comprehension and Research Agent Misha introduces Asimov as a code comprehension agent rather than another code generation tool, arguing 80% of engineer time is spent comprehending systems. Sarah immediately references the recent METR productivity study to validate this distinction.11:52–16:15 · Guest disagreement 2/10 Technical Architecture: Long Context, Neural Retrieval, and Tool Reasoning Misha details the technical stack behind Asimov, highlighting post-training small models for long-context neural retrieval and custom tool reasoning. He critiques arbitrary benchmarks like Humanity's Last Exam for lacking enterprise utility.16:16–21:52 · Guest disagreement 1/10 Customer-Driven Evaluations and Startup Advantage over Big Labs Misha explains that a startup's sole advantage over frontier labs is tight customer-informed evals versus diffuse academic benchmark optimization. Sarah synthesizes how general labs suffer from disconnected organizational layers between research and product.22:02–25:09 · Guest disagreement 1/10 Semantic Queries and Complex Enterprise Debugging Sarah asks for concrete test queries, and Misha explains complex semantic debugging scenarios like silent race conditions and flaky test triage across distributed teams.25:09–28:38 · Guest disagreement 1/10 Collaborative Org-Wide Memory and Knowledge Governance Sarah engages deeply on the mechanics of org-wide memory governance, distinguishing static hierarchical role-based access control from dynamic Git pull request review workflows. Misha agrees and compares it to Google code ownership models.28:38–33:39 · Guest disagreement 1/10 Frontier Lab Economics: Compute Scaling and RL vs Pre-Training Sarah and Misha analyze the economics of frontier research, noting RL post-training requires two orders of magnitude fewer FLOPs than pre-training, making independent labs viable without cloud provider capture.33:39–37:21 · Guest disagreement 2/10 The Reward Modeling Bottleneck and Credit Assignment in RL Misha delivers a masterclass on the reward modeling bottleneck and credit assignment failures in RL, explaining why reasoning models follow meandering garden paths when atomic step verification is missing.37:21–44:27 · Guest disagreement 4/10 Generalization Debunked and the Emergence of Jagged Superintelligence Misha drops the contrarian take that generalization is an illusion created by pulling test distribution into training data. Sarah challenges him on whether this resigns the field to brute-force data coverage rather than emergent general reasoning.44:27–48:10 · Guest disagreement 2/10 Industry Consolidation: The Windsurf Deal and Frontier Lab Verticalization Misha analyzes the Windsurf acquisition talks and verticalization across labs, warning that UI wrapper startups without in-house intelligence face an existential threat when frontier labs subsidize downstream applications.48:10–55:12 · Guest disagreement 1/10 Reinforcement Learning in Robotics versus Language Models Sarah pitches an unconventional data-capture concept for engineering workflows. Misha responds by contrasting the noisiness and hackability of multi-modal robotics rewards with the structured, verifiable nature of language tokens.55:13–57:54 · Guest disagreement 1/10 Expanding from Code Reasoning to Broader Enterprise Workflows Misha explains how code comprehension serves as the universal substrate for all enterprise agent workflows because models interface with external software via functional code calls.57:54–1:02:33 · Guest disagreement 1/10 Deployment Realities and the Multi-Decade Timeline to ASI Misha tempers short-term AGI hype by predicting a multi-decade enterprise deployment cycle, drawing parallels to DeepMind's specialized research strikes. Sarah articulates her investment thesis on funding distribution capture within this decade.0:44–3:25 · The hosts pushing back 1/10 Co-Designing Product and Research for Superintelligence Sarah sets the stage by asking Misha to distinguish between building superintelligent autonomous systems and pure academic superintelligence. Misha articulates Reflection AI's philosophy of co-designing product and frontier research against ASI-complete domains.3:25–7:36 · The hosts pushing back 0/10 Misha Laskin's Path from Theoretical Physics to Frontier RL Sarah inquires into Misha's transition from theoretical physics into deep RL at Peter Abbeel's lab and Google DeepMind. Misha explains his conviction that RL scaling atop LLMs is the final paradigm toward superintelligence.7:41–11:52 · The hosts pushing back 1/10 Introducing Asimov: The Code Comprehension and Research Agent Misha introduces Asimov as a code comprehension agent rather than another code generation tool, arguing 80% of engineer time is spent comprehending systems. Sarah immediately references the recent METR productivity study to validate this distinction.11:52–16:15 · The hosts pushing back 1/10 Technical Architecture: Long Context, Neural Retrieval, and Tool Reasoning Misha details the technical stack behind Asimov, highlighting post-training small models for long-context neural retrieval and custom tool reasoning. He critiques arbitrary benchmarks like Humanity's Last Exam for lacking enterprise utility.16:16–21:52 · The hosts pushing back 1/10 Customer-Driven Evaluations and Startup Advantage over Big Labs Misha explains that a startup's sole advantage over frontier labs is tight customer-informed evals versus diffuse academic benchmark optimization. Sarah synthesizes how general labs suffer from disconnected organizational layers between research and product.22:02–25:09 · The hosts pushing back 0/10 Semantic Queries and Complex Enterprise Debugging Sarah asks for concrete test queries, and Misha explains complex semantic debugging scenarios like silent race conditions and flaky test triage across distributed teams.25:09–28:38 · The hosts pushing back 2/10 Collaborative Org-Wide Memory and Knowledge Governance Sarah engages deeply on the mechanics of org-wide memory governance, distinguishing static hierarchical role-based access control from dynamic Git pull request review workflows. Misha agrees and compares it to Google code ownership models.28:38–33:39 · The hosts pushing back 1/10 Frontier Lab Economics: Compute Scaling and RL vs Pre-Training Sarah and Misha analyze the economics of frontier research, noting RL post-training requires two orders of magnitude fewer FLOPs than pre-training, making independent labs viable without cloud provider capture.33:39–37:21 · The hosts pushing back 1/10 The Reward Modeling Bottleneck and Credit Assignment in RL Misha delivers a masterclass on the reward modeling bottleneck and credit assignment failures in RL, explaining why reasoning models follow meandering garden paths when atomic step verification is missing.37:21–44:27 · The hosts pushing back 4/10 Generalization Debunked and the Emergence of Jagged Superintelligence Misha drops the contrarian take that generalization is an illusion created by pulling test distribution into training data. Sarah challenges him on whether this resigns the field to brute-force data coverage rather than emergent general reasoning.44:27–48:10 · The hosts pushing back 1/10 Industry Consolidation: The Windsurf Deal and Frontier Lab Verticalization Misha analyzes the Windsurf acquisition talks and verticalization across labs, warning that UI wrapper startups without in-house intelligence face an existential threat when frontier labs subsidize downstream applications.48:10–55:12 · The hosts pushing back 1/10 Reinforcement Learning in Robotics versus Language Models Sarah pitches an unconventional data-capture concept for engineering workflows. Misha responds by contrasting the noisiness and hackability of multi-modal robotics rewards with the structured, verifiable nature of language tokens.55:13–57:54 · The hosts pushing back 0/10 Expanding from Code Reasoning to Broader Enterprise Workflows Misha explains how code comprehension serves as the universal substrate for all enterprise agent workflows because models interface with external software via functional code calls.57:54–1:02:33 · The hosts pushing back 1/10 Deployment Realities and the Multi-Decade Timeline to ASI Misha tempers short-term AGI hype by predicting a multi-decade enterprise deployment cycle, drawing parallels to DeepMind's specialized research strikes. Sarah articulates her investment thesis on funding distribution capture within this decade.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 35% · guest 65%0:00 · the hosts 35% · guest 65%3:00 · the hosts 21.4% · guest 78.6%3:00 · the hosts 21.4% · guest 78.6%6:00 · the hosts 13.5% · guest 86.5%6:00 · the hosts 13.5% · guest 86.5%9:00 · the hosts 24.5% · guest 75.5%9:00 · the hosts 24.5% · guest 75.5%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 4.3% · guest 95.7%15:00 · the hosts 4.3% · guest 95.7%18:00 · the hosts 16.1% · guest 83.9%18:00 · the hosts 16.1% · guest 83.9%21:00 · the hosts 5.5% · guest 94.5%21:00 · the hosts 5.5% · guest 94.5%24:00 · the hosts 29.4% · guest 70.6%24:00 · the hosts 29.4% · guest 70.6%27:00 · the hosts 53.7% · guest 46.3%27:00 · the hosts 53.7% · guest 46.3%30:00 · the hosts 7.3% · guest 92.7%30:00 · the hosts 7.3% · guest 92.7%33:00 · the hosts 21.9% · guest 78.1%33:00 · the hosts 21.9% · guest 78.1%36:00 · the hosts 18.1% · guest 81.9%36:00 · the hosts 18.1% · guest 81.9%39:00 · the hosts 30.6% · guest 69.4%39:00 · the hosts 30.6% · guest 69.4%42:00 · the hosts 5.7% · guest 94.3%42:00 · the hosts 5.7% · guest 94.3%45:00 · the hosts 0.2% · guest 99.8%45:00 · the hosts 0.2% · guest 99.8%48:00 · the hosts 76.4% · guest 23.6%48:00 · the hosts 76.4% · guest 23.6%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 14.7% · guest 85.3%54:00 · the hosts 14.7% · guest 85.3%57:00 · the hosts 5.1% · guest 94.9%57:00 · the hosts 5.1% · guest 94.9%1:00:00 · the hosts 46.9% · guest 53.1%1:00:00 · the hosts 46.9% · guest 53.1%
Sharpest disagreement ▶ 37:45 Generalization is an illusion

Misha aggressively rejects conventional AI wisdom by claiming that generalization does not exist and is merely bringing test distribution into the training set.

Hardest push from the hosts ▶ 40:10 Pushing back on brute-force distribution vs genuine capability

Sarah directly challenges Misha's cynical view of generalization, contrasting it with top frontier researchers who believe in emergent free reasoning capabilities beyond explicit training coverage.

Biggest teaching moment ▶ 50:50 Reward hackability in vision-language vs text models

Misha leverages his robotics PhD background to educate Sarah on why RL struggles in robotics due to noisy sensory pixels, contrasting it with language's compressed, verifiable representation.

The host holds their own ▶ 26:45 Analyzing Git pull-request governance vs static RBAC

Sarah demonstrates sophisticated software systems knowledge, deconstructing how organizational memory permissions require dynamic content review rather than static role-based access control.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Co-Designing Product and Research for Superintelligence 5411 Sarah sets the stage by asking Misha to distinguish between building superintelligent autonomous systems and pure academic superintelligence. Misha articulates Reflection AI's philosophy of co-designing product and frontier research against ASI-complete domains.
Misha Laskin's Path from Theoretical Physics to Frontier RL 4300 Sarah inquires into Misha's transition from theoretical physics into deep RL at Peter Abbeel's lab and Google DeepMind. Misha explains his conviction that RL scaling atop LLMs is the final paradigm toward superintelligence.
Introducing Asimov: The Code Comprehension and Research Agent 6511 Misha introduces Asimov as a code comprehension agent rather than another code generation tool, arguing 80% of engineer time is spent comprehending systems. Sarah immediately references the recent METR productivity study to validate this distinction.
Technical Architecture: Long Context, Neural Retrieval, and Tool Reasoning 6621 Misha details the technical stack behind Asimov, highlighting post-training small models for long-context neural retrieval and custom tool reasoning. He critiques arbitrary benchmarks like Humanity's Last Exam for lacking enterprise utility.
Customer-Driven Evaluations and Startup Advantage over Big Labs 6511 Misha explains that a startup's sole advantage over frontier labs is tight customer-informed evals versus diffuse academic benchmark optimization. Sarah synthesizes how general labs suffer from disconnected organizational layers between research and product.
Semantic Queries and Complex Enterprise Debugging 5510 Sarah asks for concrete test queries, and Misha explains complex semantic debugging scenarios like silent race conditions and flaky test triage across distributed teams.
Collaborative Org-Wide Memory and Knowledge Governance 7412 Sarah engages deeply on the mechanics of org-wide memory governance, distinguishing static hierarchical role-based access control from dynamic Git pull request review workflows. Misha agrees and compares it to Google code ownership models.
Frontier Lab Economics: Compute Scaling and RL vs Pre-Training 7511 Sarah and Misha analyze the economics of frontier research, noting RL post-training requires two orders of magnitude fewer FLOPs than pre-training, making independent labs viable without cloud provider capture.
The Reward Modeling Bottleneck and Credit Assignment in RL 6721 Misha delivers a masterclass on the reward modeling bottleneck and credit assignment failures in RL, explaining why reasoning models follow meandering garden paths when atomic step verification is missing.
Generalization Debunked and the Emergence of Jagged Superintelligence 7644 Misha drops the contrarian take that generalization is an illusion created by pulling test distribution into training data. Sarah challenges him on whether this resigns the field to brute-force data coverage rather than emergent general reasoning.
Industry Consolidation: The Windsurf Deal and Frontier Lab Verticalization 6621 Misha analyzes the Windsurf acquisition talks and verticalization across labs, warning that UI wrapper startups without in-house intelligence face an existential threat when frontier labs subsidize downstream applications.
Reinforcement Learning in Robotics versus Language Models 7711 Sarah pitches an unconventional data-capture concept for engineering workflows. Misha responds by contrasting the noisiness and hackability of multi-modal robotics rewards with the structured, verifiable nature of language tokens.
Expanding from Code Reasoning to Broader Enterprise Workflows 5510 Misha explains how code comprehension serves as the universal substrate for all enterprise agent workflows because models interface with external software via functional code calls.
Deployment Realities and the Multi-Decade Timeline to ASI 7511 Misha tempers short-term AGI hype by predicting a multi-decade enterprise deployment cycle, drawing parallels to DeepMind's specialized research strikes. Sarah articulates her investment thesis on funding distribution capture within this decade.

Statements from this episode (28)

Opinion
Laskin: Narrow-domain superintelligence has already been achieved by AlphaGo
“To some extent super intelligence in that sense has already been achieved. So right, AlphaGo was a super intelligent system, and there were other systems during that time that were built that were super intelligent in narrow domains.”
Misha Laskin Jul 17, 2025 ▶ 1:22
Insight
Laskin: Building ASI requires co-designing product and research together
“As long as you pick a category that I would say is kind of big enough to be ASI complete I think, and this is kind of our approach at Reflection, is it makes a lot more sense to be focused and co-design those two things together, the product of the research.”
Misha Laskin Jul 17, 2025 ▶ 3:05
Prediction Not checkable as stated
Laskin: Scaling RL on LLMs is the final paradigm before ASI
“The next paradigm, and effectively the final paradigm that we need to have in place before a, you know, what people used to call AGI, or now I think the goalposts have shifted to ASI, is reached, is just figuring out how to scale reinforcement learning on top …”
Misha Laskin Jul 17, 2025 ▶ 7:06
Opinion
Laskin: Enterprise AI coding tool productivity impact is negligible or negative
“Within enterprises, when you know, they're adopting coding tools and you see the impact that this is having on their actual productivity. And I think it's much lower than people expect. So it's in fact, it's sometimes negative, sometimes negligible.”
Misha Laskin Jul 17, 2025 ▶ 8:40
Insight
Laskin: Engineers spend 80% of their time comprehending complex systems
“When you look at what an engineer does in an organization, 80% of their time they're spending trying to comprehend complex systems and collaborating with teammates.”
Misha Laskin Jul 17, 2025 ▶ 10:28
Opinion
Laskin: Teaching AI agents to take action is mostly solved
“To me, it seems like really, 20% of the problem is teaching these agents how to act, and it's more or less solved.”
Misha Laskin Jul 17, 2025 ▶ 11:10
Opinion
Laskin: Evaluation Is the Most Important Differentiator for Frontier AI Labs
“This is I think the least spoken about part of what frontier labs do, but Possibly the most important, which is figuring out how they evaluate, like what makes Claude magically feel better at code than you know, another model out there. They did something righ…”
Misha Laskin Jul 17, 2025 ▶ 13:18
Opinion
Laskin: Humanity's Last Exam Benchmark Barely Matters to End Users
“Now, that's great, but I think the downside of that is that does humanity's last exam actually matter in any meaningful way for an end user? And I would argue that some weak correlation, but the answer is most likely no.”
Misha Laskin Jul 17, 2025 ▶ 15:53
Assertion Contradicted
Laskin: Most Contributors on OpenAI's o1 Paper Worked on Evals
“When you look at the model card for, let's say, the O-one paper that came out, I think, last year. If you look at the distribution of what most people worked on, on that paper, it was evals.”
Misha Laskin Jul 17, 2025 ▶ 17:16
Insight
Laskin: AI Apps Without Custom Model Training Are Fundamentally Limited
“The important part, I think, is to be able to tweak every part of the system from, you know, the product features to the agent design to the model training in order to build the best overall system. And if you are capped in which parts you can change, like if …”
Misha Laskin Jul 17, 2025 ▶ 19:32
Insight
Laskin: Single-file coding questions do not require deep research agents
“If you're looking at a file, and there's like a specific thing in that file, and you're just trying to get a quick answer to it, you don't really need the hammer of like a deep research like experience. You don't need to wait, you know, like tens of seconds or…”
Misha Laskin Jul 17, 2025 ▶ 22:16
Prediction Not checkable as stated
Laskin: Team-Wide AI Memory Governance Will Mirror Pull Request Workflows
“Where if you want to change the agents, the team wide memory, then it probably is going to look something like a pull request where the person who really understands that system Approves or, you know, edits it or something like this. I don't think it's going t…”
Misha Laskin Jul 17, 2025 ▶ 27:12
Insight
Laskin: RL requires far fewer FLOPs than pre-training for frontier models
“We're in this brief period in history right now where the RL flops are still manageable. Like you can really have a best in class product if you're focused. And yes, you'll need to put, you know, you still need a decent amount of GPUs, but from a flops perspec…”
Misha Laskin Jul 17, 2025 ▶ 30:41
Insight
Laskin: New frontier labs can succeed without cloud provider ownership
“Our thought was that this was the time where you can actually start a you know, a generational frontier lab that does not need to be coupled to a, you know, to a big cloud provider because if you do it right, you'll actually be able to generate you know, suffi…”
Misha Laskin Jul 17, 2025 ▶ 31:15
Assertion Not checkable as stated
Laskin: Anthropic is generating massive revenue at an unprecedented growth rate
“When you look at how fast like Anthropics revenue is growing I think, right, they're kind of in this spot where it's like a massive revenue generating business that's growing at an unprecedented rate.”
Misha Laskin Jul 17, 2025 ▶ 31:51
Insight
Laskin: Accurately verifying arbitrary outcomes is ASI-complete
“The reward problem in itself is at the time I called, I thought it was AGI complete. Now I'd say it's ASI complete, but by the time you have a neural network that can accurately verify any outcome, that is probably a super intelligence.”
Misha Laskin Jul 17, 2025 ▶ 35:52
Insight
Laskin: Current RL algorithms lack atomic credit assignment, causing meandering reasoning
“The RL methods we have today are quite bad, I would say, exploration and credit assignment. Like they, they're sort of just like the fundamental algorithms are take the things that work and make them happen more frequently, and the things that don't work and h…”
Misha Laskin Jul 17, 2025 ▶ 36:33
Insight
Laskin: Machine learning generalization is just bringing test distribution into training
“There's no such thing as generalization. There's just bringing the test distribution into train.”
Misha Laskin Jul 17, 2025 ▶ 37:49
Prediction Not checkable as stated
Laskin: Definitive superintelligence in meaningful categories will arrive in a couple years
“I think that where I think we'll be in a couple of years from now is that there'll be kind of definitive super intelligence in, Some meaningful categories of work.”
Misha Laskin Jul 17, 2025 ▶ 38:38
What-if
Laskin: Dota and AlphaStar would have achieved superintelligence with more compute
“Dota V and AlphaStar were near super intelligent systems, and if OpenAI and DeepMind had sunk more compute into them, they would have definitely become super intelligent.”
Misha Laskin Jul 17, 2025 ▶ 39:52
Prediction Held up
Laskin: AI models will beat humans at competitive coding within a year
“Code forces and other competitive coding environments. The models are almost best in the world, and within the year will probably be just the best in the world.”
Misha Laskin Jul 17, 2025 ▶ 43:00
Insight
Laskin: Frontier labs cannot easily buy end-user distribution via acquisitions
“I don't think it's guaranteed that a big lab can, you know, buy their way to the end user because the fundamental problems of your, you know, research team being far away from your product team will still be true. And the company having, you know, a hundred di…”
Misha Laskin Jul 17, 2025 ▶ 46:17
Prediction Not checkable as stated
Laskin: AI coding startups face existential risk without in-house frontier models
“And then from the startup side, I think it actually puts companies that Are in these kind of critical path categories like search and coding in a pretty existential place if they can't build their own frontier models. Not all frontier labs will be able to vert…”
Misha Laskin Jul 17, 2025 ▶ 46:41
Insight
Laskin: Sensory and vision-language model rewards are far more hackable than LLM rewards
“The challenge is that if we, if you think that language model rewards are hackable vision language model rewards or, you know, like other sensory signal rewards are infinitely more hackable.”
Misha Laskin Jul 17, 2025 ▶ 51:50
Insight
Laskin: Reinforcement learning is the only scalable path for synthetic data
“When we're generating synthetic data there is the only scalable path is really reinforcement learning.”
Misha Laskin Jul 17, 2025 ▶ 53:44
Insight
Laskin: Code reasoning models will generalize across other enterprise work
“The reason code is special is if you believe that the way a language model will interact with almost any piece of software is through function calls and therefore code, then if you build very capable reasoners coding reasoners that, you know, are sort of purpo…”
Misha Laskin Jul 17, 2025 ▶ 55:45
Prediction Not checkable as stated
Laskin: Deploying superintelligence and reaching 10% GDP growth is multi-decade
“Actually going in and deploying it and building it for, you know, specific categories of work. There are going to be a lot of product and kind of research innovation specific to those categories that will probably make this a multi-decade thing. So I don't thi…”
Misha Laskin Jul 17, 2025 ▶ 58:35
Prediction Not checkable as stated
Laskin: Enterprise AI coding deployment is within dozens of months, not decades
“I think coding is this era as well. This one I think will take longer than people thought as well, because again, enterprise is organizational problems. There's much different than The benchmarks that we have today, but I think it will be one of the faster one…”
Misha Laskin Jul 17, 2025 ▶ 1:01:57
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.