Insight certainty 3/5 debate potential 3/5

A 90% SWE-Bench Score Can Still Fall Flat in Customer Environments

Misha Laskin · Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin] · Mar 7, 2025 · at 23:33

Misha Laskin, CEO of Reflection AI, discusses evaluating coding agent capabilities beyond standard benchmarks like SWE-bench.

0:00 / 0:12exact quote · 12.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Autonomous coding benchmarks, let's say, like Sweetbench, are useful. I'm not going to discount them. They are useful. But let's say, you know, 90% on Sweetbench could still mean something that just falls over flat within a customer setting.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Misha Laskin

Prediction Not checkable as stated
Solving Autonomous Coding Is the Direct Path to AGI
“Our core belief is that if you solve this problem, you solve the autonomous coding problem and build a super intelligent coding agent, that that thing will lead to super intelligence more broadly.”
Misha Laskin Mar 7, 2025 ▶ 3:50 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Prediction Not checkable as stated
Superintelligence Cannot Be Trained Entirely From Scratch
“In the era of language models, I don't think you'll be able to train superintelligence from scratch.”
Misha Laskin Mar 7, 2025 ▶ 5:54 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Prediction Not checkable as stated
Future UIs Will Be Built as Programmatic Interfaces for AI Models
“Over the coming years, there'll be more kind of AI friendly or language model friendly UIs. And what's friendly to a language model is, is code. So the way a model will be doing work, not just for coding and software engineering, Is by basically making functio…”
Misha Laskin Mar 7, 2025 ▶ 9:37 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Prediction Not checkable as stated
Frontier Labs May Hoard Superintelligent Models and Release Nerfed Versions
“You can imagine you know, the world converging on a few companies have really powerful coding models. They basically release a nerfed version of that to the public at large, and basically have a competitive advantage by having, you know, a super intelligent co…”
Misha Laskin Mar 7, 2025 ▶ 16:17 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Disclosure
Gemini 1 Proved GPT-4-Level Models Can Bootstrap Reinforcement Learning
“Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were basically at GPT-IV level or above, were capable en…”
Misha Laskin Mar 7, 2025 ▶ 4:34 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Opinion
Software Engineering Is Ergonomic for LLMs, Making It the Ideal Wedge
“Our belief as a company is that the correct wedge in, the correct starting point to this entire problem is decoding agent, because it's already, you know, software engineering is already what I would call kind of ergonomic for a language model.”
Misha Laskin Mar 7, 2025 ▶ 6:32 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.