Insight certainty 3/5 debate potential 2/5

Chen: AI reasoning is the key mechanism to make agents reliable

Mark Chen · Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO · Jun 7, 2025 · at 1:26:43

Mark Chen, Chief Research Officer at OpenAI, discusses why reasoning models are necessary to prevent autonomous agent errors in real-world computer tasks.

0:00 / 0:06exact quote · 6.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“And I think the reason why we care so much about reasoning is because I think that's the path that we get reliable agents through.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Mark Chen

Assertion Not checkable as stated
Chen: OpenAI has internal models matching Gemini 3 and better successors coming
“Just looking purely at the benchmarks, you know, we actually felt quite confident you know, we have models internally that perform at the level of Gemini three, and we're pretty confident that we will release them soon, and we can release successor models that…”
Mark Chen Dec 3, 2025 ▶ 10:27 The World’s Fastest Growing Defense Company, OpenAI’s Code Red, Google Strikes Back | Diet TBPN
Insight
Chen: AI reasoning only emerges at large scale requiring massive compute
“When you look at reasoning you just don't see that happen at small scale, right? There's like a certain scale at which it starts becoming signal bearing and that requires you to have resources, right?”
Mark Chen Jun 7, 2025 ▶ 1:22:48 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Prediction Not checkable as stated
Chen: 2025 will be the year of autonomous AI agents
“We see 25 as this year of agents, right? We think of it as a year where models are going to do a lot more autonomous work. You can let them Kind of be unsupervised for much longer periods of time.”
Mark Chen Jun 7, 2025 ▶ 1:05:08 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Assertion Supported
Chen: Frontier AI models consistently score over 90% on AIME
“I think one clear example here is the Amy, like probably the hardest auto gradable, like human math eval, at least in the US. And yeah, the models are consistently getting like 90 plus percent on these.”
Mark Chen Jun 7, 2025 ▶ 1:08:50 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Assertion Supported
Chen: OpenAI reached 3 million paying business users
“We hit a big milestone. We got I think three million paying business users fairly recently.”
Mark Chen Jun 7, 2025 ▶ 1:07:26 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Insight
Mark Chen: Compute Can Scale Heavily into RL Given Right Levers
“I think, like, if you find the right levers, you can really pump a lot of compute into RL as well as pre-training.”
Mark Chen Jun 7, 2025 ▶ 1:28:16 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 500 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.