Mark Chen was not on this episode. A recording of them was played into it, so these
are their words but not an appearance on TBPN. It still counts as said, and it is kept out of every score on their page.
Mark Chen, an executive at OpenAI, speaks in an interview with journalist Ashlee Vance about benchmarking OpenAI's internal AI models against Google's Gemini 3.
“Just looking purely at the benchmarks, you know, we actually felt quite confident you know, we have models internally that perform at the level of Gemini three, and we're pretty confident that we will release them soon, and we can release successor models that are even better.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Mark Chen
Insight
Chen: AI reasoning only emerges at large scale requiring massive compute
“When you look at reasoning you just don't see that happen at small scale, right? There's like a certain scale at which it starts becoming signal bearing and that requires you to have resources, right?”
Mark ChenJun 7, 2025▶ 1:22:48Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
PredictionNot checkable as stated
Chen: 2025 will be the year of autonomous AI agents
“We see 25 as this year of agents, right? We think of it as a year where models are going to do a lot more autonomous work. You can let them Kind of be unsupervised for much longer periods of time.”
Mark ChenJun 7, 2025▶ 1:05:08Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
AssertionSupported
Chen: Frontier AI models consistently score over 90% on AIME
“I think one clear example here is the Amy, like probably the hardest auto gradable, like human math eval, at least in the US. And yeah, the models are consistently getting like 90 plus percent on these.”
Mark ChenJun 7, 2025▶ 1:08:50Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Insight
Chen: AI reasoning is the key mechanism to make agents reliable
“And I think the reason why we care so much about reasoning is because I think that's the path that we get reliable agents through.”
Mark ChenJun 7, 2025▶ 1:26:43Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
AssertionSupported
Chen: OpenAI reached 3 million paying business users
“We hit a big milestone. We got I think three million paying business users fairly recently.”
Mark ChenJun 7, 2025▶ 1:07:26Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Insight
Mark Chen: Compute Can Scale Heavily into RL Given Right Levers
“I think, like, if you find the right levers, you can really pump a lot of compute into RL as well as pre-training.”
Mark ChenJun 7, 2025▶ 1:28:16Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 500 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.