Feb 27, 2025 · 24m · big-technology

OpenAI's Chief Research Officer on GPT 4.5's Debut, Scaling Laws, And Teaching EQ to Models

Mark Chen · 13m spoken Alex Kantrowitz · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Big Technology Podcast episode, OpenAI Chief Research Officer Mark Chen discusses the release of GPT-4.5, defending foundational scaling laws, outlining the synergy between pre-training and reasoning architectures, and detailing improvements in serving efficiency and model emotional intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Alex holds 41.5% of the talking time here. How this is scored →

Alex as informed peer 5.3 Guest teaching 3.3 Guest disagreement 1.4 Alex pushing back 3.6
05100:0010:0020:000:00–3:30 · Alex as informed peer 4/10 Introducing GPT-4.5 and Chief Research Officer Mark Chen Kantrowitz pushes Chen on the naming convention and public expectations, asking why the release isn't called GPT-5 given the long gap between major releases. Chen calmly reframes the release within OpenAI's predictable scaling paradigm and explains their parallel research track in reasoning.3:31–7:09 · Alex as informed peer 5/10 Testing the Scaling Hypothesis and Complementary Reasoning Systems Kantrowitz asks whether LLMs are hitting a scaling wall given recent industry debates. Chen clarifies that unsupervised pre-training and reasoning models operate on complementary axes rather than competing paradigms.7:10–9:36 · Alex as informed peer 5/10 Training Run Dynamics and Frontier Model Scaling Projections Kantrowitz presses Chen on rumors regarding halted and restarted training runs for GPT-4.5. Chen pushes back on the rumor's framing, explaining that pausing and mid-run interventions are standard exploratory procedures across all foundation model runs.9:36–12:08 · Alex as informed peer 6/10 Inference Efficiency and Mixture of Experts Architecture Implementation Kantrowitz demonstrates technical familiarity with DeepSeek's Mixture of Experts (MoE) optimizations and queries OpenAI's efficiency methods. Chen explains that inference efficiency optimizations remain decoupled from pre-training core capabilities.12:09–14:15 · Alex as informed peer 5/10 Frontier Model Capabilities Versus Specialized Niche AI Architectures Kantrowitz cites his community Discord debates advocating for specialized niche models over general-purpose frontier models. Chen justifies OpenAI's dual approach of pushing top-tier frontier capabilities while offering mini models for cost efficiency.14:15–18:12 · Alex as informed peer 6/10 Enabling Autonomous AI Agents Through Enhanced Foundation Models Kantrowitz brings up an ongoing podcast debate regarding model intelligence versus product wrapper optimization. Chen sides with Kantrowitz's model-first stance, illustrating how agentic workflows like Deep Research rely directly on underlying model power.18:12–21:59 · Alex as informed peer 6/10 Evaluating Emotional Intelligence and Qualitative Human Interaction Benchmarks Kantrowitz anticipates potential skepticism that emphasizing emotional intelligence (EQ) is a goalpost shift away from traditional hard benchmarks. Chen defends the qualitative framing, insisting the model meets standard performance metrics while unlocking subtle interpersonal capabilities.21:59–23:55 · Alex as informed peer 5/10 OpenAI Talent Bench Dynamics and GPT-4.5 Release Schedule Kantrowitz asks directly about internal morale and recent high-profile executive departures at OpenAI. Chen characterizes the turnover as natural ecosystem churn while defending the strength of the remaining research bench.0:00–3:30 · Guest teaching 3/10 Introducing GPT-4.5 and Chief Research Officer Mark Chen Kantrowitz pushes Chen on the naming convention and public expectations, asking why the release isn't called GPT-5 given the long gap between major releases. Chen calmly reframes the release within OpenAI's predictable scaling paradigm and explains their parallel research track in reasoning.3:31–7:09 · Guest teaching 4/10 Testing the Scaling Hypothesis and Complementary Reasoning Systems Kantrowitz asks whether LLMs are hitting a scaling wall given recent industry debates. Chen clarifies that unsupervised pre-training and reasoning models operate on complementary axes rather than competing paradigms.7:10–9:36 · Guest teaching 4/10 Training Run Dynamics and Frontier Model Scaling Projections Kantrowitz presses Chen on rumors regarding halted and restarted training runs for GPT-4.5. Chen pushes back on the rumor's framing, explaining that pausing and mid-run interventions are standard exploratory procedures across all foundation model runs.9:36–12:08 · Guest teaching 3/10 Inference Efficiency and Mixture of Experts Architecture Implementation Kantrowitz demonstrates technical familiarity with DeepSeek's Mixture of Experts (MoE) optimizations and queries OpenAI's efficiency methods. Chen explains that inference efficiency optimizations remain decoupled from pre-training core capabilities.12:09–14:15 · Guest teaching 3/10 Frontier Model Capabilities Versus Specialized Niche AI Architectures Kantrowitz cites his community Discord debates advocating for specialized niche models over general-purpose frontier models. Chen justifies OpenAI's dual approach of pushing top-tier frontier capabilities while offering mini models for cost efficiency.14:15–18:12 · Guest teaching 3/10 Enabling Autonomous AI Agents Through Enhanced Foundation Models Kantrowitz brings up an ongoing podcast debate regarding model intelligence versus product wrapper optimization. Chen sides with Kantrowitz's model-first stance, illustrating how agentic workflows like Deep Research rely directly on underlying model power.18:12–21:59 · Guest teaching 4/10 Evaluating Emotional Intelligence and Qualitative Human Interaction Benchmarks Kantrowitz anticipates potential skepticism that emphasizing emotional intelligence (EQ) is a goalpost shift away from traditional hard benchmarks. Chen defends the qualitative framing, insisting the model meets standard performance metrics while unlocking subtle interpersonal capabilities.21:59–23:55 · Guest teaching 2/10 OpenAI Talent Bench Dynamics and GPT-4.5 Release Schedule Kantrowitz asks directly about internal morale and recent high-profile executive departures at OpenAI. Chen characterizes the turnover as natural ecosystem churn while defending the strength of the remaining research bench.0:00–3:30 · Guest disagreement 1/10 Introducing GPT-4.5 and Chief Research Officer Mark Chen Kantrowitz pushes Chen on the naming convention and public expectations, asking why the release isn't called GPT-5 given the long gap between major releases. Chen calmly reframes the release within OpenAI's predictable scaling paradigm and explains their parallel research track in reasoning.3:31–7:09 · Guest disagreement 2/10 Testing the Scaling Hypothesis and Complementary Reasoning Systems Kantrowitz asks whether LLMs are hitting a scaling wall given recent industry debates. Chen clarifies that unsupervised pre-training and reasoning models operate on complementary axes rather than competing paradigms.7:10–9:36 · Guest disagreement 2/10 Training Run Dynamics and Frontier Model Scaling Projections Kantrowitz presses Chen on rumors regarding halted and restarted training runs for GPT-4.5. Chen pushes back on the rumor's framing, explaining that pausing and mid-run interventions are standard exploratory procedures across all foundation model runs.9:36–12:08 · Guest disagreement 1/10 Inference Efficiency and Mixture of Experts Architecture Implementation Kantrowitz demonstrates technical familiarity with DeepSeek's Mixture of Experts (MoE) optimizations and queries OpenAI's efficiency methods. Chen explains that inference efficiency optimizations remain decoupled from pre-training core capabilities.12:09–14:15 · Guest disagreement 1/10 Frontier Model Capabilities Versus Specialized Niche AI Architectures Kantrowitz cites his community Discord debates advocating for specialized niche models over general-purpose frontier models. Chen justifies OpenAI's dual approach of pushing top-tier frontier capabilities while offering mini models for cost efficiency.14:15–18:12 · Guest disagreement 1/10 Enabling Autonomous AI Agents Through Enhanced Foundation Models Kantrowitz brings up an ongoing podcast debate regarding model intelligence versus product wrapper optimization. Chen sides with Kantrowitz's model-first stance, illustrating how agentic workflows like Deep Research rely directly on underlying model power.18:12–21:59 · Guest disagreement 2/10 Evaluating Emotional Intelligence and Qualitative Human Interaction Benchmarks Kantrowitz anticipates potential skepticism that emphasizing emotional intelligence (EQ) is a goalpost shift away from traditional hard benchmarks. Chen defends the qualitative framing, insisting the model meets standard performance metrics while unlocking subtle interpersonal capabilities.21:59–23:55 · Guest disagreement 1/10 OpenAI Talent Bench Dynamics and GPT-4.5 Release Schedule Kantrowitz asks directly about internal morale and recent high-profile executive departures at OpenAI. Chen characterizes the turnover as natural ecosystem churn while defending the strength of the remaining research bench.0:00–3:30 · Alex pushing back 4/10 Introducing GPT-4.5 and Chief Research Officer Mark Chen Kantrowitz pushes Chen on the naming convention and public expectations, asking why the release isn't called GPT-5 given the long gap between major releases. Chen calmly reframes the release within OpenAI's predictable scaling paradigm and explains their parallel research track in reasoning.3:31–7:09 · Alex pushing back 3/10 Testing the Scaling Hypothesis and Complementary Reasoning Systems Kantrowitz asks whether LLMs are hitting a scaling wall given recent industry debates. Chen clarifies that unsupervised pre-training and reasoning models operate on complementary axes rather than competing paradigms.7:10–9:36 · Alex pushing back 4/10 Training Run Dynamics and Frontier Model Scaling Projections Kantrowitz presses Chen on rumors regarding halted and restarted training runs for GPT-4.5. Chen pushes back on the rumor's framing, explaining that pausing and mid-run interventions are standard exploratory procedures across all foundation model runs.9:36–12:08 · Alex pushing back 3/10 Inference Efficiency and Mixture of Experts Architecture Implementation Kantrowitz demonstrates technical familiarity with DeepSeek's Mixture of Experts (MoE) optimizations and queries OpenAI's efficiency methods. Chen explains that inference efficiency optimizations remain decoupled from pre-training core capabilities.12:09–14:15 · Alex pushing back 4/10 Frontier Model Capabilities Versus Specialized Niche AI Architectures Kantrowitz cites his community Discord debates advocating for specialized niche models over general-purpose frontier models. Chen justifies OpenAI's dual approach of pushing top-tier frontier capabilities while offering mini models for cost efficiency.14:15–18:12 · Alex pushing back 3/10 Enabling Autonomous AI Agents Through Enhanced Foundation Models Kantrowitz brings up an ongoing podcast debate regarding model intelligence versus product wrapper optimization. Chen sides with Kantrowitz's model-first stance, illustrating how agentic workflows like Deep Research rely directly on underlying model power.18:12–21:59 · Alex pushing back 5/10 Evaluating Emotional Intelligence and Qualitative Human Interaction Benchmarks Kantrowitz anticipates potential skepticism that emphasizing emotional intelligence (EQ) is a goalpost shift away from traditional hard benchmarks. Chen defends the qualitative framing, insisting the model meets standard performance metrics while unlocking subtle interpersonal capabilities.21:59–23:55 · Alex pushing back 3/10 OpenAI Talent Bench Dynamics and GPT-4.5 Release Schedule Kantrowitz asks directly about internal morale and recent high-profile executive departures at OpenAI. Chen characterizes the turnover as natural ecosystem churn while defending the strength of the remaining research bench.

speaking balance: gold is Alex, purple is the guest (3 minute bins)

0:00 · Alex 55.9% · guest 44.1%0:00 · Alex 55.9% · guest 44.1%3:00 · Alex 33.6% · guest 66.4%3:00 · Alex 33.6% · guest 66.4%6:00 · Alex 44.3% · guest 55.7%6:00 · Alex 44.3% · guest 55.7%9:00 · Alex 44.8% · guest 55.2%9:00 · Alex 44.8% · guest 55.2%12:00 · Alex 44.3% · guest 55.7%12:00 · Alex 44.3% · guest 55.7%15:00 · Alex 25.1% · guest 74.9%15:00 · Alex 25.1% · guest 74.9%18:00 · Alex 42.3% · guest 57.7%18:00 · Alex 42.3% · guest 57.7%21:00 · Alex 42.3% · guest 57.7%21:00 · Alex 42.3% · guest 57.7%24:00 · Alex 0% · guest 0%24:00 · Alex 0% · guest 0%
Sharpest disagreement ▶ 8:51 Chen reframes paused training run reports

Chen firmly rejects the premise that starting and stopping training runs is unique to GPT-4.5 or indicates failure, asserting that intermediate interventions are standard across all OpenAI foundation models.

Hardest push from Alex ▶ 20:50 Kantrowitz challenges the focus on emotional intelligence

Kantrowitz directly confronts Chen with the criticism that OpenAI is moving the goalposts by spotlighting soft EQ vibes instead of traditional hard benchmark gains.

Biggest teaching moment ▶ 4:22 Chen explains reasoning and knowledge complementarity

Chen breaks down how massive unsupervised knowledge bases are a prerequisite for reasoning models, educating the audience and host on the symbiotic feedback loops between both paradigms.

Alex holds their own ▶ 9:36 Kantrowitz details MoE architecture and DeepSeek efficiency

Kantrowitz showcases deep technical domain knowledge by detailing how routing queries via Mixture of Experts reduces compute burdens compared to activating the full parameter space.

the scores for every segment, with the reasoning behind each
ChapterTopicAlex as informed peerGuest teachingGuest disagreementAlex pushing backWhy
Introducing GPT-4.5 and Chief Research Officer Mark Chen 4314 Kantrowitz pushes Chen on the naming convention and public expectations, asking why the release isn't called GPT-5 given the long gap between major releases. Chen calmly reframes the release within OpenAI's predictable scaling paradigm and explains their parallel research track in reasoning.
Testing the Scaling Hypothesis and Complementary Reasoning Systems 5423 Kantrowitz asks whether LLMs are hitting a scaling wall given recent industry debates. Chen clarifies that unsupervised pre-training and reasoning models operate on complementary axes rather than competing paradigms.
Training Run Dynamics and Frontier Model Scaling Projections 5424 Kantrowitz presses Chen on rumors regarding halted and restarted training runs for GPT-4.5. Chen pushes back on the rumor's framing, explaining that pausing and mid-run interventions are standard exploratory procedures across all foundation model runs.
Inference Efficiency and Mixture of Experts Architecture Implementation 6313 Kantrowitz demonstrates technical familiarity with DeepSeek's Mixture of Experts (MoE) optimizations and queries OpenAI's efficiency methods. Chen explains that inference efficiency optimizations remain decoupled from pre-training core capabilities.
Frontier Model Capabilities Versus Specialized Niche AI Architectures 5314 Kantrowitz cites his community Discord debates advocating for specialized niche models over general-purpose frontier models. Chen justifies OpenAI's dual approach of pushing top-tier frontier capabilities while offering mini models for cost efficiency.
Enabling Autonomous AI Agents Through Enhanced Foundation Models 6313 Kantrowitz brings up an ongoing podcast debate regarding model intelligence versus product wrapper optimization. Chen sides with Kantrowitz's model-first stance, illustrating how agentic workflows like Deep Research rely directly on underlying model power.
Evaluating Emotional Intelligence and Qualitative Human Interaction Benchmarks 6425 Kantrowitz anticipates potential skepticism that emphasizing emotional intelligence (EQ) is a goalpost shift away from traditional hard benchmarks. Chen defends the qualitative framing, insisting the model meets standard performance metrics while unlocking subtle interpersonal capabilities.
OpenAI Talent Bench Dynamics and GPT-4.5 Release Schedule 5213 Kantrowitz asks directly about internal morale and recent high-profile executive departures at OpenAI. Chen characterizes the turnover as natural ecosystem churn while defending the strength of the remaining research bench.

Statements from this episode (12)

Assertion Not checkable as stated
Chen: GPT-4.5 performance jump matches leap from GPT-3.5 to GPT-4
“It signifies an order of magnitude improvement over the last models, kind of commensurate with the jump from 3.5 to four.”
Mark Chen Feb 27, 2025 ▶ 1:03
Disclosure
Chen: Focus on reasoning models caused the longer gap before GPT-4.5
“Why there seems to be, you know, a little bit bigger of a gap in release time between four and 4.5, we've been really largely focused on developing the reasoning parallel paradigm as well.”
Mark Chen Feb 27, 2025 ▶ 2:54
Prediction Not checkable as stated
Chen: GPT-5 could combine unsupervised scaling with reasoning paradigms
“And so I think, like GPT-V really could be the culmination of a lot of these things coming together.”
Mark Chen Feb 27, 2025 ▶ 3:24
Insight
Chen: AI models cannot learn reasoning from scratch without pre-trained knowledge
“You need knowledge in order to build reasoning on top of it. Right. a model can't kind of go in blind and just learn reasoning from scratch. So we find these two paradigms to be fairly complementary and we think, you know, they have feedback loops on each oth…”
Mark Chen Feb 27, 2025 ▶ 4:45
Assertion Partly supported
Chen: Users prefer GPT-4.5 over GPT-4o by 60% to 70% margins
“When we look at, kind of, comparisons against GPT-FORO you'll see that everyday use cases, people prefer, you know, by a margin of 60% for actually productivity and knowledge work against GPT-FORO, there's almost like a 70% preference rate.”
Mark Chen Feb 27, 2025 ▶ 5:14
Opinion
Chen: GPT-4.5 outshines reasoning models like o1 in creative writing
“And, you know, we find that in a lot of areas like creative writing, for instance Again, this is stuff that we want to test over the next one or two months but we find that there are areas like creative writing where this model outshines reasoning models.”
Mark Chen Feb 27, 2025 ▶ 6:35
Assertion Not checkable as stated
Chen: GPT-4.5 scaling returns remain consistent with OpenAI's prior projections
“You know, we are seeing the same returns, and I do want to stress that GPT-D 4.5 is that next point on this unsupervised learning paradigm, and, you know, we're very rigorous about how we do this. We make projections based on all the models we've trained befor…”
Mark Chen Feb 27, 2025 ▶ 7:55
Assertion Not checkable as stated
Chen: Pausing and restarting training runs is standard across OpenAI models
“Actually, so I think it's interesting that this gets is a point that's attributed to this model because actually in, in, in developing all of our foundation models, right they're all experiments, right? I think you know, running all of the foundation models of…”
Mark Chen Feb 27, 2025 ▶ 8:51
Assertion Not checkable as stated
Chen: OpenAI inference costs dropped orders of magnitude since GPT-4
“The costs have dropped, you know, many orders of magnitude since we first launched GPT-IV.”
Mark Chen Feb 27, 2025 ▶ 11:21
Assertion Not checkable as stated
Chen: Nearly all large language models today utilize mixture of experts
“I think pretty much all large language models today use, utilize mixture of experts.”
Mark Chen Feb 27, 2025 ▶ 11:38
Assertion Not checkable as stated
Chen: GPT-4.5 creates ASCII art almost flawlessly, unlike previous models
“If you ask any of the previous models to create ASCII art for you, right? Actually, they mostly just fall down. This one can do it Almost flawless.”
Mark Chen Feb 27, 2025 ▶ 19:54
Assertion Not checkable as stated
Chen: GPT-4.5 hits expected benchmark progression consistent with OpenAI's trajectory
“Well, I really don't think that the accurate characterization is that it doesn't hit the benchmarks that, that we expect it to. So when you look at kind of the development of three to 3.5 to four to 4.5 this does hit the benchmarks that we expect.”
Mark Chen Feb 27, 2025 ▶ 21:08
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.