Jun 25, 2026 · 41m · latent-space

Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen

Mark Chen · 22m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space cooking series, OpenAI Chief Research Officer Mark Chen joins host Alan to cook Korean tofu stew and flambéed prawns while discussing the genesis of the o1 reasoning model, scaling laws, benchmark evaluation crises, and the path to autonomous AGI research.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 3.6 Guest teaching 5.1 Guest disagreement 1.2 The hosts pushing back 1.1
05100:0015:0030:001:33–5:21 · The hosts as informed peer 3/10 From Trading to AI: Developing Research Taste Alan prompts Mark on how trading experience translates to AI research and how non-PhDs build research taste. Mark breaks down how replication of landmark papers like ResNet and PixelCNN is the practical key to building research instincts rather than relying on pedigree.5:21–8:12 · The hosts as informed peer 4/10 RL Horizons and Evaluating Superhuman Capabilities Alan asks about reinforcement learning boundaries across subjective domains and evaluating superhuman benchmarks beyond IMO. Mark explains that RL struggles where grading is subjective (like creative writing) versus objective ground truths (like math and coding).8:13–11:12 · The hosts as informed peer 3/10 Scaling Laws and the Genesis of OpenAI o1 Alan asks Mark about common contrarian takes like 'pre-training is dead.' Mark firmly rejects bearish views on scaling laws, explaining how recurring historical bottlenecks were consistently overcome and describing the early internal conviction needed to spawn o1.11:12–13:47 · The hosts as informed peer 3/10 Research Leadership and Roadmap Prioritization Alan inquires about OpenAI's unchanged research roadmap and how leadership balances top-down steering with bottom-up researcher ideas. Mark describes OpenAI's meritocratic management culture and how compute allocation checkpoints force periodic re-evaluations.13:47–18:54 · The hosts as informed peer 4/10 Compute Allocation and Identifying Top Research Talent Alan asks how OpenAI sifts through hundreds of research proposals and identifies standout talent. Mark explains directive compute allocation where managers receive dedicated large pools alongside flexible discretion to support diverse researcher profiles.18:55–23:39 · The hosts as informed peer 4/10 Navigating the Evals Crisis and Adversarial Benchmarking Alan raises the discrepancy between benchmark scores and vibe checks. Mark details the industry's evals crisis due to benchmark saturation and advocates for strictly separating eval design teams from model optimization teams to maintain an adversarial dynamic.23:39–27:12 · The hosts as informed peer 4/10 Jakub Pachocki Dynamics, Jagged Frontiers, and Context Alan pushes on why models excel at complex Olympiad problems but struggle with basic human tasks, citing context bloat and context rot. Mark discusses the jagged frontier of intelligence and contrasts naive context window expansion with state compaction techniques.27:12–31:20 · The hosts as informed peer 3/10 Flambéing Prawns, AGI Horizons, and Autonomous Research Alan asks whether AGI requires multiple drastic paradigm breakthroughs such as continual learning. Mark gently pushes back on that framing, arguing that continual learning is an approachable primitive with multiple viable shots on goal rather than an insurmountable barrier.31:20–37:07 · The hosts as informed peer 4/10 Multimodal Architectures, Vibe Researchers, and High-Risk Bets Alan probes unified multimodal architectures, vibe researching, and managing researchers whose bets repeatedly fail. Mark details why shared infrastructure stacks are favored and how a high-risk portfolio philosophy accommodates stringed failures before a breakthrough.1:33–5:21 · Guest teaching 5/10 From Trading to AI: Developing Research Taste Alan prompts Mark on how trading experience translates to AI research and how non-PhDs build research taste. Mark breaks down how replication of landmark papers like ResNet and PixelCNN is the practical key to building research instincts rather than relying on pedigree.5:21–8:12 · Guest teaching 5/10 RL Horizons and Evaluating Superhuman Capabilities Alan asks about reinforcement learning boundaries across subjective domains and evaluating superhuman benchmarks beyond IMO. Mark explains that RL struggles where grading is subjective (like creative writing) versus objective ground truths (like math and coding).8:13–11:12 · Guest teaching 6/10 Scaling Laws and the Genesis of OpenAI o1 Alan asks Mark about common contrarian takes like 'pre-training is dead.' Mark firmly rejects bearish views on scaling laws, explaining how recurring historical bottlenecks were consistently overcome and describing the early internal conviction needed to spawn o1.11:12–13:47 · Guest teaching 4/10 Research Leadership and Roadmap Prioritization Alan inquires about OpenAI's unchanged research roadmap and how leadership balances top-down steering with bottom-up researcher ideas. Mark describes OpenAI's meritocratic management culture and how compute allocation checkpoints force periodic re-evaluations.13:47–18:54 · Guest teaching 5/10 Compute Allocation and Identifying Top Research Talent Alan asks how OpenAI sifts through hundreds of research proposals and identifies standout talent. Mark explains directive compute allocation where managers receive dedicated large pools alongside flexible discretion to support diverse researcher profiles.18:55–23:39 · Guest teaching 6/10 Navigating the Evals Crisis and Adversarial Benchmarking Alan raises the discrepancy between benchmark scores and vibe checks. Mark details the industry's evals crisis due to benchmark saturation and advocates for strictly separating eval design teams from model optimization teams to maintain an adversarial dynamic.23:39–27:12 · Guest teaching 5/10 Jakub Pachocki Dynamics, Jagged Frontiers, and Context Alan pushes on why models excel at complex Olympiad problems but struggle with basic human tasks, citing context bloat and context rot. Mark discusses the jagged frontier of intelligence and contrasts naive context window expansion with state compaction techniques.27:12–31:20 · Guest teaching 5/10 Flambéing Prawns, AGI Horizons, and Autonomous Research Alan asks whether AGI requires multiple drastic paradigm breakthroughs such as continual learning. Mark gently pushes back on that framing, arguing that continual learning is an approachable primitive with multiple viable shots on goal rather than an insurmountable barrier.31:20–37:07 · Guest teaching 5/10 Multimodal Architectures, Vibe Researchers, and High-Risk Bets Alan probes unified multimodal architectures, vibe researching, and managing researchers whose bets repeatedly fail. Mark details why shared infrastructure stacks are favored and how a high-risk portfolio philosophy accommodates stringed failures before a breakthrough.1:33–5:21 · Guest disagreement 1/10 From Trading to AI: Developing Research Taste Alan prompts Mark on how trading experience translates to AI research and how non-PhDs build research taste. Mark breaks down how replication of landmark papers like ResNet and PixelCNN is the practical key to building research instincts rather than relying on pedigree.5:21–8:12 · Guest disagreement 1/10 RL Horizons and Evaluating Superhuman Capabilities Alan asks about reinforcement learning boundaries across subjective domains and evaluating superhuman benchmarks beyond IMO. Mark explains that RL struggles where grading is subjective (like creative writing) versus objective ground truths (like math and coding).8:13–11:12 · Guest disagreement 2/10 Scaling Laws and the Genesis of OpenAI o1 Alan asks Mark about common contrarian takes like 'pre-training is dead.' Mark firmly rejects bearish views on scaling laws, explaining how recurring historical bottlenecks were consistently overcome and describing the early internal conviction needed to spawn o1.11:12–13:47 · Guest disagreement 1/10 Research Leadership and Roadmap Prioritization Alan inquires about OpenAI's unchanged research roadmap and how leadership balances top-down steering with bottom-up researcher ideas. Mark describes OpenAI's meritocratic management culture and how compute allocation checkpoints force periodic re-evaluations.13:47–18:54 · Guest disagreement 1/10 Compute Allocation and Identifying Top Research Talent Alan asks how OpenAI sifts through hundreds of research proposals and identifies standout talent. Mark explains directive compute allocation where managers receive dedicated large pools alongside flexible discretion to support diverse researcher profiles.18:55–23:39 · Guest disagreement 1/10 Navigating the Evals Crisis and Adversarial Benchmarking Alan raises the discrepancy between benchmark scores and vibe checks. Mark details the industry's evals crisis due to benchmark saturation and advocates for strictly separating eval design teams from model optimization teams to maintain an adversarial dynamic.23:39–27:12 · Guest disagreement 1/10 Jakub Pachocki Dynamics, Jagged Frontiers, and Context Alan pushes on why models excel at complex Olympiad problems but struggle with basic human tasks, citing context bloat and context rot. Mark discusses the jagged frontier of intelligence and contrasts naive context window expansion with state compaction techniques.27:12–31:20 · Guest disagreement 2/10 Flambéing Prawns, AGI Horizons, and Autonomous Research Alan asks whether AGI requires multiple drastic paradigm breakthroughs such as continual learning. Mark gently pushes back on that framing, arguing that continual learning is an approachable primitive with multiple viable shots on goal rather than an insurmountable barrier.31:20–37:07 · Guest disagreement 1/10 Multimodal Architectures, Vibe Researchers, and High-Risk Bets Alan probes unified multimodal architectures, vibe researching, and managing researchers whose bets repeatedly fail. Mark details why shared infrastructure stacks are favored and how a high-risk portfolio philosophy accommodates stringed failures before a breakthrough.1:33–5:21 · The hosts pushing back 1/10 From Trading to AI: Developing Research Taste Alan prompts Mark on how trading experience translates to AI research and how non-PhDs build research taste. Mark breaks down how replication of landmark papers like ResNet and PixelCNN is the practical key to building research instincts rather than relying on pedigree.5:21–8:12 · The hosts pushing back 1/10 RL Horizons and Evaluating Superhuman Capabilities Alan asks about reinforcement learning boundaries across subjective domains and evaluating superhuman benchmarks beyond IMO. Mark explains that RL struggles where grading is subjective (like creative writing) versus objective ground truths (like math and coding).8:13–11:12 · The hosts pushing back 1/10 Scaling Laws and the Genesis of OpenAI o1 Alan asks Mark about common contrarian takes like 'pre-training is dead.' Mark firmly rejects bearish views on scaling laws, explaining how recurring historical bottlenecks were consistently overcome and describing the early internal conviction needed to spawn o1.11:12–13:47 · The hosts pushing back 1/10 Research Leadership and Roadmap Prioritization Alan inquires about OpenAI's unchanged research roadmap and how leadership balances top-down steering with bottom-up researcher ideas. Mark describes OpenAI's meritocratic management culture and how compute allocation checkpoints force periodic re-evaluations.13:47–18:54 · The hosts pushing back 1/10 Compute Allocation and Identifying Top Research Talent Alan asks how OpenAI sifts through hundreds of research proposals and identifies standout talent. Mark explains directive compute allocation where managers receive dedicated large pools alongside flexible discretion to support diverse researcher profiles.18:55–23:39 · The hosts pushing back 1/10 Navigating the Evals Crisis and Adversarial Benchmarking Alan raises the discrepancy between benchmark scores and vibe checks. Mark details the industry's evals crisis due to benchmark saturation and advocates for strictly separating eval design teams from model optimization teams to maintain an adversarial dynamic.23:39–27:12 · The hosts pushing back 2/10 Jakub Pachocki Dynamics, Jagged Frontiers, and Context Alan pushes on why models excel at complex Olympiad problems but struggle with basic human tasks, citing context bloat and context rot. Mark discusses the jagged frontier of intelligence and contrasts naive context window expansion with state compaction techniques.27:12–31:20 · The hosts pushing back 1/10 Flambéing Prawns, AGI Horizons, and Autonomous Research Alan asks whether AGI requires multiple drastic paradigm breakthroughs such as continual learning. Mark gently pushes back on that framing, arguing that continual learning is an approachable primitive with multiple viable shots on goal rather than an insurmountable barrier.31:20–37:07 · The hosts pushing back 1/10 Multimodal Architectures, Vibe Researchers, and High-Risk Bets Alan probes unified multimodal architectures, vibe researching, and managing researchers whose bets repeatedly fail. Mark details why shared infrastructure stacks are favored and how a high-risk portfolio philosophy accommodates stringed failures before a breakthrough.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 9:13 Firm rejection of scaling law skepticism

Mark strongly disagrees with prevailing bearish takes on pre-training, noting that bottleneck claims have consistently proven false across ten orders of magnitude.

Hardest push from the hosts ▶ 25:54 Challenging naive context expansion

Alan challenges the simplistic low-hanging fruit solution of merely increasing context windows, arguing it introduces context bloat and rot.

Biggest teaching moment ▶ 22:15 Architecting adversarial eval teams

Mark explains the structural necessity of separating evaluation teams from model training teams to eliminate perverse incentives and benchmark gaming.

The host holds their own ▶ 17:42 Contrasting engineering vs research lifecycle

Alan articulates the distinct transition paths of early ideas to end-user products between production engineers and frontier researchers.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
From Trading to AI: Developing Research Taste 3511 Alan prompts Mark on how trading experience translates to AI research and how non-PhDs build research taste. Mark breaks down how replication of landmark papers like ResNet and PixelCNN is the practical key to building research instincts rather than relying on pedigree.
RL Horizons and Evaluating Superhuman Capabilities 4511 Alan asks about reinforcement learning boundaries across subjective domains and evaluating superhuman benchmarks beyond IMO. Mark explains that RL struggles where grading is subjective (like creative writing) versus objective ground truths (like math and coding).
Scaling Laws and the Genesis of OpenAI o1 3621 Alan asks Mark about common contrarian takes like 'pre-training is dead.' Mark firmly rejects bearish views on scaling laws, explaining how recurring historical bottlenecks were consistently overcome and describing the early internal conviction needed to spawn o1.
Research Leadership and Roadmap Prioritization 3411 Alan inquires about OpenAI's unchanged research roadmap and how leadership balances top-down steering with bottom-up researcher ideas. Mark describes OpenAI's meritocratic management culture and how compute allocation checkpoints force periodic re-evaluations.
Compute Allocation and Identifying Top Research Talent 4511 Alan asks how OpenAI sifts through hundreds of research proposals and identifies standout talent. Mark explains directive compute allocation where managers receive dedicated large pools alongside flexible discretion to support diverse researcher profiles.
Navigating the Evals Crisis and Adversarial Benchmarking 4611 Alan raises the discrepancy between benchmark scores and vibe checks. Mark details the industry's evals crisis due to benchmark saturation and advocates for strictly separating eval design teams from model optimization teams to maintain an adversarial dynamic.
Jakub Pachocki Dynamics, Jagged Frontiers, and Context 4512 Alan pushes on why models excel at complex Olympiad problems but struggle with basic human tasks, citing context bloat and context rot. Mark discusses the jagged frontier of intelligence and contrasts naive context window expansion with state compaction techniques.
Flambéing Prawns, AGI Horizons, and Autonomous Research 3521 Alan asks whether AGI requires multiple drastic paradigm breakthroughs such as continual learning. Mark gently pushes back on that framing, arguing that continual learning is an approachable primitive with multiple viable shots on goal rather than an insurmountable barrier.
Multimodal Architectures, Vibe Researchers, and High-Risk Bets 4511 Alan probes unified multimodal architectures, vibe researching, and managing researchers whose bets repeatedly fail. Mark details why shared infrastructure stacks are favored and how a high-risk portfolio philosophy accommodates stringed failures before a breakthrough.

Statements from this episode (16)

Disclosure
Mark Chen brought soup to researchers to counter Zuckerberg's talent poaching
“Oh, you know, it's absolutely a true story. And I have brought soup to our own researchers.”
Mark Chen Jun 25, 2026 ▶ 0:39
Opinion
Chen: Meta poaching has calmed down and OpenAI came out on top
“I think that met us calmed down a little bit. I think we came out on top”
Mark Chen Jun 25, 2026 ▶ 0:43
Insight
Mark Chen: A PhD is not necessary to excel in AI research
“There are a lot of researchers who just started out without formal training in machine learning or AI research. We've very much believed in training people up to do this. I think the real hard thing is the ability to creatively solve problems and think outside…”
Mark Chen Jun 25, 2026 ▶ 2:14
Insight
Mark Chen: Paper replication is the best way to develop AI research taste
“The best mechanism I've found for developing that is really just replication. So I think you should take papers that you really look up to, and just try to fully replicate it.”
Mark Chen Jun 25, 2026 ▶ 3:36
Opinion
Mark Chen: AI models are producing 'Move 37' breakthroughs in math and coding
“There's move-thirty-sevens in, in math. There's in computer science and coding. I think even, yeah, just, it feels like a lot of people woke up at the start of this year and were like, man, agents are working in my profession. And you know, they're essentially…”
Mark Chen Jun 25, 2026 ▶ 4:56
Insight
OpenAI's Chen: Reinforcement learning struggles in subjective, hard-to-grade fields
“RLs traditionally had headwinds when it's come to fields that, you know, it's more kind of, Subjective than objective. So if you kind of think of, you know, one kind of, you know example of this is creative writing, where, you know, you could take two pieces o…”
Mark Chen Jun 25, 2026 ▶ 5:54
Assertion Not checkable as stated
OpenAI's Chen: AI models already discover novel theorems and advance sciences
“The initial direction we took was you should move it to real world research, right? And we've seen that the models, they've gotten a lot better at just kind of discovering novel theorems and pushing the frontiers of hard sciences. Even today, right, that's no …”
Mark Chen Jun 25, 2026 ▶ 7:27
Prediction Not checkable as stated
Mark Chen: AI scaling laws will continue to hold
“And so I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering, more careful scaling, and it always unlocks that next ability to scale further. So I mean, it's held for You know, almost 10 order…”
Mark Chen Jun 25, 2026 ▶ 9:58
Assertion Not checkable as stated
Jakob Pachocki and Ilya Sutskever overcame internal inertia to build OpenAI o1
“Even at a company like OpenAI, you would have people ask naturally, why do something when you have a machine that works? And fundamentally, you know, it's to the credit of, you know, Jakob, Ilya, many of the people who really had conviction and vision in this …”
Mark Chen Jun 25, 2026 ▶ 10:46
Disclosure
OpenAI's three research pillars are pre-training, RL, and alignment
“At the very highest level, right, we have an org that focuses on pre-training, right, which is, you know, giving models a lot of world knowledge. We focus on RL, like, teaching the models how to reason with that knowledge, how to chain the little insights toge…”
Mark Chen Jun 25, 2026 ▶ 14:03
Assertion Not checkable as stated
OpenAI's Mark Chen: AI is in an evals crisis with saturated benchmarks
“So I think beyond that, the other scary thing in the field is the number of canonical gold standard benchmarks is low. And we really are kind of in an evals crisis, right? Where all the really great evals that we all know, like growing up, like taking the SAT …”
Mark Chen Jun 25, 2026 ▶ 20:23
Insight
Chen: Eval teams and model optimization teams must remain separate
“Yeah, I think there's a kind of interesting philosophy of separate the teams that are creating the evals from the teams that are optimizing the models themselves, because that way you don't, like, co-incentivize them, right? Like, the way the evals theme can w…”
Mark Chen Jun 25, 2026 ▶ 22:34
Insight
Chen: Compaction shortcuts expensive native long context in coding products
“Many, many coding products today have features like compaction, right? Where you can compress kind of either insights or working state and stuff like that, you know, it just shortcuts a lot of the very brutally difficult and expensive permit is that you have t…”
Mark Chen Jun 25, 2026 ▶ 26:51
Disclosure
Chen: OpenAI favors unifying modalities in as few architectures as possible
“For a research lab, I think there are a lot of advantages for it to being under one. So you just have to maintain one infrastructure stack, for instance. I think the cost to, like, maintaining and scaling many infrastructure stacks at once I think that's somet…”
Mark Chen Jun 25, 2026 ▶ 31:58
Disclosure
Chen: OpenAI's three-year goal is models conducting end-to-end research
“When we look at our kind of three-year roadmap, right the end goal that we want to reach is one where You know, the models are just doing end-to-end research, and I think a part of that problem is just being able to have the model come up with good taste.”
Mark Chen Jun 25, 2026 ▶ 34:15
Opinion
Chen: Pre-training is not dead and remains underrated in AI research
“Well, I think if you still have a pre-training is dead view of the world I think pre-training is definitely yeah, yeah, not, not dead. It's underrated.”
Mark Chen Jun 25, 2026 ▶ 38:01
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.