Jun 25, 2026 · 41m · latent-space
Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Latent Space cooking series, OpenAI Chief Research Officer Mark Chen joins host Alan to cook Korean tofu stew and flambéed prawns while discussing the genesis of the o1 reasoning model, scaling laws, benchmark evaluation crises, and the path to autonomous AGI research.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Mark strongly disagrees with prevailing bearish takes on pre-training, noting that bottleneck claims have consistently proven false across ten orders of magnitude.
Hardest push from the hosts ▶ 25:54 Challenging naive context expansionAlan challenges the simplistic low-hanging fruit solution of merely increasing context windows, arguing it introduces context bloat and rot.
Biggest teaching moment ▶ 22:15 Architecting adversarial eval teamsMark explains the structural necessity of separating evaluation teams from model training teams to eliminate perverse incentives and benchmark gaming.
The host holds their own ▶ 17:42 Contrasting engineering vs research lifecycleAlan articulates the distinct transition paths of early ideas to end-user products between production engineers and frontier researchers.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| From Trading to AI: Developing Research Taste | 3 | 5 | 1 | 1 | Alan prompts Mark on how trading experience translates to AI research and how non-PhDs build research taste. Mark breaks down how replication of landmark papers like ResNet and PixelCNN is the practical key to building research instincts rather than relying on pedigree. | |
| RL Horizons and Evaluating Superhuman Capabilities | 4 | 5 | 1 | 1 | Alan asks about reinforcement learning boundaries across subjective domains and evaluating superhuman benchmarks beyond IMO. Mark explains that RL struggles where grading is subjective (like creative writing) versus objective ground truths (like math and coding). | |
| Scaling Laws and the Genesis of OpenAI o1 | 3 | 6 | 2 | 1 | Alan asks Mark about common contrarian takes like 'pre-training is dead.' Mark firmly rejects bearish views on scaling laws, explaining how recurring historical bottlenecks were consistently overcome and describing the early internal conviction needed to spawn o1. | |
| Research Leadership and Roadmap Prioritization | 3 | 4 | 1 | 1 | Alan inquires about OpenAI's unchanged research roadmap and how leadership balances top-down steering with bottom-up researcher ideas. Mark describes OpenAI's meritocratic management culture and how compute allocation checkpoints force periodic re-evaluations. | |
| Compute Allocation and Identifying Top Research Talent | 4 | 5 | 1 | 1 | Alan asks how OpenAI sifts through hundreds of research proposals and identifies standout talent. Mark explains directive compute allocation where managers receive dedicated large pools alongside flexible discretion to support diverse researcher profiles. | |
| Navigating the Evals Crisis and Adversarial Benchmarking | 4 | 6 | 1 | 1 | Alan raises the discrepancy between benchmark scores and vibe checks. Mark details the industry's evals crisis due to benchmark saturation and advocates for strictly separating eval design teams from model optimization teams to maintain an adversarial dynamic. | |
| Jakub Pachocki Dynamics, Jagged Frontiers, and Context | 4 | 5 | 1 | 2 | Alan pushes on why models excel at complex Olympiad problems but struggle with basic human tasks, citing context bloat and context rot. Mark discusses the jagged frontier of intelligence and contrasts naive context window expansion with state compaction techniques. | |
| Flambéing Prawns, AGI Horizons, and Autonomous Research | 3 | 5 | 2 | 1 | Alan asks whether AGI requires multiple drastic paradigm breakthroughs such as continual learning. Mark gently pushes back on that framing, arguing that continual learning is an approachable primitive with multiple viable shots on goal rather than an insurmountable barrier. | |
| Multimodal Architectures, Vibe Researchers, and High-Risk Bets | 4 | 5 | 1 | 1 | Alan probes unified multimodal architectures, vibe researching, and managing researchers whose bets repeatedly fail. Mark details why shared infrastructure stacks are favored and how a high-risk portfolio philosophy accommodates stringed failures before a breakthrough. |