Sep 25, 2025 · 53m · a16z
From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
OpenAI leaders Jakub Pachocki and Mark Chen join The a16z Podcast to discuss the evolution of reasoning models, the future of AI-driven scientific discovery, and the strategic management behind frontier AI research.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Jakub explicitly rejects the host's framing that research conviction and truth-seeking exist in zero-sum tension, clarifying that deep belief in an idea can coexist with objective progress tracking.
Hardest push from the host ▶ 36:33 Host introduces Google's Nano Banana counter-exampleThe host directly challenges OpenAI's research priorities by raising Google's Nano Banana image model, questioning whether OpenAI risks ignoring valuable media generation breakthroughs.
Biggest teaching moment ▶ 2:44 Jakub reframes eval metrics and RL reasoningJakub educates the host on how pre-training benchmark evaluations have reached saturation, explaining how RL-driven domain reasoning fundamentally changes progress measurement.
The host holds their own ▶ 24:11 Host draws on bioinformatics grad school experienceThe host cites her own background as a bioinformatics researcher in graduate school to frame a sophisticated question about the psychological trap of going native on a research problem.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| GPT-V and Bringing Reasoning to the Mainstream | 3 | 4 | 1 | 1 | The host demonstrates familiarity with eval saturation percentages and AtCoder benchmark rankings. Jakub educates the host on how RL training in specific reasoning domains differs from traditional pre-training generalization. | |
| Lightbulb Moments and Daily Utility in Hard Sciences | 1 | 3 | 0 | 0 | The host prompts the guests for surprising lightbulb moments during internal testing. The guests explain how physicists and mathematicians experienced breakthrough utility in formula derivations. | |
| The Research Roadmap: Building the Automated Researcher | 4 | 3 | 1 | 2 | The host exhibits technical depth by questioning whether multi-step agentic tool use creates quality regressions compared to single-step execution. Jakub and Mark clarify that core reasoning capability provides the necessary stability across long horizons. | |
| Applying Reasoning to Open-Ended and Unverifiable Domains | 3 | 4 | 1 | 2 | The host brings up industry skepticism regarding RL plateaus, synthetic data mode collapse, and eval saturation. Jakub reframes the issue by detailing OpenAI's historical approach to RL environments and language pre-training integration. | |
| Reward Modeling and the Evolution of AI Learning | 2 | 3 | 0 | 0 | The host asks practical questions about enterprise reward modeling and latency presets for coding models. Jakub advises shifting mindsets toward simpler, more human-like learning paradigms. | |
| From Competitive Coding to Vibe Coding and Vibe Researching | 4 | 2 | 0 | 0 | The host demonstrates domain context by comparing AI coding adoption to Lee Sedol's retirement after losing to AlphaGo. The guests reflect collaboratively on personal coding shifts and the rise of high school vibe coding. | |
| Research Mindset, Problem Selection, and Overcoming Obstacles | 5 | 4 | 2 | 2 | The host leverages her bioinformatics research background to ask about the tension between research conviction and truth-seeking. Jakub politely rejects the premise, asserting that conviction and truth-seeking are not in zero-sum tension. | |
| Recruiting, Retention, and Fostering Research Culture | 3 | 3 | 1 | 1 | The host introduces external commentary by citing Elon Musk's tweet on the researcher versus engineer distinction. Mark nuances the topic by outlining distinct research archetypes, such as ideators versus rigorous experimenters. | |
| Aligning Research Strategy with Product Vision | 2 | 2 | 0 | 1 | The host explores how leadership protects fundamental research while integrating top product executives. Mark and Jakub outline structural mandate clarity and shared company-wide alignment. | |
| Compute Allocation and Portfolio Prioritization | 4 | 3 | 1 | 2 | The host challenges research priorities by citing Google's viral Nano Banana image model as a potential distraction from pure reasoning. Mark responds by emphasizing strict compute portfolio management and clear focus on winning core bets. | |
| Academia, Frontier AI, and the OpenAI Residency | 3 | 2 | 0 | 0 | The host contrasts the historical role of university research labs with modern corporate frontier AI orgs. Mark discusses the OpenAI Residency program designed to accelerate PhD-level AI intuition. | |
| Navigating External Perception and Long-Term Research Conviction | 2 | 3 | 1 | 1 | The host asks whether short-term external product reception influences long-term research roadmaps. Jakub explains that core research operates from strong internal conviction rather than external perception cycles. |