Nov 14, 2024 · 35m · y-combinator

Why The Next AI Breakthroughs Will Be In Reasoning, Not Scaling · Y Combinator

Garry Tan · 10m spoken Diana Hu · 8m spoken Harj Taggar · 6m spoken Jared Friedman · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Y Combinator managing partners analyze the fundamental shift in AI development from parameter scaling to step-by-step inference reasoning. Through live startup demos and technical breakdowns of OpenAI's o1 model, they demonstrate how reasoning models are unlocking physical-world engineering, higher enterprise precision, and new defensive moats for startups.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The partners hold 63.6% of the talking time here. How this is scored →

The partners as informed peer 7.3 Guest teaching 0.3 Guest disagreement 0.0 The partners pushing back 0.0
05100:0010:0020:0030:001:34–4:31 · The partners as informed peer 6/10 OpenAI's Origins and Sam Altman's Techno-Optimistic Vision Jared and Gary provide historical perspective on OpenAI's inception at Y Combinator, framing Sam Altman's long-term focus on reasoning and scientific acceleration. Harj introduces the emerging capability of AI in chip design.4:31–8:22 · The partners as informed peer 8/10 Startup Demo 1: Diode Computer Automating Circuit and PCB Design Diana provides detailed domain context on PCB architecture, NP-complete routing, and component selection limits under GPT-4. The Diode Computer founder demonstrates automated board layout and schematic generation using o1 reasoning.8:22–10:25 · The partners as informed peer 7/10 Multi-Model AI Workflows and Real Capability Unlocks Gary and Diana analyze multi-model pipeline architecture, explaining why extracting PDF datasheets with GPT-4o mini and reasoning with o1 works synergistically. Jared notes how this unlocks real capabilities for funded startups.10:25–14:26 · The partners as informed peer 7/10 Startup Demo 2: Camfer Automating 3D CAD Design in SolidWorks The partners review Camfer's natural language CAD workflow and its ability to solve Navier-Stokes equations for airfoils. Harj and Gary explore how shifting compute to the inference stage mirrors human scientific iteration.14:26–16:53 · The partners as informed peer 8/10 Reinforcement Learning Evolution: From Dota 2 to OpenAI o1 Diana and Jared trace OpenAI's technical heritage from Dota 2 self-play and Q-learning through to reward modeling in o1. Diana explains why factual verification domains like math and science benefit most from reinforcement learning.16:53–21:35 · The partners as informed peer 8/10 Dual Vectors of AI Progress: Base Scaling vs. Inference Reasoning Jared explains the divergence between pre-training model scale and test-time RL compute. Gary outlines how proprietary eval datasets from unindexed enterprise workflows form the primary defensive moat for vertical AI startups.21:35–31:50 · The partners as informed peer 7/10 High-Precision AI and Case Study: GigaML with Zepto Harj and Gary dissect GigaML's business evolution from model fine-tuning to automated customer support at Zepto. Diana reports on GigaML's performance jump from a 70 percent error rate to 5 percent by pairing o1 with structured evals.31:50–35:03 · The partners as informed peer 7/10 Deprecated Startup Ideas vs. Physical World Breakthroughs Harj and Gary evaluate the risks facing coding agent startups that rely on proprietary prompt wrappers without directability. Diana points to physical-world engineering disciplines as the highest-upside beneficiaries of reasoning models.1:34–4:31 · Guest teaching 0/10 OpenAI's Origins and Sam Altman's Techno-Optimistic Vision Jared and Gary provide historical perspective on OpenAI's inception at Y Combinator, framing Sam Altman's long-term focus on reasoning and scientific acceleration. Harj introduces the emerging capability of AI in chip design.4:31–8:22 · Guest teaching 2/10 Startup Demo 1: Diode Computer Automating Circuit and PCB Design Diana provides detailed domain context on PCB architecture, NP-complete routing, and component selection limits under GPT-4. The Diode Computer founder demonstrates automated board layout and schematic generation using o1 reasoning.8:22–10:25 · Guest teaching 0/10 Multi-Model AI Workflows and Real Capability Unlocks Gary and Diana analyze multi-model pipeline architecture, explaining why extracting PDF datasheets with GPT-4o mini and reasoning with o1 works synergistically. Jared notes how this unlocks real capabilities for funded startups.10:25–14:26 · Guest teaching 0/10 Startup Demo 2: Camfer Automating 3D CAD Design in SolidWorks The partners review Camfer's natural language CAD workflow and its ability to solve Navier-Stokes equations for airfoils. Harj and Gary explore how shifting compute to the inference stage mirrors human scientific iteration.14:26–16:53 · Guest teaching 0/10 Reinforcement Learning Evolution: From Dota 2 to OpenAI o1 Diana and Jared trace OpenAI's technical heritage from Dota 2 self-play and Q-learning through to reward modeling in o1. Diana explains why factual verification domains like math and science benefit most from reinforcement learning.16:53–21:35 · Guest teaching 0/10 Dual Vectors of AI Progress: Base Scaling vs. Inference Reasoning Jared explains the divergence between pre-training model scale and test-time RL compute. Gary outlines how proprietary eval datasets from unindexed enterprise workflows form the primary defensive moat for vertical AI startups.21:35–31:50 · Guest teaching 0/10 High-Precision AI and Case Study: GigaML with Zepto Harj and Gary dissect GigaML's business evolution from model fine-tuning to automated customer support at Zepto. Diana reports on GigaML's performance jump from a 70 percent error rate to 5 percent by pairing o1 with structured evals.31:50–35:03 · Guest teaching 0/10 Deprecated Startup Ideas vs. Physical World Breakthroughs Harj and Gary evaluate the risks facing coding agent startups that rely on proprietary prompt wrappers without directability. Diana points to physical-world engineering disciplines as the highest-upside beneficiaries of reasoning models.1:34–4:31 · Guest disagreement 0/10 OpenAI's Origins and Sam Altman's Techno-Optimistic Vision Jared and Gary provide historical perspective on OpenAI's inception at Y Combinator, framing Sam Altman's long-term focus on reasoning and scientific acceleration. Harj introduces the emerging capability of AI in chip design.4:31–8:22 · Guest disagreement 0/10 Startup Demo 1: Diode Computer Automating Circuit and PCB Design Diana provides detailed domain context on PCB architecture, NP-complete routing, and component selection limits under GPT-4. The Diode Computer founder demonstrates automated board layout and schematic generation using o1 reasoning.8:22–10:25 · Guest disagreement 0/10 Multi-Model AI Workflows and Real Capability Unlocks Gary and Diana analyze multi-model pipeline architecture, explaining why extracting PDF datasheets with GPT-4o mini and reasoning with o1 works synergistically. Jared notes how this unlocks real capabilities for funded startups.10:25–14:26 · Guest disagreement 0/10 Startup Demo 2: Camfer Automating 3D CAD Design in SolidWorks The partners review Camfer's natural language CAD workflow and its ability to solve Navier-Stokes equations for airfoils. Harj and Gary explore how shifting compute to the inference stage mirrors human scientific iteration.14:26–16:53 · Guest disagreement 0/10 Reinforcement Learning Evolution: From Dota 2 to OpenAI o1 Diana and Jared trace OpenAI's technical heritage from Dota 2 self-play and Q-learning through to reward modeling in o1. Diana explains why factual verification domains like math and science benefit most from reinforcement learning.16:53–21:35 · Guest disagreement 0/10 Dual Vectors of AI Progress: Base Scaling vs. Inference Reasoning Jared explains the divergence between pre-training model scale and test-time RL compute. Gary outlines how proprietary eval datasets from unindexed enterprise workflows form the primary defensive moat for vertical AI startups.21:35–31:50 · Guest disagreement 0/10 High-Precision AI and Case Study: GigaML with Zepto Harj and Gary dissect GigaML's business evolution from model fine-tuning to automated customer support at Zepto. Diana reports on GigaML's performance jump from a 70 percent error rate to 5 percent by pairing o1 with structured evals.31:50–35:03 · Guest disagreement 0/10 Deprecated Startup Ideas vs. Physical World Breakthroughs Harj and Gary evaluate the risks facing coding agent startups that rely on proprietary prompt wrappers without directability. Diana points to physical-world engineering disciplines as the highest-upside beneficiaries of reasoning models.1:34–4:31 · The partners pushing back 0/10 OpenAI's Origins and Sam Altman's Techno-Optimistic Vision Jared and Gary provide historical perspective on OpenAI's inception at Y Combinator, framing Sam Altman's long-term focus on reasoning and scientific acceleration. Harj introduces the emerging capability of AI in chip design.4:31–8:22 · The partners pushing back 0/10 Startup Demo 1: Diode Computer Automating Circuit and PCB Design Diana provides detailed domain context on PCB architecture, NP-complete routing, and component selection limits under GPT-4. The Diode Computer founder demonstrates automated board layout and schematic generation using o1 reasoning.8:22–10:25 · The partners pushing back 0/10 Multi-Model AI Workflows and Real Capability Unlocks Gary and Diana analyze multi-model pipeline architecture, explaining why extracting PDF datasheets with GPT-4o mini and reasoning with o1 works synergistically. Jared notes how this unlocks real capabilities for funded startups.10:25–14:26 · The partners pushing back 0/10 Startup Demo 2: Camfer Automating 3D CAD Design in SolidWorks The partners review Camfer's natural language CAD workflow and its ability to solve Navier-Stokes equations for airfoils. Harj and Gary explore how shifting compute to the inference stage mirrors human scientific iteration.14:26–16:53 · The partners pushing back 0/10 Reinforcement Learning Evolution: From Dota 2 to OpenAI o1 Diana and Jared trace OpenAI's technical heritage from Dota 2 self-play and Q-learning through to reward modeling in o1. Diana explains why factual verification domains like math and science benefit most from reinforcement learning.16:53–21:35 · The partners pushing back 0/10 Dual Vectors of AI Progress: Base Scaling vs. Inference Reasoning Jared explains the divergence between pre-training model scale and test-time RL compute. Gary outlines how proprietary eval datasets from unindexed enterprise workflows form the primary defensive moat for vertical AI startups.21:35–31:50 · The partners pushing back 0/10 High-Precision AI and Case Study: GigaML with Zepto Harj and Gary dissect GigaML's business evolution from model fine-tuning to automated customer support at Zepto. Diana reports on GigaML's performance jump from a 70 percent error rate to 5 percent by pairing o1 with structured evals.31:50–35:03 · The partners pushing back 0/10 Deprecated Startup Ideas vs. Physical World Breakthroughs Harj and Gary evaluate the risks facing coding agent startups that rely on proprietary prompt wrappers without directability. Diana points to physical-world engineering disciplines as the highest-upside beneficiaries of reasoning models.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 42.3% · guest 57.7%0:00 · the partners 42.3% · guest 57.7%3:00 · the partners 100% · guest 0%3:00 · the partners 100% · guest 0%6:00 · the partners 36.4% · guest 63.6%6:00 · the partners 36.4% · guest 63.6%9:00 · the partners 87.8% · guest 12.2%9:00 · the partners 87.8% · guest 12.2%12:00 · the partners 61.9% · guest 38.1%12:00 · the partners 61.9% · guest 38.1%15:00 · the partners 99.8% · guest 0.2%15:00 · the partners 99.8% · guest 0.2%18:00 · the partners 9.2% · guest 90.8%18:00 · the partners 9.2% · guest 90.8%21:00 · the partners 60.4% · guest 39.6%21:00 · the partners 60.4% · guest 39.6%24:00 · the partners 94.1% · guest 5.9%24:00 · the partners 94.1% · guest 5.9%27:00 · the partners 60.9% · guest 39.1%27:00 · the partners 60.9% · guest 39.1%30:00 · the partners 56.8% · guest 43.2%30:00 · the partners 56.8% · guest 43.2%33:00 · the partners 50.5% · guest 49.5%33:00 · the partners 50.5% · guest 49.5%
Sharpest disagreement ▶ 32:18 Harj reframes pivot suggestion

Harj moderates Diana's suggestion that certain startup categories should pivot, arguing instead that coding agent teams need to rethink their underlying value proposition.

Hardest push from the partners ▶ 22:33 Pushback on commoditization of engineering

Harj and Jared directly challenge the popular narrative that reasoning models commoditize software development, arguing that elite technical teams will capture the remaining high-margin accuracy gap.

Biggest teaching moment ▶ 6:25 Diode founder demonstrates live circuit compilation

The Diode Computer founder demonstrates the operational software generating full board layouts and running auto-routing from natural language constraints.

The partners hold their own ▶ 14:44 Diana connects Q-learning to o1 reasoning

Diana breaks down the lineage of reinforcement learning from early self-play algorithms in Dota to modern verifiability reward functions in o1.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
OpenAI's Origins and Sam Altman's Techno-Optimistic Vision 6000 Jared and Gary provide historical perspective on OpenAI's inception at Y Combinator, framing Sam Altman's long-term focus on reasoning and scientific acceleration. Harj introduces the emerging capability of AI in chip design.
Startup Demo 1: Diode Computer Automating Circuit and PCB Design 8200 Diana provides detailed domain context on PCB architecture, NP-complete routing, and component selection limits under GPT-4. The Diode Computer founder demonstrates automated board layout and schematic generation using o1 reasoning.
Multi-Model AI Workflows and Real Capability Unlocks 7000 Gary and Diana analyze multi-model pipeline architecture, explaining why extracting PDF datasheets with GPT-4o mini and reasoning with o1 works synergistically. Jared notes how this unlocks real capabilities for funded startups.
Startup Demo 2: Camfer Automating 3D CAD Design in SolidWorks 7000 The partners review Camfer's natural language CAD workflow and its ability to solve Navier-Stokes equations for airfoils. Harj and Gary explore how shifting compute to the inference stage mirrors human scientific iteration.
Reinforcement Learning Evolution: From Dota 2 to OpenAI o1 8000 Diana and Jared trace OpenAI's technical heritage from Dota 2 self-play and Q-learning through to reward modeling in o1. Diana explains why factual verification domains like math and science benefit most from reinforcement learning.
Dual Vectors of AI Progress: Base Scaling vs. Inference Reasoning 8000 Jared explains the divergence between pre-training model scale and test-time RL compute. Gary outlines how proprietary eval datasets from unindexed enterprise workflows form the primary defensive moat for vertical AI startups.
High-Precision AI and Case Study: GigaML with Zepto 7000 Harj and Gary dissect GigaML's business evolution from model fine-tuning to automated customer support at Zepto. Diana reports on GigaML's performance jump from a 70 percent error rate to 5 percent by pairing o1 with structured evals.
Deprecated Startup Ideas vs. Physical World Breakthroughs 7000 Harj and Gary evaluate the risks facing coding agent startups that rely on proprietary prompt wrappers without directability. Diana points to physical-world engineering disciplines as the highest-upside beneficiaries of reasoning models.

Statements from this episode (18)

Assertion Not checkable as stated
Gary Tan: Sam Altman directly estimated AGI is 4 to 15 years away
“Seeing him on Monday, he actually directly estimated, you know, between four and 15 years.”
Garry Tan Nov 14, 2024 ▶ 1:26
Assertion Not checkable as stated
Friedman: Altman founded OpenAI believing AGI would accelerate all scientific progress
“I remember back when he was starting open AI, one of the things that really motivated him to do it was he believed that when we actually had AGI basically be better at doing science than humans were, and therefore would accelerate the rate of all scientific pr…”
Jared Friedman Nov 14, 2024 ▶ 2:50
Prediction Not checkable as stated
Taggar: AI Is on Track to Design Chips Better Than Humans
“At some point, the AI will get good enough to just, like, design chips better than, like, humans can, and then it will just, like, eliminate one of its bottlenecks for, like, getting greater intelligence, and so it feels like that's already kind of, like, We'r…”
Harj Taggar Nov 14, 2024 ▶ 4:00
Assertion Supported
Hu: Diode demonstrated that OpenAI's o1 can automate PCB system design
“But the thing that they demonstrated now with O-one was actually able to do the system design and component selection, which is crazy. So it would be able to read all the data sheets and select the right components.”
Diana Hu Nov 14, 2024 ▶ 5:50
Insight
Hu: Top AI Products Use Multi-Model Pipelines Splitting Extraction and Reasoning
“I think this is actually a very common pattern that we're seeing a lot of the interesting products built with AI. You use different kinds of models. So, yes, four O mini is for PDF extraction, and then O one for the reasoning, because it's actually very hard t…”
Diana Hu Nov 14, 2024 ▶ 8:55
Assertion Not checkable as stated
Friedman: Diode Failed on GPT-4 but Worked on o1 With Same Prompts
“The other thing I think is interesting about this example is like during the batch before a one came out, diode had tried to do this with GPT four. Oh, and it just flat out didn't work. And then they basically tried the same thing, the same prompts, but fed it…”
Jared Friedman Nov 14, 2024 ▶ 9:24
Assertion Not publicly verifiable
Hu: OpenAI's o1 solves Navier-Stokes equations for airfoil design
“O-one was actually able to write All of these equations, all these partial differential equations, and solve basically Naive Stokes questions to actually solve airfoil.”
Diana Hu Nov 14, 2024 ▶ 11:37
Assertion Not checkable as stated
Tan: Sam Altman aims to scale AI compute to $1 trillion spend
“Since then, Sam has told me that he actually wants to go to four orders of magnitude to get to a trillion dollars in you know, sort of spend.”
Garry Tan Nov 14, 2024 ▶ 12:00
Insight
Hu: OpenAI o1 Combines Next-Token LLMs with RL Reward Functions
“GPT is all generative based on predicting the next token and patterns and then getting those results to check that they're correct. So I think a lot of it is you had to have a lot of data that was factually correct and Fed into probably the model and the train…”
Diana Hu Nov 14, 2024 ▶ 16:09
Assertion Supported
Friedman: Full OpenAI o1 Model Is a Huge Step Function Above o1-Preview
“Like the full O-one model, which is coming out any day now is a huge step function above even O-one preview, which is what enabled all these incredible results at the hackathon.”
Jared Friedman Nov 14, 2024 ▶ 17:43
Assertion Not checkable as stated
Friedman: Sam Altman Says OpenAI o2 and o3 Are Not Far Behind
“Sam was just telling us that like O-two and O-three are not far behind.”
Jared Friedman Nov 14, 2024 ▶ 17:54
Prediction Not checkable as stated
Tan: OpenAI o1 reasoning will replace manual workflow prompt engineering
“It sounds like basically with O-one, the chain of thoughts will replace the workflow. So you might not need to break it down into steps yourself, but the evals are still really important.”
Garry Tan Nov 14, 2024 ▶ 18:57
Insight
Tan: AI startup moats come from proprietary data for domain evals
“You can almost Argue that anything that is consumer and publicly available on the internet, that's going to be in the base model. So then your moat ultimately is for all of the other things that are not already online, whether it's, you know, for case techs be…”
Garry Tan Nov 14, 2024 ▶ 20:57
Assertion Open · timeframe Nov 2027
Tan: GigaML Automated 30,000 Zepto Customer Support Tickets Daily
“So last time I did office hours with them, they said that they automated 30,000 tickets per day.”
Garry Tan Nov 14, 2024 ▶ 27:53
Assertion Supported
Tan: OpenAI Uses Fake Model to Hide Raw o1 Chain of Thought
“If you use O-one in ChatGPT, it looks like it will tell you what's really going on, but apparently they have a fake model that just spits out things to give you the impression that it's breaking it up into steps. And they've actually You know, hidden it, becau…”
Garry Tan Nov 14, 2024 ▶ 30:32
Opinion
Taggar: AI coding startups risk obsolescence from OpenAI o1's native reasoning
“I wouldn't go all the way and suggest they should pivot, but I do think companies that are building AI coding agents or AI program engineers are potentially have stuff to think about here, because it seems like O-One in particular is like outperforming on just…”
Harj Taggar Nov 14, 2024 ▶ 32:19
Assertion Not checkable as stated
Friedman: AI phone calling startups are booming after failing a year ago
“Like a year ago doing startup ideas where like the AI agent would talk on the phone just like didn't work. We had a bunch of companies that tried and all the companies didn't work. And over the summer it really started working under the trends from the past tw…”
Jared Friedman Nov 14, 2024 ▶ 33:32
Opinion
Hu: OpenAI o1 unlocks startups in physical engineering and biology
“To connect to Sam's essay is a lot of things that are going to make the atom world, physical world better because it's really good at math and physics. So any startup that's working around mechanical engineering, electrical engineering, chemical engineering, b…”
Diana Hu Nov 14, 2024 ▶ 34:00
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.