Feb 28, 2023 · 20m · y-combinator

The REAL potential of generative AI · Y Combinator

Raza Habib · 15m spoken Ali Rowghani · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Y Combinator Podcast episode, host Ali Rowghani interviews Raza Habib, Co-Founder and CEO of Humanloop, exploring how developers build commercial applications with large language models. The conversation covers fine-tuning workflows, prompt context grounding, evolving software engineering roles, AI safety, and the roadmap toward Artificial General Intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The partners as informed peer 3.8 Guest teaching 6.3 Guest disagreement 1.2 The partners pushing back 1.5
05100:0010:0020:001:30–4:32 · The partners as informed peer 3/10 Defining Large Language Models and Scaling Ali asks introductory high-level questions about what large language models are and why they have exploded in popularity. Raza provides deep educational explanations on statistical prediction, scaling laws, and hallucination reduction via context injection.4:32–6:42 · The partners as informed peer 4/10 Understanding Fine-Tuning and RLHF Ali prompts Raza to explain fine-tuning and provides conversational examples like email logs. Raza gives a thorough technical explanation of instruction tuning, RLHF, and Anthropic's scalable automated feedback methods.6:42–9:54 · The partners as informed peer 4/10 Production Data Capture and Humanloop Fine-Tuning Demo Ali asks practical developer questions about fine-tuning pipelines and summarizes the key stages. Raza outlines Humanloop's workflow addressing prototyping, evaluation, and customization.9:54–12:33 · The partners as informed peer 3/10 How Generative AI Changes the Developer Role Ali inquires about future developer workflows and upcoming technical breakthroughs. Raza explains how LLMs augment coding today and why developers might be among the first knowledge workers automated under AGI.12:33–17:35 · The partners as informed peer 6/10 AI Safety, Ethics, and Existential Risk Ali pushes back directly against Raza's skepticism regarding OpenAI's data flywheel advantage, arguing that a two-year head start and thousands of apps provide a strong moat. Raza counters that feedback data is hard to maintain across a generalized model without performance degradation.17:35–19:36 · The partners as informed peer 3/10 The Startup Explosion and Building Unique AI Products Ali asks about startup opportunities and invites Raza to pitch Humanloop's open hiring roles. The discussion is entirely collaborative and supportive.1:30–4:32 · Guest teaching 7/10 Defining Large Language Models and Scaling Ali asks introductory high-level questions about what large language models are and why they have exploded in popularity. Raza provides deep educational explanations on statistical prediction, scaling laws, and hallucination reduction via context injection.4:32–6:42 · Guest teaching 7/10 Understanding Fine-Tuning and RLHF Ali prompts Raza to explain fine-tuning and provides conversational examples like email logs. Raza gives a thorough technical explanation of instruction tuning, RLHF, and Anthropic's scalable automated feedback methods.6:42–9:54 · Guest teaching 6/10 Production Data Capture and Humanloop Fine-Tuning Demo Ali asks practical developer questions about fine-tuning pipelines and summarizes the key stages. Raza outlines Humanloop's workflow addressing prototyping, evaluation, and customization.9:54–12:33 · Guest teaching 7/10 How Generative AI Changes the Developer Role Ali inquires about future developer workflows and upcoming technical breakthroughs. Raza explains how LLMs augment coding today and why developers might be among the first knowledge workers automated under AGI.12:33–17:35 · Guest teaching 6/10 AI Safety, Ethics, and Existential Risk Ali pushes back directly against Raza's skepticism regarding OpenAI's data flywheel advantage, arguing that a two-year head start and thousands of apps provide a strong moat. Raza counters that feedback data is hard to maintain across a generalized model without performance degradation.17:35–19:36 · Guest teaching 5/10 The Startup Explosion and Building Unique AI Products Ali asks about startup opportunities and invites Raza to pitch Humanloop's open hiring roles. The discussion is entirely collaborative and supportive.1:30–4:32 · Guest disagreement 1/10 Defining Large Language Models and Scaling Ali asks introductory high-level questions about what large language models are and why they have exploded in popularity. Raza provides deep educational explanations on statistical prediction, scaling laws, and hallucination reduction via context injection.4:32–6:42 · Guest disagreement 1/10 Understanding Fine-Tuning and RLHF Ali prompts Raza to explain fine-tuning and provides conversational examples like email logs. Raza gives a thorough technical explanation of instruction tuning, RLHF, and Anthropic's scalable automated feedback methods.6:42–9:54 · Guest disagreement 1/10 Production Data Capture and Humanloop Fine-Tuning Demo Ali asks practical developer questions about fine-tuning pipelines and summarizes the key stages. Raza outlines Humanloop's workflow addressing prototyping, evaluation, and customization.9:54–12:33 · Guest disagreement 1/10 How Generative AI Changes the Developer Role Ali inquires about future developer workflows and upcoming technical breakthroughs. Raza explains how LLMs augment coding today and why developers might be among the first knowledge workers automated under AGI.12:33–17:35 · Guest disagreement 3/10 AI Safety, Ethics, and Existential Risk Ali pushes back directly against Raza's skepticism regarding OpenAI's data flywheel advantage, arguing that a two-year head start and thousands of apps provide a strong moat. Raza counters that feedback data is hard to maintain across a generalized model without performance degradation.17:35–19:36 · Guest disagreement 0/10 The Startup Explosion and Building Unique AI Products Ali asks about startup opportunities and invites Raza to pitch Humanloop's open hiring roles. The discussion is entirely collaborative and supportive.1:30–4:32 · The partners pushing back 1/10 Defining Large Language Models and Scaling Ali asks introductory high-level questions about what large language models are and why they have exploded in popularity. Raza provides deep educational explanations on statistical prediction, scaling laws, and hallucination reduction via context injection.4:32–6:42 · The partners pushing back 1/10 Understanding Fine-Tuning and RLHF Ali prompts Raza to explain fine-tuning and provides conversational examples like email logs. Raza gives a thorough technical explanation of instruction tuning, RLHF, and Anthropic's scalable automated feedback methods.6:42–9:54 · The partners pushing back 1/10 Production Data Capture and Humanloop Fine-Tuning Demo Ali asks practical developer questions about fine-tuning pipelines and summarizes the key stages. Raza outlines Humanloop's workflow addressing prototyping, evaluation, and customization.9:54–12:33 · The partners pushing back 1/10 How Generative AI Changes the Developer Role Ali inquires about future developer workflows and upcoming technical breakthroughs. Raza explains how LLMs augment coding today and why developers might be among the first knowledge workers automated under AGI.12:33–17:35 · The partners pushing back 5/10 AI Safety, Ethics, and Existential Risk Ali pushes back directly against Raza's skepticism regarding OpenAI's data flywheel advantage, arguing that a two-year head start and thousands of apps provide a strong moat. Raza counters that feedback data is hard to maintain across a generalized model without performance degradation.17:35–19:36 · The partners pushing back 0/10 The Startup Explosion and Building Unique AI Products Ali asks about startup opportunities and invites Raza to pitch Humanloop's open hiring roles. The discussion is entirely collaborative and supportive.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 0% · guest 100%0:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%
Sharpest disagreement ▶ 14:58 Raza challenges the general model moat hypothesis

Raza firmly disagrees with the premise that general foundation model creators have an insurmountable feedback moat, arguing feedback fine-tuning only confers advantages in narrow domains.

Hardest push from the partners ▶ 14:47 Ali challenges Raza on data network effects

Ali directly challenges Raza's skepticism, pressing him on why a two-year lead and thousands of production applications wouldn't generate a dominant data flywheel.

Biggest teaching moment ▶ 4:37 Raza explains RLHF vs base models

Raza clearly breaks down how InstructGPT and RLHF enabled smaller models to outperform 100x larger unaligned models, detailing the technical mechanisms behind ChatGPT's breakthrough.

The partners hold their own ▶ 14:47 Ali presses on developer ecosystem moat

Ali demonstrates domain intuition by articulating a concrete counterargument regarding accumulated application data advantages.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Defining Large Language Models and Scaling 3711 Ali asks introductory high-level questions about what large language models are and why they have exploded in popularity. Raza provides deep educational explanations on statistical prediction, scaling laws, and hallucination reduction via context injection.
Understanding Fine-Tuning and RLHF 4711 Ali prompts Raza to explain fine-tuning and provides conversational examples like email logs. Raza gives a thorough technical explanation of instruction tuning, RLHF, and Anthropic's scalable automated feedback methods.
Production Data Capture and Humanloop Fine-Tuning Demo 4611 Ali asks practical developer questions about fine-tuning pipelines and summarizes the key stages. Raza outlines Humanloop's workflow addressing prototyping, evaluation, and customization.
How Generative AI Changes the Developer Role 3711 Ali inquires about future developer workflows and upcoming technical breakthroughs. Raza explains how LLMs augment coding today and why developers might be among the first knowledge workers automated under AGI.
AI Safety, Ethics, and Existential Risk 6635 Ali pushes back directly against Raza's skepticism regarding OpenAI's data flywheel advantage, arguing that a two-year head start and thousands of apps provide a strong moat. Raza counters that feedback data is hard to maintain across a generalized model without performance degradation.
The Startup Explosion and Building Unique AI Products 3500 Ali asks about startup opportunities and invites Raza to pitch Humanloop's open hiring roles. The discussion is entirely collaborative and supportive.

Statements from this episode (14)

Insight
Habib: Scaling LLM parameters and data consistently improves prediction accuracy
“As you scale the language models, both in terms of the number of parameters they have, but also in the size of the data set that they're trained on, it turns out that they continue to get better and better at this prediction task.”
Raza Habib Feb 28, 2023 ▶ 2:02
Insight
Habib: Next-word prediction at scale forces LLMs to develop reasoning
“They are able to do this task extremely well. And the only way to do that is to have gotten better at under, you know, some form of reasoning and some form of knowledge.”
Raza Habib Feb 28, 2023 ▶ 2:55
Insight
Habib: ChatGPT's obsequious tone drove the need for custom, use-case-specific models
“I think when ChatGPT came out, there was a lot of frustration from people who didn't like its personality. The tone was a bit obsequious and it's, you know, it'll defer. It doesn't want to give strong opinions on things. And to me that demonstrates the need fo…”
Raza Habib Feb 28, 2023 ▶ 4:11
Assertion Supported
Habib: Fine-tuning is what separated ChatGPT from earlier GPT models
“If you look at what the difference is between ChatGPT or the most recent OpenAI Text DaVinci Three model, and what's been in the platform for two years and has not gotten as much attention, the difference is fine tuning. Like it's the same base model, more or …”
Raza Habib Feb 28, 2023 ▶ 4:38
Assertion Supported
Habib: OpenAI's 1.3B InstructGPT beat the 100x larger GPT-3 via RLHF
“In the InstructGPT paper that OpenAI released, they compared, you know, a one or two billion parameter model with instruction tuning and RHF to the full GPT-III model and people preferred that despite the fact it was a hundred times smaller.”
Raza Habib Feb 28, 2023 ▶ 5:54
Assertion Supported
Habib: Anthropic matched RLHF performance using AI-generated feedback
“Anthropic had this very exciting paper just a couple of weeks ago where actually we're able to get similar results to RLHF without the H. So just actually having a second model provide the evaluation feedback as well. And that's obviously a lot more scalable.”
Raza Habib Feb 28, 2023 ▶ 6:07
Insight
Habib: User Edits and Send Actions Serve as Model Improvement Signals
“They probably edit it, so you can capture the edited text, and they maybe get a response or they don't get a response. So all of those bits of feedback are things we would capture and then use to drive improvements of the underlying model.”
Raza Habib Feb 28, 2023 ▶ 7:27
Insight
Habib: LLM evaluation is harder than traditional ML due to subjectivity
“Then the use cases that people are building now tend to be a lot more subjective than you might have done with machine learning before. And so evaluation is a lot harder. You can't just calculate accuracy on a test set.”
Raza Habib Feb 28, 2023 ▶ 8:12
Insight
Habib: Senior developers benefit more from GitHub Copilot than junior developers
“One thing that is surprising to me is that the people who say to me they use it the most are some of the people I consider to be better or more senior developers. You might've thought this tool would help juniors more, but I think people who are more accustome…”
Raza Habib Feb 28, 2023 ▶ 10:30
Prediction Not checkable as stated
Habib: Software engineers will be among the first jobs automated by AGI
“When we do get towards things that look like AGI, I suspect that developers will actually be one of the first jobs to see large fractions of their job be automated, which I think is very counterintuitive but also predicting the future is hard, so.”
Raza Habib Feb 28, 2023 ▶ 11:19
Insight
Habib: Action-Taking Capabilities Shift LLMs from Text Generators to Agents
“One thing that I'm really excited about is actually augmenting large language models with the ability to take actions. And so we've seen a few examples of this at the startup called adapt AI that are doing this and a few others where you essentially let the la…”
Raza Habib Feb 28, 2023 ▶ 11:59
Insight
Habib: AI models bake in creator and dataset biases
“The models bake in biases and preferences that were in the model and the data and the team that built it at the time that it was being constructed.”
Raza Habib Feb 28, 2023 ▶ 13:25
Opinion
Habib: Foundation model barriers are capital and talent, not secret sauce
“Like to me, the barriers to entry of training, one of these models are mostly capital and talent. Like the people needed are still very specialized and very smart and you need lots of money to pay for GPUs. But beyond that, I don't see that much secret sauce, …”
Raza Habib Feb 28, 2023 ▶ 14:03
Opinion
Habib: Skeptical feedback data gives general model creators an insurmountable flywheel
“There's some question about, you know, whether or not the feedback data might give them a flywheel. I'm a little bit skeptical of that, that it would give them so much that no one could catch up.”
Raza Habib Feb 28, 2023 ▶ 14:39
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.