Feb 6, 2026 · 45m · y-combinator

We're All Addicted To Claude Code · Y Combinator

Calvin French-Owen · 23m spoken Garry Tan · 11m spoken Diana Hu · 2m spoken Harj Taggar · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The Light Cone, hosts Garry Tan, Diana Hu, Harj Taggar, and Jared Friedman interview Segment co-founder and former OpenAI Codex developer Calvin French-Owen about the revolutionary impact of AI coding agents like Claude Code. The panel explores technical context engineering, agent architectures, bottom-up developer tool distribution, and how software engineers are transitioning into strategic creative directors.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The partners hold 12.5% of the talking time here. How this is scored →

The partners as informed peer 5.7 Guest teaching 5.1 Guest disagreement 1.3 The partners pushing back 1.4
05100:0015:0030:0045:000:00–2:07 · The partners as informed peer 5/10 Highlights: Flying Through Code with AI Agents Gary Tan sets the stage using his personal marathon analogy and hands-on coding experience with Claude Code to frame the discussion. Kelvin reciprocates cordially while grounding the conversation in startup realities versus enterprise risk.2:07–4:19 · The partners as informed peer 4/10 From IDE Plugins to CLI Coding Agents Kelvin explains the architectural shift from IDE-based tools like Cursor to CLI sub-agent workflows in Claude Code, demonstrating how Anthropic splits context across file-traversal sub-agents running Haiku. The hosts listen attentively to his firsthand developer insights.4:19–6:25 · The partners as informed peer 6/10 The Retro-Future Triumph of Terminal CLIs Diana Hu and Gary Tan contribute strong technical observations about the terminal being the purest composable interface and how CLI agents seamlessly tap local environments. Kelvin and Jarek agree on the retro-future nature of terminal CLIs beating IDEs.6:25–8:28 · The partners as informed peer 6/10 Bottom-Up vs. Top-Down Enterprise Software Distribution Jarek and Kelvin debate the bottom-up versus top-down enterprise software adoption dynamics. Gary Tan pushes back with a historical parallel to Netscape Navigator's non-commercial licensing discovery playbook.8:28–12:28 · The partners as informed peer 6/10 How LLMs Select Dev Tools & The Supabase Case Study Kelvin shares insights into generative engine optimization (GEO), while Diana Hu cites Supabase's open-source documentation dominance in LLM recommendations. Kelvin details how coding agents leverage grep over semantic search because of code density.12:28–15:12 · The partners as informed peer 4/10 Boilerplate Reduction, Automated Reviews & Context Poisoning Kelvin provides an extensive breakdown of tactical practices to be a top 1 percent agent user, detailing minimal boilerplate stacks, code review bots, and active context clearing to avoid context poisoning.15:12–17:55 · The partners as informed peer 7/10 Context Window Degradation & Canary Verification Tricks Diana Hu actively introduces the canary verification technique for context window degradation, which Kelvin validates. Kelvin contrasts Claude Code's session context limits with Codex's rolling turn compaction mechanism.17:55–20:46 · The partners as informed peer 6/10 Anthropic vs. OpenAI DNA: Tooling vs. Autonomous Intelligence Gary Tan prompts a high-level comparison between Anthropic and OpenAI architectures. Kelvin provides a sharp contrast between Anthropic's human-tooling DNA and OpenAI's autonomous RL intelligence focus, which Gary complements with an AlphaGo analogy.20:46–22:48 · The partners as informed peer 6/10 Engineering Management, Mental Models & Product Kernels Gary raises the challenge of teaching engineering management and error triage to younger engineers who rely on AI. Kelvin uses Slack's early primitives as an example of foundational product kernels that define downstream agent expansions.22:48–25:07 · The partners as informed peer 5/10 Y Combinator Batch Application Announcement Following a brief YC batch announcement, Harj Taggar asks which engineers benefit most from AI tooling. Kelvin explains that senior engineers with strong architectural intuition gain leverage over junior engineers.25:07–28:14 · The partners as informed peer 5/10 Modern CS Curriculum: Core Systems & Agent Tinkering Harj Taggar asks how to structure a modern computer science curriculum. Kelvin stresses systems fundamentals paired with rapid iterative agent prototyping, while Harj and Gary reflect on builder taste and ADHD-style multi-tasking execution.28:14–31:20 · The partners as informed peer 6/10 Hyper-Personalized Applications & Personal Cloud Agents Gary, Jarek, and Diana discuss personal cloud agent forks and customized agentic interfaces. Harj and Jarek note how agents eliminate the mental context-loading cost that previously barred managers from coding in short intervals.31:20–34:46 · The partners as informed peer 6/10 Zero-Value Integration Plumbing vs. Campaign Execution Diana Hu asks Kelvin how he would rebuild Segment today. Kelvin points out that foundational plumbing and integration value has plummeted to zero, shifting product value up to campaign-level execution and autonomous customer intelligence.34:46–39:55 · The partners as informed peer 7/10 Test-Driven Development as Evals for Coding Agents Gary Tan and Diana Hu highlight test-driven development as the essential eval framework that prevents agent regression. Gary shares a detailed war story debugging an ActiveJob serialization bug across Rails source code.39:55–41:41 · The partners as informed peer 6/10 Claudebots, Codex Scripting & The Shift to AI Directors The conversation covers Claudebot multi-agent networks, Codex sandbox security philosophies, and prompt injection defense. Gary Tan challenges runtime support omissions for Ruby on Rails in Codex, and Kelvin contextualizes OpenAI's Python mono-repo bias.0:00–2:07 · Guest teaching 3/10 Highlights: Flying Through Code with AI Agents Gary Tan sets the stage using his personal marathon analogy and hands-on coding experience with Claude Code to frame the discussion. Kelvin reciprocates cordially while grounding the conversation in startup realities versus enterprise risk.2:07–4:19 · Guest teaching 6/10 From IDE Plugins to CLI Coding Agents Kelvin explains the architectural shift from IDE-based tools like Cursor to CLI sub-agent workflows in Claude Code, demonstrating how Anthropic splits context across file-traversal sub-agents running Haiku. The hosts listen attentively to his firsthand developer insights.4:19–6:25 · Guest teaching 4/10 The Retro-Future Triumph of Terminal CLIs Diana Hu and Gary Tan contribute strong technical observations about the terminal being the purest composable interface and how CLI agents seamlessly tap local environments. Kelvin and Jarek agree on the retro-future nature of terminal CLIs beating IDEs.6:25–8:28 · Guest teaching 3/10 Bottom-Up vs. Top-Down Enterprise Software Distribution Jarek and Kelvin debate the bottom-up versus top-down enterprise software adoption dynamics. Gary Tan pushes back with a historical parallel to Netscape Navigator's non-commercial licensing discovery playbook.8:28–12:28 · Guest teaching 6/10 How LLMs Select Dev Tools & The Supabase Case Study Kelvin shares insights into generative engine optimization (GEO), while Diana Hu cites Supabase's open-source documentation dominance in LLM recommendations. Kelvin details how coding agents leverage grep over semantic search because of code density.12:28–15:12 · Guest teaching 7/10 Boilerplate Reduction, Automated Reviews & Context Poisoning Kelvin provides an extensive breakdown of tactical practices to be a top 1 percent agent user, detailing minimal boilerplate stacks, code review bots, and active context clearing to avoid context poisoning.15:12–17:55 · Guest teaching 5/10 Context Window Degradation & Canary Verification Tricks Diana Hu actively introduces the canary verification technique for context window degradation, which Kelvin validates. Kelvin contrasts Claude Code's session context limits with Codex's rolling turn compaction mechanism.17:55–20:46 · Guest teaching 6/10 Anthropic vs. OpenAI DNA: Tooling vs. Autonomous Intelligence Gary Tan prompts a high-level comparison between Anthropic and OpenAI architectures. Kelvin provides a sharp contrast between Anthropic's human-tooling DNA and OpenAI's autonomous RL intelligence focus, which Gary complements with an AlphaGo analogy.20:46–22:48 · Guest teaching 5/10 Engineering Management, Mental Models & Product Kernels Gary raises the challenge of teaching engineering management and error triage to younger engineers who rely on AI. Kelvin uses Slack's early primitives as an example of foundational product kernels that define downstream agent expansions.22:48–25:07 · Guest teaching 6/10 Y Combinator Batch Application Announcement Following a brief YC batch announcement, Harj Taggar asks which engineers benefit most from AI tooling. Kelvin explains that senior engineers with strong architectural intuition gain leverage over junior engineers.25:07–28:14 · Guest teaching 5/10 Modern CS Curriculum: Core Systems & Agent Tinkering Harj Taggar asks how to structure a modern computer science curriculum. Kelvin stresses systems fundamentals paired with rapid iterative agent prototyping, while Harj and Gary reflect on builder taste and ADHD-style multi-tasking execution.28:14–31:20 · Guest teaching 4/10 Hyper-Personalized Applications & Personal Cloud Agents Gary, Jarek, and Diana discuss personal cloud agent forks and customized agentic interfaces. Harj and Jarek note how agents eliminate the mental context-loading cost that previously barred managers from coding in short intervals.31:20–34:46 · Guest teaching 6/10 Zero-Value Integration Plumbing vs. Campaign Execution Diana Hu asks Kelvin how he would rebuild Segment today. Kelvin points out that foundational plumbing and integration value has plummeted to zero, shifting product value up to campaign-level execution and autonomous customer intelligence.34:46–39:55 · Guest teaching 5/10 Test-Driven Development as Evals for Coding Agents Gary Tan and Diana Hu highlight test-driven development as the essential eval framework that prevents agent regression. Gary shares a detailed war story debugging an ActiveJob serialization bug across Rails source code.39:55–41:41 · Guest teaching 6/10 Claudebots, Codex Scripting & The Shift to AI Directors The conversation covers Claudebot multi-agent networks, Codex sandbox security philosophies, and prompt injection defense. Gary Tan challenges runtime support omissions for Ruby on Rails in Codex, and Kelvin contextualizes OpenAI's Python mono-repo bias.0:00–2:07 · Guest disagreement 1/10 Highlights: Flying Through Code with AI Agents Gary Tan sets the stage using his personal marathon analogy and hands-on coding experience with Claude Code to frame the discussion. Kelvin reciprocates cordially while grounding the conversation in startup realities versus enterprise risk.2:07–4:19 · Guest disagreement 2/10 From IDE Plugins to CLI Coding Agents Kelvin explains the architectural shift from IDE-based tools like Cursor to CLI sub-agent workflows in Claude Code, demonstrating how Anthropic splits context across file-traversal sub-agents running Haiku. The hosts listen attentively to his firsthand developer insights.4:19–6:25 · Guest disagreement 1/10 The Retro-Future Triumph of Terminal CLIs Diana Hu and Gary Tan contribute strong technical observations about the terminal being the purest composable interface and how CLI agents seamlessly tap local environments. Kelvin and Jarek agree on the retro-future nature of terminal CLIs beating IDEs.6:25–8:28 · Guest disagreement 2/10 Bottom-Up vs. Top-Down Enterprise Software Distribution Jarek and Kelvin debate the bottom-up versus top-down enterprise software adoption dynamics. Gary Tan pushes back with a historical parallel to Netscape Navigator's non-commercial licensing discovery playbook.8:28–12:28 · Guest disagreement 1/10 How LLMs Select Dev Tools & The Supabase Case Study Kelvin shares insights into generative engine optimization (GEO), while Diana Hu cites Supabase's open-source documentation dominance in LLM recommendations. Kelvin details how coding agents leverage grep over semantic search because of code density.12:28–15:12 · Guest disagreement 1/10 Boilerplate Reduction, Automated Reviews & Context Poisoning Kelvin provides an extensive breakdown of tactical practices to be a top 1 percent agent user, detailing minimal boilerplate stacks, code review bots, and active context clearing to avoid context poisoning.15:12–17:55 · Guest disagreement 1/10 Context Window Degradation & Canary Verification Tricks Diana Hu actively introduces the canary verification technique for context window degradation, which Kelvin validates. Kelvin contrasts Claude Code's session context limits with Codex's rolling turn compaction mechanism.17:55–20:46 · Guest disagreement 2/10 Anthropic vs. OpenAI DNA: Tooling vs. Autonomous Intelligence Gary Tan prompts a high-level comparison between Anthropic and OpenAI architectures. Kelvin provides a sharp contrast between Anthropic's human-tooling DNA and OpenAI's autonomous RL intelligence focus, which Gary complements with an AlphaGo analogy.20:46–22:48 · Guest disagreement 1/10 Engineering Management, Mental Models & Product Kernels Gary raises the challenge of teaching engineering management and error triage to younger engineers who rely on AI. Kelvin uses Slack's early primitives as an example of foundational product kernels that define downstream agent expansions.22:48–25:07 · Guest disagreement 1/10 Y Combinator Batch Application Announcement Following a brief YC batch announcement, Harj Taggar asks which engineers benefit most from AI tooling. Kelvin explains that senior engineers with strong architectural intuition gain leverage over junior engineers.25:07–28:14 · Guest disagreement 1/10 Modern CS Curriculum: Core Systems & Agent Tinkering Harj Taggar asks how to structure a modern computer science curriculum. Kelvin stresses systems fundamentals paired with rapid iterative agent prototyping, while Harj and Gary reflect on builder taste and ADHD-style multi-tasking execution.28:14–31:20 · Guest disagreement 1/10 Hyper-Personalized Applications & Personal Cloud Agents Gary, Jarek, and Diana discuss personal cloud agent forks and customized agentic interfaces. Harj and Jarek note how agents eliminate the mental context-loading cost that previously barred managers from coding in short intervals.31:20–34:46 · Guest disagreement 1/10 Zero-Value Integration Plumbing vs. Campaign Execution Diana Hu asks Kelvin how he would rebuild Segment today. Kelvin points out that foundational plumbing and integration value has plummeted to zero, shifting product value up to campaign-level execution and autonomous customer intelligence.34:46–39:55 · Guest disagreement 1/10 Test-Driven Development as Evals for Coding Agents Gary Tan and Diana Hu highlight test-driven development as the essential eval framework that prevents agent regression. Gary shares a detailed war story debugging an ActiveJob serialization bug across Rails source code.39:55–41:41 · Guest disagreement 2/10 Claudebots, Codex Scripting & The Shift to AI Directors The conversation covers Claudebot multi-agent networks, Codex sandbox security philosophies, and prompt injection defense. Gary Tan challenges runtime support omissions for Ruby on Rails in Codex, and Kelvin contextualizes OpenAI's Python mono-repo bias.0:00–2:07 · The partners pushing back 1/10 Highlights: Flying Through Code with AI Agents Gary Tan sets the stage using his personal marathon analogy and hands-on coding experience with Claude Code to frame the discussion. Kelvin reciprocates cordially while grounding the conversation in startup realities versus enterprise risk.2:07–4:19 · The partners pushing back 1/10 From IDE Plugins to CLI Coding Agents Kelvin explains the architectural shift from IDE-based tools like Cursor to CLI sub-agent workflows in Claude Code, demonstrating how Anthropic splits context across file-traversal sub-agents running Haiku. The hosts listen attentively to his firsthand developer insights.4:19–6:25 · The partners pushing back 2/10 The Retro-Future Triumph of Terminal CLIs Diana Hu and Gary Tan contribute strong technical observations about the terminal being the purest composable interface and how CLI agents seamlessly tap local environments. Kelvin and Jarek agree on the retro-future nature of terminal CLIs beating IDEs.6:25–8:28 · The partners pushing back 3/10 Bottom-Up vs. Top-Down Enterprise Software Distribution Jarek and Kelvin debate the bottom-up versus top-down enterprise software adoption dynamics. Gary Tan pushes back with a historical parallel to Netscape Navigator's non-commercial licensing discovery playbook.8:28–12:28 · The partners pushing back 1/10 How LLMs Select Dev Tools & The Supabase Case Study Kelvin shares insights into generative engine optimization (GEO), while Diana Hu cites Supabase's open-source documentation dominance in LLM recommendations. Kelvin details how coding agents leverage grep over semantic search because of code density.12:28–15:12 · The partners pushing back 1/10 Boilerplate Reduction, Automated Reviews & Context Poisoning Kelvin provides an extensive breakdown of tactical practices to be a top 1 percent agent user, detailing minimal boilerplate stacks, code review bots, and active context clearing to avoid context poisoning.15:12–17:55 · The partners pushing back 2/10 Context Window Degradation & Canary Verification Tricks Diana Hu actively introduces the canary verification technique for context window degradation, which Kelvin validates. Kelvin contrasts Claude Code's session context limits with Codex's rolling turn compaction mechanism.17:55–20:46 · The partners pushing back 2/10 Anthropic vs. OpenAI DNA: Tooling vs. Autonomous Intelligence Gary Tan prompts a high-level comparison between Anthropic and OpenAI architectures. Kelvin provides a sharp contrast between Anthropic's human-tooling DNA and OpenAI's autonomous RL intelligence focus, which Gary complements with an AlphaGo analogy.20:46–22:48 · The partners pushing back 1/10 Engineering Management, Mental Models & Product Kernels Gary raises the challenge of teaching engineering management and error triage to younger engineers who rely on AI. Kelvin uses Slack's early primitives as an example of foundational product kernels that define downstream agent expansions.22:48–25:07 · The partners pushing back 1/10 Y Combinator Batch Application Announcement Following a brief YC batch announcement, Harj Taggar asks which engineers benefit most from AI tooling. Kelvin explains that senior engineers with strong architectural intuition gain leverage over junior engineers.25:07–28:14 · The partners pushing back 1/10 Modern CS Curriculum: Core Systems & Agent Tinkering Harj Taggar asks how to structure a modern computer science curriculum. Kelvin stresses systems fundamentals paired with rapid iterative agent prototyping, while Harj and Gary reflect on builder taste and ADHD-style multi-tasking execution.28:14–31:20 · The partners pushing back 1/10 Hyper-Personalized Applications & Personal Cloud Agents Gary, Jarek, and Diana discuss personal cloud agent forks and customized agentic interfaces. Harj and Jarek note how agents eliminate the mental context-loading cost that previously barred managers from coding in short intervals.31:20–34:46 · The partners pushing back 1/10 Zero-Value Integration Plumbing vs. Campaign Execution Diana Hu asks Kelvin how he would rebuild Segment today. Kelvin points out that foundational plumbing and integration value has plummeted to zero, shifting product value up to campaign-level execution and autonomous customer intelligence.34:46–39:55 · The partners pushing back 1/10 Test-Driven Development as Evals for Coding Agents Gary Tan and Diana Hu highlight test-driven development as the essential eval framework that prevents agent regression. Gary shares a detailed war story debugging an ActiveJob serialization bug across Rails source code.39:55–41:41 · The partners pushing back 2/10 Claudebots, Codex Scripting & The Shift to AI Directors The conversation covers Claudebot multi-agent networks, Codex sandbox security philosophies, and prompt injection defense. Gary Tan challenges runtime support omissions for Ruby on Rails in Codex, and Kelvin contextualizes OpenAI's Python mono-repo bias.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 5.9% · guest 94.1%0:00 · the partners 5.9% · guest 94.1%3:00 · the partners 12.1% · guest 87.9%3:00 · the partners 12.1% · guest 87.9%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 20.4% · guest 79.6%9:00 · the partners 20.4% · guest 79.6%12:00 · the partners 7.4% · guest 92.6%12:00 · the partners 7.4% · guest 92.6%15:00 · the partners 19.8% · guest 80.2%15:00 · the partners 19.8% · guest 80.2%18:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%21:00 · the partners 5.5% · guest 94.5%21:00 · the partners 5.5% · guest 94.5%24:00 · the partners 14% · guest 86%24:00 · the partners 14% · guest 86%27:00 · the partners 25.3% · guest 74.7%27:00 · the partners 25.3% · guest 74.7%30:00 · the partners 37.5% · guest 62.5%30:00 · the partners 37.5% · guest 62.5%33:00 · the partners 5.1% · guest 94.9%33:00 · the partners 5.1% · guest 94.9%36:00 · the partners 13.7% · guest 86.3%36:00 · the partners 13.7% · guest 86.3%39:00 · the partners 20.4% · guest 79.6%39:00 · the partners 20.4% · guest 79.6%42:00 · the partners 0.7% · guest 99.3%42:00 · the partners 0.7% · guest 99.3%45:00 · the partners 8.5% · guest 91.5%45:00 · the partners 8.5% · guest 91.5%
Sharpest disagreement ▶ 7:26 Jarek Contrarian Stance on Bottom-Up Beating Top-Down CTO Sales

Jarek forcefully dismisses top-down enterprise software sales as obsolete and slow due to CTO security concerns compared to bottom-up CLI adoption.

Hardest push from the partners ▶ 7:44 Gary Tan Rebuts Bottom-Up Exclusivity with Netscape Licensing Case

Gary Tan challenges the premise that top-down enterprise enforcement is dead by citing Netscape Navigator's historic conversion of rogue internal downloads into commercial licenses.

Biggest teaching moment ▶ 31:44 Kelvin Dismantles Traditional SaaS Plumbing Value

Kelvin directly schools the panel on how AI coding agents have wiped out the core integration plumbing value proposition that built multi-billion dollar companies like Segment.

The partners hold their own ▶ 15:59 Diana Hu Educates Panel on Canary Context Degradation Probing

Diana Hu demonstrates deep practitioner expertise by introducing an esoteric prompt engineering canary trick for context degradation that Kelvin had not tried.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Highlights: Flying Through Code with AI Agents 5311 Gary Tan sets the stage using his personal marathon analogy and hands-on coding experience with Claude Code to frame the discussion. Kelvin reciprocates cordially while grounding the conversation in startup realities versus enterprise risk.
From IDE Plugins to CLI Coding Agents 4621 Kelvin explains the architectural shift from IDE-based tools like Cursor to CLI sub-agent workflows in Claude Code, demonstrating how Anthropic splits context across file-traversal sub-agents running Haiku. The hosts listen attentively to his firsthand developer insights.
The Retro-Future Triumph of Terminal CLIs 6412 Diana Hu and Gary Tan contribute strong technical observations about the terminal being the purest composable interface and how CLI agents seamlessly tap local environments. Kelvin and Jarek agree on the retro-future nature of terminal CLIs beating IDEs.
Bottom-Up vs. Top-Down Enterprise Software Distribution 6323 Jarek and Kelvin debate the bottom-up versus top-down enterprise software adoption dynamics. Gary Tan pushes back with a historical parallel to Netscape Navigator's non-commercial licensing discovery playbook.
How LLMs Select Dev Tools & The Supabase Case Study 6611 Kelvin shares insights into generative engine optimization (GEO), while Diana Hu cites Supabase's open-source documentation dominance in LLM recommendations. Kelvin details how coding agents leverage grep over semantic search because of code density.
Boilerplate Reduction, Automated Reviews & Context Poisoning 4711 Kelvin provides an extensive breakdown of tactical practices to be a top 1 percent agent user, detailing minimal boilerplate stacks, code review bots, and active context clearing to avoid context poisoning.
Context Window Degradation & Canary Verification Tricks 7512 Diana Hu actively introduces the canary verification technique for context window degradation, which Kelvin validates. Kelvin contrasts Claude Code's session context limits with Codex's rolling turn compaction mechanism.
Anthropic vs. OpenAI DNA: Tooling vs. Autonomous Intelligence 6622 Gary Tan prompts a high-level comparison between Anthropic and OpenAI architectures. Kelvin provides a sharp contrast between Anthropic's human-tooling DNA and OpenAI's autonomous RL intelligence focus, which Gary complements with an AlphaGo analogy.
Engineering Management, Mental Models & Product Kernels 6511 Gary raises the challenge of teaching engineering management and error triage to younger engineers who rely on AI. Kelvin uses Slack's early primitives as an example of foundational product kernels that define downstream agent expansions.
Y Combinator Batch Application Announcement 5611 Following a brief YC batch announcement, Harj Taggar asks which engineers benefit most from AI tooling. Kelvin explains that senior engineers with strong architectural intuition gain leverage over junior engineers.
Modern CS Curriculum: Core Systems & Agent Tinkering 5511 Harj Taggar asks how to structure a modern computer science curriculum. Kelvin stresses systems fundamentals paired with rapid iterative agent prototyping, while Harj and Gary reflect on builder taste and ADHD-style multi-tasking execution.
Hyper-Personalized Applications & Personal Cloud Agents 6411 Gary, Jarek, and Diana discuss personal cloud agent forks and customized agentic interfaces. Harj and Jarek note how agents eliminate the mental context-loading cost that previously barred managers from coding in short intervals.
Zero-Value Integration Plumbing vs. Campaign Execution 6611 Diana Hu asks Kelvin how he would rebuild Segment today. Kelvin points out that foundational plumbing and integration value has plummeted to zero, shifting product value up to campaign-level execution and autonomous customer intelligence.
Test-Driven Development as Evals for Coding Agents 7511 Gary Tan and Diana Hu highlight test-driven development as the essential eval framework that prevents agent regression. Gary shares a detailed war story debugging an ActiveJob serialization bug across Rails source code.
Claudebots, Codex Scripting & The Shift to AI Directors 6622 The conversation covers Claudebot multi-agent networks, Codex sandbox security philosophies, and prompt injection defense. Gary Tan challenges runtime support omissions for Ruby on Rails in Codex, and Kelvin contextualizes OpenAI's Python mono-repo bias.

Statements from this episode (46)

Opinion
Garry Tan: Claude Code is an unlock that enables 5x faster coding
“I recently got very, very addicted to Claude Code, and I would describe it as, like, 10 years ago I was a marathon runner, and I loved doing it, and then I suffered a catastrophic knee injury, which is called manager mode, and I stopped coding, which is tragic…”
Garry Tan Feb 6, 2026 ▶ 1:19
Prediction Not checkable as stated
French-Owen: Future coding will feel like delegating tasks to coworkers
“Hey, in the future coding is really going to feel more like talking to a coworker. Like you're going to send off a question and then they'll go off and do something and come back to you with a PR.”
Calvin French-Owen Feb 6, 2026 ▶ 2:26
Prediction Not checkable as stated
French-Owen: Software developers will all become AI managers in the future
“I think in some sense, you're right that like everyone is going to become a manager in the future, or at least that's my hot take. But in order to get there are steps along the way and you have to really build a lot of trust in the model and understand what it…”
Calvin French-Owen Feb 6, 2026 ▶ 2:55
Assertion Supported
French-Owen: Claude Code spawns Haiku sub-agents with dedicated context windows
“When you ask Claude code to do something, it will typically spawn and explore sub agent or like multiple ones. And basically each of those are running haiku to traverse the file system and kind of like explore what's there. And they're doing it in their own co…”
Calvin French-Owen Feb 6, 2026 ▶ 3:46
Insight
French-Owen: Claude Code Succeeds by Distancing Devs From Low-Level Code
“And I think it's important actually to quad code that it's not an IDE because it sort of distances you from the code that's being written. Like IDEs are all about exploring files, right? And you're like trying to keep all the state in your head and understand …”
Calvin French-Owen Feb 6, 2026 ▶ 4:57
Disclosure
Garry Tan grants Claude Code direct access to his production database
“I mean, I'm not sure if I'm supposed to do this, but I've actually also had it access my production database.”
Garry Tan Feb 6, 2026 ▶ 5:56
Assertion Not checkable as stated
Garry Tan: Claude Code can debug nested delayed jobs five levels deep
“Oh my God, like this thing can debug nested delayed jobs, like five levels in and figure out what the bug was and then write a test for it. And it never happens again.”
Garry Tan Feb 6, 2026 ▶ 6:11
Insight
French-Owen: Bypassing IT approval gives developer tools a major adoption edge
“Like thinking about a cursor or a cloud code or a codex CLI, the fact that you can just download it and use it without having to get IT permissions or anything makes a huge difference.”
Calvin French-Owen Feb 6, 2026 ▶ 6:26
Assertion Supported
Tan: Netscape tracked enterprise IPs on free downloads to enforce licenses
“That was the original Netscape Navigator. It was free for non-commercial use, and then people would just download it and use it for commercial use, and then they could just track down the IPs and figure out exactly how many clients were in all of these differe…”
Garry Tan Feb 6, 2026 ▶ 7:44
Insight
Tan: Developers adopt vendors like PostHog simply because Claude Code recommends them
“Now people are probably just making Architecture decisions about what to use directly in CloudCode. Like, they might not even know what, you know, what analytics to use. And it's like, oh yeah, as long as CloudCode says use PostHog, like, they're using PostHog…”
Garry Tan Feb 6, 2026 ▶ 8:08
Insight
French-Owen: Strong public docs and Reddit mentions drive LLM recommendations
“I think, yeah, if you're selling a developer tool, like having good docs that are out there, like having social proof, like Maybe being posted on Reddit a little bit more. All of that helps your case tremendously.”
Calvin French-Owen Feb 6, 2026 ▶ 9:01
Assertion Not checkable as stated
Diana Hu: Supabase is the default backend recommendation across all LLMs
“Whenever someone asks how to set up anything that you need, some sort of backend, Firebase type of transaction, the default answer from all the LLMs is actually a Superbase.”
Diana Hu Feb 6, 2026 ▶ 9:26
Opinion
French-Owen: Coding agents disproportionately benefit open-source software
“I will say it does help open source disproportionately, I would say. Like, I don't know if you all saw there's a ramp blog post that they recently published about building their own coding agent, and they were mentioning that they use open code as a harness be…”
Calvin French-Owen Feb 6, 2026 ▶ 9:52
Disclosure
French-Owen: OpenAI used reinforcement learning on o3 checkpoints for coding tasks
“Well, basically we kind of had like a checkpoint. For I think it was O three, like one of the reasoning models. And then we did a bunch of fine tuning on it and reinforcement learning where it's like, oh, you're given a bunch of questions to like solve these c…”
Calvin French-Owen Feb 6, 2026 ▶ 10:30
Assertion Supported
French-Owen: Claude Code and Codex Use Grep Rather Than Semantic Embeddings
“Like I think cursor takes an approach where they actually do semantic search, where they embed everything and figure out like, Hey, what query is closest to this? If you look at a codex or a cloud code they actually just use like grep.”
Calvin French-Owen Feb 6, 2026 ▶ 11:16
Insight
French-Owen: Providing automated verification loops drastically improves coding agent performance
“Obviously giving the model a way to check its work helps improve performance drastically. So the more that you can run tests and lint CI, et cetera.”
Calvin French-Owen Feb 6, 2026 ▶ 14:13
Opinion
French-Owen: Cursor BugBot and Codex perform very well at code review correctness
“I use the cursor bug bot has gotten quite good and I actually like codex for code review as well. I find it does a very good job on correctness.”
Calvin French-Owen Feb 6, 2026 ▶ 14:28
Insight
French-Owen: Coding agents suffer from context poisoning and need active context resets
“I think context poisoning is a real thing where it kind of like goes down one loop and it will continue because it has this persistence, but it's referring back to tokens, which are like not Right in terms of pursuing a solution. And so one thing that I often …”
Calvin French-Owen Feb 6, 2026 ▶ 14:54
Insight
French-Owen: LLMs enter the 'dumb zone' after exceeding context thresholds
“He has this concept of, like, the LLMs reaching the dumb zone, where it's like, after a certain amount of tokens It just starts, like, degrading in quality. And I actually think that's very true, especially if you think about, like, how the reinforcement learn…”
Calvin French-Owen Feb 6, 2026 ▶ 15:27
Insight
Diana Hu: Inserting canary facts tests if an LLM's context has degraded
“One of the tricks that I think founders use is you put like a cannery at the beginning of the context. There's something very esoteric that it would only help. It's like something really funny. It's like, I don't know. My name is Calvin and blah, blah, blah. I…”
Diana Hu Feb 6, 2026 ▶ 16:00
Assertion Supported
French-Owen: Codex Uses Periodic Compaction to Support Long-Running Tasks
“Where it will run compaction, like, periodically after each turn, and so Codex can continue to run for a very long time, and if you look at the percentage in the CLI, you'll see it, like, move up and down as compaction runs.”
Calvin French-Owen Feb 6, 2026 ▶ 17:23
Opinion
Tan: Current AI coding agents cannot run autonomously for long periods
“The coding agents right now are really, really smart, but not smart enough to run on their own for long periods of time”
Garry Tan Feb 6, 2026 ▶ 18:06
Opinion
French-Owen: OpenAI targets long-horizon autonomous intelligence for AGI
“Whereas OpenAI really leans into this idea of just, like, we are going to train the best model and reinforce over time and get it to do longer and longer horizon things in this pursuit of artificial general intelligence.”
Calvin French-Owen Feb 6, 2026 ▶ 18:54
Opinion
Garry Tan: Claude Code enables one person to do five people's daily work
“It's like, I can do five people's worth of work in like a single day. It's like rocket boosters.”
Garry Tan Feb 6, 2026 ▶ 19:51
Prediction Not checkable as stated
French-Owen: Solo engineers using AI will outperform large enterprise teams
“And I think it's going to be very strange as, like, these individual teams of, like, one person are like, hey, that team over there isn't doing the right thing, like, let me just build a prototype that, like, works better. I think at some point it's going to s…”
Calvin French-Owen Feb 6, 2026 ▶ 20:27
Insight
French-Owen: Core product primitives must be set early for coding agents to build on
“It's, like, difficult to change the user's mental model, and so I, at least for myself building products, it's like, you have to think about that Very carefully from an early stage, because again, whatever you supply to the coding agents as that kind of kernel…”
Calvin French-Owen Feb 6, 2026 ▶ 22:34
Insight
French-Owen: More senior engineers benefit most from AI coding agents
“In general, I think that kind of the more senior, senior you are, the more you benefit because the agents are so good at taking some sort of idea and then putting it into action.”
Calvin French-Owen Feb 6, 2026 ▶ 23:13
Opinion
French-Owen: AI models still struggle with software architecture
“And I think having a smell for what the right architecture is, is still the area where the models, like, don't do the best job.”
Calvin French-Owen Feb 6, 2026 ▶ 25:00
Insight
French-Owen: Systems fundamentals like HTTP and databases remain essential in CS
“Personally, I think still understanding systems is very important and just having some conception of, like, how, like, Git works, you know, or, like, HTTP, or databases, like, queues, like, all of these different systems. I think that those fundamentals are st…”
Calvin French-Owen Feb 6, 2026 ▶ 25:18
Prediction Not checkable as stated
Taggar: Top 18-to-22-Year-Olds in 5 Years Will Have 'Off the Charts' Taste
“I wonder if, like, the best 18 to twenty-two-year-olds, like, five years from now will just have, like, off the charts taste and everything. Because they'll just be so much more prolific. They should be, right? Like, they should just be launching and touching …”
Harj Taggar Feb 6, 2026 ▶ 26:27
Insight
Taggar: Claude Code Lets Multi-Project Builders Actually Finish Half-Done Ideas
“Now I just think, like, you kind of, like, there's certain types of brains that just have, like, 10 branches going in their heads, but you never have enough hours in the day to actually, like, see any of them through, so they're always, like, half complete, an…”
Harj Taggar Feb 6, 2026 ▶ 27:44
Prediction Not checkable as stated
French-Owen: Every Worker Will Eventually Have Dedicated Cloud AI Agents
“I mean, sort of what I've been thinking, I don't know how far this future is, but like eventually every. Person who's working, like has their own sort of like cloud computer and like set of cloud agents who are running for them.”
Calvin French-Owen Feb 6, 2026 ▶ 29:14
Prediction Not checkable as stated
French-Owen: AI Will Make Average Companies Smaller and More Numerous
“I think the average company is probably going to get like a little smaller and there's going to be many more of them doing more things.”
Calvin French-Owen Feb 6, 2026 ▶ 29:52
Insight
Diana Hu: Opportunity exists for an agentic-first system of record
“And the system of record, there's opportunity for something that's kind of agentic first, because right now we're still kind of integrate very much with databases and SQL or NoSQL queries at a very low level. But imagine something that generates all the data t…”
Diana Hu Feb 6, 2026 ▶ 30:57
Opinion
French-Owen: Slack restricted API access to prevent third-party agentic exfiltration
“Like I think Slack locked down their API a little bit because they didn't want people just exfiltrating everything from Slack and then building agentic experiences on top of it.”
Calvin French-Owen Feb 6, 2026 ▶ 31:27
Insight
French-Owen: The value of writing basic integration code has dropped to zero
“Segment is a funny business in that where we started was building these integrations, right? And so it's like, oh, you need to wire up, like, the same data going to, like, Mixpanel and Kissmetrics and Google Analytics, et cetera. And I think just writing that …”
Calvin French-Owen Feb 6, 2026 ▶ 31:45
Insight
French-Owen: Context window limits remain the primary bottleneck for coding agents
“I mean, I still think context window is like probably the number one limit. Like if you look at cloud code executing, it's delegating to all these different context windows. At the end of the day, when each one comes back, it's like getting some sort of summar…”
Calvin French-Owen Feb 6, 2026 ▶ 34:27
Disclosure
Garry Tan: Reaching 100% Test Coverage Supercharged Coding Velocity with Claude Code
“I was operating for, like, the first two or three days of my nine days in the wilderness. Like no tests or very few tests, and then one day I was like, all right, today's refactor day. I'm gonna do, get to a hundred percent test coverage, and then I just sped …”
Garry Tan Feb 6, 2026 ▶ 35:54
Insight
Diana Hu: Getting Good Prompts Requires Test-Driven Development via Evals
“The way you get a good prompt is all test-driven, just like evals, right? In a sense, the test cases are your evals.”
Diana Hu Feb 6, 2026 ▶ 36:30
Insight
Taggar: Never give personal AI agents access to emails due to prompt injection
“Do not give it access to emails would be my number one piece of advice or probably anything. Cause it's not clear how safe it is. And it's probably almost certainly gonna probably a lot of people being prompted injected by it right now”
Harj Taggar Feb 6, 2026 ▶ 39:09
Insight
French-Owen: Codex approaches coding in unintuitive, AlphaGo-like ways
“It's interesting to see like Codex's personality shine through when writing code. I would say it does the most stuff that humans don't do kind of in this alpha go sense where it's like, oh, it'll write a Python script to like modify some part of the file syste…”
Calvin French-Owen Feb 6, 2026 ▶ 39:36
Opinion
French-Owen: Claude Opus Misses Multi-File Bugs That Codex Catches
“Refreshing complex UI state and Opus often will miss it if there's many files, but Codex seems to catch it.”
Calvin French-Owen Feb 6, 2026 ▶ 40:37
Prediction Not checkable as stated
French-Owen: Future Coding Agent Power Users Will Act Like Managers and Designers
“I do feel like the best or the people who will get the most out of coding agents in the future are going to be kind of like more manager, like where they're focusing on directing flows in certain ways. They're probably going to be a little bit more like design…”
Calvin French-Owen Feb 6, 2026 ▶ 41:16
Assertion Not checkable as stated
French-Owen: OpenAI Codex performs particularly well on Python monorepos
“Codex works very well on Python mono repos.”
Calvin French-Owen Feb 6, 2026 ▶ 42:46
Opinion
French-Owen: OpenAI takes sandboxing and security more seriously than competitors
“The sandboxing, it's such an interesting question because I think OpenAI actually takes the like sandboxing and security question more seriously than almost anyone else.”
Calvin French-Owen Feb 6, 2026 ▶ 43:55
Assertion Not checkable as stated
Taggar: Half of YC's engineering team skips coding agent permissions
“It's about fifty-fifty on the YC engineering team.”
Harj Taggar Feb 6, 2026 ▶ 45:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.