Jul 24, 2026 · 44m · startup-ideas

Most Valuable Skill of 2026: Managing AI Agents

Ryan Carson · 32m spoken Greg Isenberg · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Host Greg Isenberg and veteran operator Ryan Carson explore how knowledge workers can transition into elite engineering managers of autonomous AI agent fleets. The conversation provides an actionable blueprint covering cloud-based software factories, mobile orchestration, autonomous testing loops, and token cost optimization.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Greg holds 17.9% of the talking time here. How this is scored →

Greg as informed peer 3.0 Guest teaching 5.0 Guest disagreement 2.0 Greg pushing back 0.9
05100:0015:0030:001:00–3:58 · Greg as informed peer 1/10 The Shift to Becoming an Agent Operator The host sets up the premise of managing AI agent teams and openly admits he conducts zero percent of his work on mobile. The guest establishes the mindset shift required for founders to act as engineering managers scaling solo operations.3:59–7:23 · Greg as informed peer 1/10 Ryan Carson's Hardware and Software Command Center The guest gives a detailed operational walkthrough of his eight-screen setup, voice-to-text integration, and production key separation in 1Password. The host remains in pure listening mode as the guest presents tactical security protocols.7:24–14:36 · Greg as informed peer 3/10 Cloud Agents Versus Local Development Environments The guest aggressively dismisses local development proponents as 'cavemen' who hold back output. The host challenges this view by bringing up the pervasive sentiment from local maxis on X who dismiss cloud environments as amateurish.14:37–24:22 · Greg as informed peer 4/10 Operating Cadence and High-Stakes Decision Making The guest shows live GitHub data demonstrating 20 to 40 PRs shipped per day and introduces the 25-minute decision cadence. The host draws meaningful parallels between modern agent management and classic pre-AI organizational executive functions.24:24–34:38 · Greg as informed peer 3/10 Architecting Autonomous Testing and Self-Improvement Loops The guest educates the audience on multi-tier automations, from browser QA playbooks to automated rubric evaluations of paralegal chats. The host asks sharp questions clarifying the boundary between functional bug QA and UX design review.34:38–40:32 · Greg as informed peer 5/10 Token Economics, Model Routing, and Independent Labs The guest passionately criticizes founders who lock their software workflows into single frontier labs like Anthropic or OpenAI. The host intervenes to clarify that indie labs still route to frontier models and articulates the mortgage broker analogy.40:33–44:45 · Greg as informed peer 4/10 Building in Public and Final Takeaways The host provides a concrete case study of Sahil Bloom to illustrate the mechanics and compounding benefits of learning in public on X. The guest completely concurs as they wrap up key takeaways.1:00–3:58 · Guest teaching 4/10 The Shift to Becoming an Agent Operator The host sets up the premise of managing AI agent teams and openly admits he conducts zero percent of his work on mobile. The guest establishes the mindset shift required for founders to act as engineering managers scaling solo operations.3:59–7:23 · Guest teaching 5/10 Ryan Carson's Hardware and Software Command Center The guest gives a detailed operational walkthrough of his eight-screen setup, voice-to-text integration, and production key separation in 1Password. The host remains in pure listening mode as the guest presents tactical security protocols.7:24–14:36 · Guest teaching 6/10 Cloud Agents Versus Local Development Environments The guest aggressively dismisses local development proponents as 'cavemen' who hold back output. The host challenges this view by bringing up the pervasive sentiment from local maxis on X who dismiss cloud environments as amateurish.14:37–24:22 · Guest teaching 5/10 Operating Cadence and High-Stakes Decision Making The guest shows live GitHub data demonstrating 20 to 40 PRs shipped per day and introduces the 25-minute decision cadence. The host draws meaningful parallels between modern agent management and classic pre-AI organizational executive functions.24:24–34:38 · Guest teaching 7/10 Architecting Autonomous Testing and Self-Improvement Loops The guest educates the audience on multi-tier automations, from browser QA playbooks to automated rubric evaluations of paralegal chats. The host asks sharp questions clarifying the boundary between functional bug QA and UX design review.34:38–40:32 · Guest teaching 6/10 Token Economics, Model Routing, and Independent Labs The guest passionately criticizes founders who lock their software workflows into single frontier labs like Anthropic or OpenAI. The host intervenes to clarify that indie labs still route to frontier models and articulates the mortgage broker analogy.40:33–44:45 · Guest teaching 2/10 Building in Public and Final Takeaways The host provides a concrete case study of Sahil Bloom to illustrate the mechanics and compounding benefits of learning in public on X. The guest completely concurs as they wrap up key takeaways.1:00–3:58 · Guest disagreement 1/10 The Shift to Becoming an Agent Operator The host sets up the premise of managing AI agent teams and openly admits he conducts zero percent of his work on mobile. The guest establishes the mindset shift required for founders to act as engineering managers scaling solo operations.3:59–7:23 · Guest disagreement 1/10 Ryan Carson's Hardware and Software Command Center The guest gives a detailed operational walkthrough of his eight-screen setup, voice-to-text integration, and production key separation in 1Password. The host remains in pure listening mode as the guest presents tactical security protocols.7:24–14:36 · Guest disagreement 4/10 Cloud Agents Versus Local Development Environments The guest aggressively dismisses local development proponents as 'cavemen' who hold back output. The host challenges this view by bringing up the pervasive sentiment from local maxis on X who dismiss cloud environments as amateurish.14:37–24:22 · Guest disagreement 2/10 Operating Cadence and High-Stakes Decision Making The guest shows live GitHub data demonstrating 20 to 40 PRs shipped per day and introduces the 25-minute decision cadence. The host draws meaningful parallels between modern agent management and classic pre-AI organizational executive functions.24:24–34:38 · Guest disagreement 1/10 Architecting Autonomous Testing and Self-Improvement Loops The guest educates the audience on multi-tier automations, from browser QA playbooks to automated rubric evaluations of paralegal chats. The host asks sharp questions clarifying the boundary between functional bug QA and UX design review.34:38–40:32 · Guest disagreement 5/10 Token Economics, Model Routing, and Independent Labs The guest passionately criticizes founders who lock their software workflows into single frontier labs like Anthropic or OpenAI. The host intervenes to clarify that indie labs still route to frontier models and articulates the mortgage broker analogy.40:33–44:45 · Guest disagreement 0/10 Building in Public and Final Takeaways The host provides a concrete case study of Sahil Bloom to illustrate the mechanics and compounding benefits of learning in public on X. The guest completely concurs as they wrap up key takeaways.1:00–3:58 · Greg pushing back 0/10 The Shift to Becoming an Agent Operator The host sets up the premise of managing AI agent teams and openly admits he conducts zero percent of his work on mobile. The guest establishes the mindset shift required for founders to act as engineering managers scaling solo operations.3:59–7:23 · Greg pushing back 0/10 Ryan Carson's Hardware and Software Command Center The guest gives a detailed operational walkthrough of his eight-screen setup, voice-to-text integration, and production key separation in 1Password. The host remains in pure listening mode as the guest presents tactical security protocols.7:24–14:36 · Greg pushing back 2/10 Cloud Agents Versus Local Development Environments The guest aggressively dismisses local development proponents as 'cavemen' who hold back output. The host challenges this view by bringing up the pervasive sentiment from local maxis on X who dismiss cloud environments as amateurish.14:37–24:22 · Greg pushing back 1/10 Operating Cadence and High-Stakes Decision Making The guest shows live GitHub data demonstrating 20 to 40 PRs shipped per day and introduces the 25-minute decision cadence. The host draws meaningful parallels between modern agent management and classic pre-AI organizational executive functions.24:24–34:38 · Greg pushing back 1/10 Architecting Autonomous Testing and Self-Improvement Loops The guest educates the audience on multi-tier automations, from browser QA playbooks to automated rubric evaluations of paralegal chats. The host asks sharp questions clarifying the boundary between functional bug QA and UX design review.34:38–40:32 · Greg pushing back 2/10 Token Economics, Model Routing, and Independent Labs The guest passionately criticizes founders who lock their software workflows into single frontier labs like Anthropic or OpenAI. The host intervenes to clarify that indie labs still route to frontier models and articulates the mortgage broker analogy.40:33–44:45 · Greg pushing back 0/10 Building in Public and Final Takeaways The host provides a concrete case study of Sahil Bloom to illustrate the mechanics and compounding benefits of learning in public on X. The guest completely concurs as they wrap up key takeaways.

speaking balance: gold is Greg, purple is the guest (3 minute bins)

0:00 · Greg 50.8% · guest 49.2%0:00 · Greg 50.8% · guest 49.2%3:00 · Greg 3.6% · guest 96.4%3:00 · Greg 3.6% · guest 96.4%6:00 · Greg 7.8% · guest 92.2%6:00 · Greg 7.8% · guest 92.2%9:00 · Greg 0% · guest 100%9:00 · Greg 0% · guest 100%12:00 · Greg 20.2% · guest 79.8%12:00 · Greg 20.2% · guest 79.8%15:00 · Greg 0% · guest 100%15:00 · Greg 0% · guest 100%18:00 · Greg 25.1% · guest 74.9%18:00 · Greg 25.1% · guest 74.9%21:00 · Greg 31.9% · guest 68.1%21:00 · Greg 31.9% · guest 68.1%24:00 · Greg 0% · guest 100%24:00 · Greg 0% · guest 100%27:00 · Greg 17.6% · guest 82.4%27:00 · Greg 17.6% · guest 82.4%30:00 · Greg 10.9% · guest 89.1%30:00 · Greg 10.9% · guest 89.1%33:00 · Greg 8.8% · guest 91.2%33:00 · Greg 8.8% · guest 91.2%36:00 · Greg 4.2% · guest 95.8%36:00 · Greg 4.2% · guest 95.8%39:00 · Greg 36.7% · guest 63.3%39:00 · Greg 36.7% · guest 63.3%42:00 · Greg 55.6% · guest 44.4%42:00 · Greg 55.6% · guest 44.4%
Sharpest disagreement ▶ 12:56 Local dev maxis called cavemen

The guest bluntly rejects the conventional developer ethos, claiming anyone developing locally instead of using cloud VMs is a caveman holding back output by ten times.

Hardest push from Greg ▶ 13:15 Host relays local-first counterarguments

The host confronts the guest's stance by highlighting prevalent industry pushback from local maxis on social media who consider cloud VM users amateurish.

Biggest teaching moment ▶ 32:42 Self-improving agent evaluation architecture

The guest details how to construct an automated evaluation pipeline where daily conversations are graded against a rubric to trigger autonomous bug-fix PRs.

Greg holds their own ▶ 39:17 Mortgage broker model routing analogy

The host reframes the value proposition of independent agent labs by comparing them to mortgage brokers finding the best model rates, earning immediate praise from the guest.

the scores for every segment, with the reasoning behind each
ChapterTopicGreg as informed peerGuest teachingGuest disagreementGreg pushing backWhy
The Shift to Becoming an Agent Operator 1410 The host sets up the premise of managing AI agent teams and openly admits he conducts zero percent of his work on mobile. The guest establishes the mindset shift required for founders to act as engineering managers scaling solo operations.
Ryan Carson's Hardware and Software Command Center 1510 The guest gives a detailed operational walkthrough of his eight-screen setup, voice-to-text integration, and production key separation in 1Password. The host remains in pure listening mode as the guest presents tactical security protocols.
Cloud Agents Versus Local Development Environments 3642 The guest aggressively dismisses local development proponents as 'cavemen' who hold back output. The host challenges this view by bringing up the pervasive sentiment from local maxis on X who dismiss cloud environments as amateurish.
Operating Cadence and High-Stakes Decision Making 4521 The guest shows live GitHub data demonstrating 20 to 40 PRs shipped per day and introduces the 25-minute decision cadence. The host draws meaningful parallels between modern agent management and classic pre-AI organizational executive functions.
Architecting Autonomous Testing and Self-Improvement Loops 3711 The guest educates the audience on multi-tier automations, from browser QA playbooks to automated rubric evaluations of paralegal chats. The host asks sharp questions clarifying the boundary between functional bug QA and UX design review.
Token Economics, Model Routing, and Independent Labs 5652 The guest passionately criticizes founders who lock their software workflows into single frontier labs like Anthropic or OpenAI. The host intervenes to clarify that indie labs still route to frontier models and articulates the mortgage broker analogy.
Building in Public and Final Takeaways 4200 The host provides a concrete case study of Sahil Bloom to illustrate the mechanics and compounding benefits of learning in public on X. The guest completely concurs as they wrap up key takeaways.

Statements from this episode (15)

Insight
Carson: Knowledge workers across all roles are becoming AI agent managers
“Essentially you are a manager of agents now. So no matter what you used to do, whether it was a people manager, an IC, you are going to become a manager of agents now, and you need to be the best in the world.”
Ryan Carson Jul 24, 2026 ▶ 1:51
Disclosure
Carson does over half of his startup work on his phone
“You know, I do almost 50%, probably more of my work from my phone.”
Ryan Carson Jul 24, 2026 ▶ 3:25
Insight
Carson: Never give AI agents direct production write keys
“It's very important to keep your keys safe, secure, and separated from the agent, right? So I've got all my prod write keys in one password. I do not give my agents a prod write keys. So production writing is dangerous, right? And your agents will do something…”
Ryan Carson Jul 24, 2026 ▶ 5:26
Opinion
Carson: Devin is one of the best software factories in tech
“I think it's one of the best software factories in the industry. It's not cheap. But it's good.”
Ryan Carson Jul 24, 2026 ▶ 6:07
Insight
Carson: Managing AI agents makes developers more technical, not less
“The more you become a better agent manager actually the more technical you become.”
Ryan Carson Jul 24, 2026 ▶ 9:40
Disclosure
Carson runs up to ten cloud coding agents simultaneously without collisions
“So I often have, you know, at least five cloud agents working at once, often 10, and there's no risk that the code is going to collide.”
Ryan Carson Jul 24, 2026 ▶ 12:14
Opinion
Carson: Engineers coding locally are cavemen shipping 10x less code
“So I think if you are working locally, I honestly think you are a caveman. Like, I think you're, you are holding yourself back, and you are shipping 10 X less than you could be, and it is not smart. No matter if you think the cool people work locally and all t…”
Ryan Carson Jul 24, 2026 ▶ 12:53
Insight
Carson: The AI agent era requires working more, not less
“And in in order to survive in this new world, you're going to work a lot more, not a lot less.”
Ryan Carson Jul 24, 2026 ▶ 17:20
Assertion Not checkable as stated
Carson averages 22 to 25 pull requests daily using AI agents
“And you can see what, you know, the average here is, is sort of this 22 to 25 PRs a day, right? And sometimes, you know, 40 a day.”
Ryan Carson Jul 24, 2026 ▶ 17:42
Disclosure
Carson spends $60 per run for automated Devin browser tests
“And I have a it's called end-to-end sign-up test, and that runs three times a week because it is expensive. It's probably 60 bucks in tokens because it's doing a lot of, it's doing a lot of browser testing”
Ryan Carson Jul 24, 2026 ▶ 25:49
Disclosure
Carson ships three PRs daily from an autonomous AI grading loop
“Every day I have an automation that looks at these chats and then grades them on a rubric. And so you, and again, just talk to your agent about this. You pick the most important part of your app that you want to self-improve and say, here's how you agent judge…”
Ryan Carson Jul 24, 2026 ▶ 33:27
Disclosure
Carson spent $20,000 on AI tokens in a single month
“So last month I spent probably 20 grand in, in tokens. Which is just too much. Like, it's not viable.”
Ryan Carson Jul 24, 2026 ▶ 34:55
Prediction Not checkable as stated
Carson: AI engineering token costs will hit $5,000 monthly per employee
“I think all of us are in a place where we, we're getting to the spot where it's realizing, okay, for real engineering work per employee, you're looking at probably five grand a month. Like, that's probably where we're gonna shake out here.”
Ryan Carson Jul 24, 2026 ▶ 35:02
Opinion
Carson warns against building workflows on frontier models to avoid lock-in
“Like, if you are, if your whole engineering you know, motion is happening inside of clock code or inside of codex, What are you doing? Like, because they are not incentivized to make it reasonable for you long term, right? They're going to lock you into their …”
Ryan Carson Jul 24, 2026 ▶ 36:59
Prediction Not checkable as stated
Carson: AI agents will write, review, and ship 100% of code
“You will be using a software factory. Like, and what I mean is, the agents are going to be writing a hundred percent of your code, reviewing a hundred percent of your code, shipping a hundred percent of your code. Like, that's where we're going.”
Ryan Carson Jul 24, 2026 ▶ 41:18
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.