Nov 10, 2025 · 35m · latent-space

⚡ Inside GitHub’s AI Revolution: Jared Palmer Reveals Agent HQ & The Future of Coding Agents

Jared Palmer · 23m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this in-depth interview on Latent Space, GitHub Senior Vice President Jared Palmer discusses the launch of GitHub Agent HQ, sharing key technical insights into the evolution of coding agents, runtime sandboxing, reliability engineering, and the modernization of core developer workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.9 Guest teaching 3.0 Guest disagreement 1.0 The hosts pushing back 0.9
05100:0010:0020:0030:001:36–10:21 · The hosts as informed peer 4/10 Scaling from Niche Tools to GitHub's Broad Ecosystem The conversation is very friendly and collaborative. The host prompts Jared on the origin story of v0 and AI SDK while contributing historical milestones (like GPT-4 context limits and Code Interpreter).10:21–14:04 · The hosts as informed peer 6/10 Synthetic Composite Models vs. Multi-Agent Ecosystems The host draws on his Cognition engineering background to discuss composite models vs. model pickers, pushing the argument that model switching at the base API layer is the wrong abstraction for coding agents.14:04–16:24 · The hosts as informed peer 6/10 Redefining Coding Agents, Tool Execution, and Skills When Jared asks about skills versus MCPs, the host takes the floor to explain DXT and how skills serve as an LLM-friendly file-system-first interface.16:24–18:29 · The hosts as informed peer 3/10 GitHub Agent HQ and Unified AI Developer Workflows Jared explains the vision behind GitHub Agent HQ and the closer integration between VS Code, Azure, and GitHub. The host reacts enthusiastically with minimal challenge.18:29–21:36 · The hosts as informed peer 6/10 Sandboxes, Dev Containers, and Repo Environment Setup The host brings up container sandboxing setups, noting Cognition's Kubernetes pods and his previous attempt at Netlify to open-source standardized framework auto-detection.21:38–23:55 · The hosts as informed peer 6/10 Protocol Standardization, MCPs, and Computer Use The host educates Jared on ACP payments and shares technical details on computer use advances driven by open OCR models like DeepSeek OCR.23:56–26:30 · The hosts as informed peer 3/10 Reliability Engineering: The Climb from 90% to 99% Accuracy Jared explains the harsh reality of reliability engineering in AI products, asserting that many teams live in 'la-la land' regarding infra failure rates and error-free multi-turn sessions.26:30–29:04 · The hosts as informed peer 6/10 Applying Coding Agents to Knowledge Work and Browsers The discussion covers applying coding agents to knowledge work and browsers. The host demonstrates technical depth by describing his custom open-source tool Chrome Dump and attempts to build within Tauri.29:04–34:44 · The hosts as informed peer 4/10 Overhauling the GitHub Homepage Experience Jared delivers an in-depth breakdown of stacked diffs, comparing Facebook's Buck/Mercurial workflow against Git pull requests and explaining why large repos demand stacked workflows.1:36–10:21 · Guest teaching 3/10 Scaling from Niche Tools to GitHub's Broad Ecosystem The conversation is very friendly and collaborative. The host prompts Jared on the origin story of v0 and AI SDK while contributing historical milestones (like GPT-4 context limits and Code Interpreter).10:21–14:04 · Guest teaching 2/10 Synthetic Composite Models vs. Multi-Agent Ecosystems The host draws on his Cognition engineering background to discuss composite models vs. model pickers, pushing the argument that model switching at the base API layer is the wrong abstraction for coding agents.14:04–16:24 · Guest teaching 2/10 Redefining Coding Agents, Tool Execution, and Skills When Jared asks about skills versus MCPs, the host takes the floor to explain DXT and how skills serve as an LLM-friendly file-system-first interface.16:24–18:29 · Guest teaching 3/10 GitHub Agent HQ and Unified AI Developer Workflows Jared explains the vision behind GitHub Agent HQ and the closer integration between VS Code, Azure, and GitHub. The host reacts enthusiastically with minimal challenge.18:29–21:36 · Guest teaching 2/10 Sandboxes, Dev Containers, and Repo Environment Setup The host brings up container sandboxing setups, noting Cognition's Kubernetes pods and his previous attempt at Netlify to open-source standardized framework auto-detection.21:38–23:55 · Guest teaching 2/10 Protocol Standardization, MCPs, and Computer Use The host educates Jared on ACP payments and shares technical details on computer use advances driven by open OCR models like DeepSeek OCR.23:56–26:30 · Guest teaching 5/10 Reliability Engineering: The Climb from 90% to 99% Accuracy Jared explains the harsh reality of reliability engineering in AI products, asserting that many teams live in 'la-la land' regarding infra failure rates and error-free multi-turn sessions.26:30–29:04 · Guest teaching 2/10 Applying Coding Agents to Knowledge Work and Browsers The discussion covers applying coding agents to knowledge work and browsers. The host demonstrates technical depth by describing his custom open-source tool Chrome Dump and attempts to build within Tauri.29:04–34:44 · Guest teaching 6/10 Overhauling the GitHub Homepage Experience Jared delivers an in-depth breakdown of stacked diffs, comparing Facebook's Buck/Mercurial workflow against Git pull requests and explaining why large repos demand stacked workflows.1:36–10:21 · Guest disagreement 1/10 Scaling from Niche Tools to GitHub's Broad Ecosystem The conversation is very friendly and collaborative. The host prompts Jared on the origin story of v0 and AI SDK while contributing historical milestones (like GPT-4 context limits and Code Interpreter).10:21–14:04 · Guest disagreement 1/10 Synthetic Composite Models vs. Multi-Agent Ecosystems The host draws on his Cognition engineering background to discuss composite models vs. model pickers, pushing the argument that model switching at the base API layer is the wrong abstraction for coding agents.14:04–16:24 · Guest disagreement 1/10 Redefining Coding Agents, Tool Execution, and Skills When Jared asks about skills versus MCPs, the host takes the floor to explain DXT and how skills serve as an LLM-friendly file-system-first interface.16:24–18:29 · Guest disagreement 0/10 GitHub Agent HQ and Unified AI Developer Workflows Jared explains the vision behind GitHub Agent HQ and the closer integration between VS Code, Azure, and GitHub. The host reacts enthusiastically with minimal challenge.18:29–21:36 · Guest disagreement 1/10 Sandboxes, Dev Containers, and Repo Environment Setup The host brings up container sandboxing setups, noting Cognition's Kubernetes pods and his previous attempt at Netlify to open-source standardized framework auto-detection.21:38–23:55 · Guest disagreement 1/10 Protocol Standardization, MCPs, and Computer Use The host educates Jared on ACP payments and shares technical details on computer use advances driven by open OCR models like DeepSeek OCR.23:56–26:30 · Guest disagreement 2/10 Reliability Engineering: The Climb from 90% to 99% Accuracy Jared explains the harsh reality of reliability engineering in AI products, asserting that many teams live in 'la-la land' regarding infra failure rates and error-free multi-turn sessions.26:30–29:04 · Guest disagreement 1/10 Applying Coding Agents to Knowledge Work and Browsers The discussion covers applying coding agents to knowledge work and browsers. The host demonstrates technical depth by describing his custom open-source tool Chrome Dump and attempts to build within Tauri.29:04–34:44 · Guest disagreement 1/10 Overhauling the GitHub Homepage Experience Jared delivers an in-depth breakdown of stacked diffs, comparing Facebook's Buck/Mercurial workflow against Git pull requests and explaining why large repos demand stacked workflows.1:36–10:21 · The hosts pushing back 1/10 Scaling from Niche Tools to GitHub's Broad Ecosystem The conversation is very friendly and collaborative. The host prompts Jared on the origin story of v0 and AI SDK while contributing historical milestones (like GPT-4 context limits and Code Interpreter).10:21–14:04 · The hosts pushing back 2/10 Synthetic Composite Models vs. Multi-Agent Ecosystems The host draws on his Cognition engineering background to discuss composite models vs. model pickers, pushing the argument that model switching at the base API layer is the wrong abstraction for coding agents.14:04–16:24 · The hosts pushing back 1/10 Redefining Coding Agents, Tool Execution, and Skills When Jared asks about skills versus MCPs, the host takes the floor to explain DXT and how skills serve as an LLM-friendly file-system-first interface.16:24–18:29 · The hosts pushing back 0/10 GitHub Agent HQ and Unified AI Developer Workflows Jared explains the vision behind GitHub Agent HQ and the closer integration between VS Code, Azure, and GitHub. The host reacts enthusiastically with minimal challenge.18:29–21:36 · The hosts pushing back 1/10 Sandboxes, Dev Containers, and Repo Environment Setup The host brings up container sandboxing setups, noting Cognition's Kubernetes pods and his previous attempt at Netlify to open-source standardized framework auto-detection.21:38–23:55 · The hosts pushing back 1/10 Protocol Standardization, MCPs, and Computer Use The host educates Jared on ACP payments and shares technical details on computer use advances driven by open OCR models like DeepSeek OCR.23:56–26:30 · The hosts pushing back 1/10 Reliability Engineering: The Climb from 90% to 99% Accuracy Jared explains the harsh reality of reliability engineering in AI products, asserting that many teams live in 'la-la land' regarding infra failure rates and error-free multi-turn sessions.26:30–29:04 · The hosts pushing back 0/10 Applying Coding Agents to Knowledge Work and Browsers The discussion covers applying coding agents to knowledge work and browsers. The host demonstrates technical depth by describing his custom open-source tool Chrome Dump and attempts to build within Tauri.29:04–34:44 · The hosts pushing back 1/10 Overhauling the GitHub Homepage Experience Jared delivers an in-depth breakdown of stacked diffs, comparing Facebook's Buck/Mercurial workflow against Git pull requests and explaining why large repos demand stacked workflows.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 24:35 Jared calls out industry complacency on reliability

Jared bluntly asserts that AI builders are living in 'la-la land' if they don't rigorously track dropped requests, provider uptime, and error-free multi-turn agent sessions.

Hardest push from the hosts ▶ 13:37 Host challenges model picker abstraction layer

The host pushes back against traditional model-switching interfaces, arguing that loosely bound generic abstractions produce lowest-common-denominator agent behavior.

Biggest teaching moment ▶ 31:20 Jared explains Facebook stacked diffs vs Git pull requests

Jared provides a masterclass on monorepo engineering, contrasting Facebook's internal Mercurial stacked commits with traditional Git PR paradigms.

The host holds their own ▶ 15:15 Host explains skills as the universal file system interface

When Jared admits unfamiliarity with Anthropic skills, the host steps up to articulate how skills package tool execution into a clean, markdown-driven file system abstraction.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Scaling from Niche Tools to GitHub's Broad Ecosystem 4311 The conversation is very friendly and collaborative. The host prompts Jared on the origin story of v0 and AI SDK while contributing historical milestones (like GPT-4 context limits and Code Interpreter).
Synthetic Composite Models vs. Multi-Agent Ecosystems 6212 The host draws on his Cognition engineering background to discuss composite models vs. model pickers, pushing the argument that model switching at the base API layer is the wrong abstraction for coding agents.
Redefining Coding Agents, Tool Execution, and Skills 6211 When Jared asks about skills versus MCPs, the host takes the floor to explain DXT and how skills serve as an LLM-friendly file-system-first interface.
GitHub Agent HQ and Unified AI Developer Workflows 3300 Jared explains the vision behind GitHub Agent HQ and the closer integration between VS Code, Azure, and GitHub. The host reacts enthusiastically with minimal challenge.
Sandboxes, Dev Containers, and Repo Environment Setup 6211 The host brings up container sandboxing setups, noting Cognition's Kubernetes pods and his previous attempt at Netlify to open-source standardized framework auto-detection.
Protocol Standardization, MCPs, and Computer Use 6211 The host educates Jared on ACP payments and shares technical details on computer use advances driven by open OCR models like DeepSeek OCR.
Reliability Engineering: The Climb from 90% to 99% Accuracy 3521 Jared explains the harsh reality of reliability engineering in AI products, asserting that many teams live in 'la-la land' regarding infra failure rates and error-free multi-turn sessions.
Applying Coding Agents to Knowledge Work and Browsers 6210 The discussion covers applying coding agents to knowledge work and browsers. The host demonstrates technical depth by describing his custom open-source tool Chrome Dump and attempts to build within Tauri.
Overhauling the GitHub Homepage Experience 4611 Jared delivers an in-depth breakdown of stacked diffs, comparing Facebook's Buck/Mercurial workflow against Git pull requests and explaining why large repos demand stacked workflows.

Statements from this episode (14)

Assertion Supported
GitHub launches Agent HQ at the GitHub Universe conference
“And today we launched Asian HQ among other things here at universe.”
Jared Palmer Nov 10, 2025 ▶ 1:21
Assertion Not checkable as stated
Vercel's v0 took nine months to reach $1M ARR
“That launches, and then probably like nine months later it took us like nine months to get to like a million ARR. This little team, but then the models progressed.”
Jared Palmer Nov 10, 2025 ▶ 9:05
Assertion Not checkable as stated
Vercel's v0 added $1M MRR every 14 days after chat rewrite
“When we launched V-Zero, the chat version, or the new V-Zero, whatever you want to call it it's like, 14 days, another million MRR, 14 days, another million MRR, it was like a rocket ship after that.”
Jared Palmer Nov 10, 2025 ▶ 9:48
Insight
Constraining v0 to Next.js and shadcn enabled faster execution
“What was really liberating for us was actually the focus on just one stack, or one framework, whatever else was trying to do, General purpose coding agent. We were like, no, we're just going to focus on Next.js front-end and Chad CNN. And that was really, that…”
Jared Palmer Nov 10, 2025 ▶ 9:58
Disclosure
Vercel shared its post-training harness with frontier AI model labs
“We also started working with all the frontier model labs to help because it was in our Vercel's best interest. To have them be created Next.js. And also because of the post train models, and you can read about this on the Vercel blog, like we can the post trai…”
Jared Palmer Nov 10, 2025 ▶ 10:33
Disclosure
GitHub Agent HQ supports third-party agents like Claude Code and Cognition
“Switching gears to GitHub. Like we are all about model choice now and like making sure that and what's cool is that like, it's, we also have copilot, which is our harness. And co-pilot CLI, but we also have third party partners like Cloud Code and Codex and Co…”
Jared Palmer Nov 10, 2025 ▶ 13:13
Assertion Supported
Microsoft CoreAI unites Visual Studio, VS Code, GitHub, and Azure
“And so you think about the new core core AI organization, you've got Visual Studio Code, GitHub, and parts of Azure, all in one.”
Jared Palmer Nov 10, 2025 ▶ 16:40
Disclosure
Microsoft has competing internal runtime sandbox solutions for AI agents
“There's work and discussion about what that runtime should be even internally at Microsoft. We've got a couple of different competing things, so we'll figure it out in the next, you know, in the next cycle here.”
Jared Palmer Nov 10, 2025 ▶ 19:29
Opinion
MCP is the primary method enterprise customers use to add agent context
“The MCP is huge, it seems like. It is the way that a lot of the, especially when it comes to, like, digital transformation or some of our enterprise customers, it's where they're going to be able to be at, or where they are able to add context.”
Jared Palmer Nov 10, 2025 ▶ 22:16
Disclosure
GitHub announces custom agent support with prompts and MCPs in Agent HQ
“In addition to that, we also have custom agents that we announced today, too. So, like, you can work with prompts and stuff within your agent HQ and customize these agents for different tasks, and those can have MCPs and such.”
Jared Palmer Nov 10, 2025 ▶ 22:27
Insight
AI builders are blind to poor quality without tracking error-free sessions
“Most people are blind, like living in la la land about how poor quality their AI product lately is, unless they're really measuring like the number of error free sessions, like how many errors are coming from the infra providers, like, you know, how, how many …”
Jared Palmer Nov 10, 2025 ▶ 24:55
Prediction Not checkable as stated
Autonomous AI agents require vastly more reliable infrastructure to succeed
“Agents will really only work where we get to not only like the more intelligent models, but better reliability of the infrastructure providers, right?”
Jared Palmer Nov 10, 2025 ▶ 25:44
Assertion Not checkable as stated
GitHub rejected stacked diff prototypes in 2020 and 2022 as too risky
“And there's been multiple attempts at this internally going back to like, 20, 20. And there was one very, very, very polished attempt to in 20, 22. And it just, I don't have all the context, so, but it was, there was a pretty good implementation. All of the wo…”
Jared Palmer Nov 10, 2025 ▶ 33:43
Disclosure
GitHub is actively exploring stacked diff support for its product roadmap
“So anyway, we're, we had cold meetings internally already, and we're trying to weave it into the planning and the roadmap, and so hopefully we'll be able to share more updates soon, but like, it's a Top of the list known feature. And so, and like, we're workin…”
Jared Palmer Nov 10, 2025 ▶ 34:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.