Jan 17, 2026 · 1h 13m · latent-space

Brex’s AI Hail Mary — With CTO James Reggio (acquired for $5B by Capital One!)

James Reggio · 57m spoken Shawn Wang · 6m spoken Alessio Fanelli · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Brex CTO James Reggio details the fintech company's three-pillar AI strategy—Corporate, Operational, and Product AI—highlighting their multi-agent orchestration architecture, internal developer tooling, and founder-driven engineering culture.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.2% of the talking time here. How this is scored →

The hosts as informed peer 5.9 Guest teaching 6.2 Guest disagreement 1.5 The hosts pushing back 2.4
05100:0015:0030:0045:001:00:002:36–5:10 · The hosts as informed peer 6/10 Hiring Ex-Founders and Brex's 'Quitters Welcome' Philosophy Swyx and Alessio challenge the trendy practice of hiring ex-founders, questioning whether founder tenure is too short or an anti-signal. James acknowledges their concern but counters with Brex's deliberate 'Quitters Welcome' philosophy and how instant distribution attracts top talent.5:11–7:33 · The hosts as informed peer 5/10 Engineering Organization and the Dedicated AI Center of Excellence Alessio asks about engineering organization structure and AI adoption disparities across teams. James explains their product domain divisions and the deliberate creation of a centralized 10-person AI center of excellence designed like an internal disruptor startup.7:34–11:40 · The hosts as informed peer 5/10 Pod Composition and Managing Internal AI Perceptions Alessio raises the common cultural problem of non-AI engineers feeling alienated or unvalued. James reframes this, explaining that because Brex engineering heavily rewards revenue impact, core product engineers remain proud of driving direct business metrics over speculative AI projects.11:40–15:55 · The hosts as informed peer 6/10 Brex Agent Platform Architecture, TypeScript, and Mastra Swyx presses on Brex's choice of Mastra over established tooling like LangChain. James defends the decision by noting ergonomics and historical limitations of early LangChain before explaining their greenfield TypeScript stack.15:55–20:31 · The hosts as informed peer 6/10 Multi-Agent Architecture and the Brex Assistant James gives an in-depth breakdown of Brex's multi-agent architecture for the Brex Assistant, explaining how single-agent tool overloading failed and why encapsulated, inter-agent direct messaging modeled after human org charts worked better.20:31–22:44 · The hosts as informed peer 7/10 Model Context Protocol vs. Multi-Turn Agent Conversations Alessio and James discuss whether Anthropic's Model Context Protocol (MCP) should govern agent-to-agent communication. James clarifies that MCP suits single-turn imperative tool calls, whereas multi-agent collaboration requires conversational, multi-turn dialogue.22:44–27:32 · The hosts as informed peer 5/10 Deep Dive into Brex's Three AI Strategic Pillars James details Brex's three AI strategic pillars (corporate, operational, product). He explains how operational AI directly slashed costs in compliance and underwriting while transforming support staff into prompt and eval authors.27:32–31:23 · The hosts as informed peer 6/10 ConductorOne Self-Service Tooling and Model Flexibility Swyx expresses enthusiasm for Brex's ConductorOne Slack-based model provisioning. James explains their philosophy of avoiding vendor lock-in by letting employees choose models dynamically and using usage metrics in enterprise renewals.31:23–33:58 · The hosts as informed peer 7/10 Managing Code Slop, Review Rigor, and Shared Understanding Swyx challenges the idea that AI review tools alone can fix code slop, arguing human ownership must scale. James agrees with the downside of codebase drift and unrigorous reviews while Swyx pushes back on tool vendors claiming automated cures.33:59–37:10 · The hosts as informed peer 6/10 The Evolution of Engineering Craft and Junior Workflows James shares findings from a college dinner where new grads used LLMs for design docs and architecture rather than blind code generation. Alessio and James discuss how junior vs senior workflows differ in design and schema specification.37:11–40:00 · The hosts as informed peer 7/10 Code Consistency Rules, Greptile, and Semantic Review in CI Alessio brings up historical parallels like Danger Systems for semantic linting in CI. James validates this by detailing their adoption of Greptile and automated Claude Code review checks in GitHub Actions.40:01–45:06 · The hosts as informed peer 5/10 Operational Lessons: From Reinforcement Learning to SOP Agents James shares a key technical failure and pivot: Brex invested heavily in reinforcement learning for credit underwriting, only to find simple web research agents and prompt-based SOPs vastly outperformed complex RL models.45:06–48:53 · The hosts as informed peer 6/10 Ideal Customer Profiles and Retool-Driven Prompt Iteration Swyx presses James on whether Brex's expanded ICP represents a retreat back to the SMB segment. James clarifies the specific revenue and transaction thresholds defining their commercial segment, and highlights their Retool-based prompt management tooling.48:53–52:13 · The hosts as informed peer 6/10 Knowledge Base Grounding and the Decision to Partner with Sierra Swyx questions why Brex bought Sierra instead of building customer support agents in-house. James explains that building low-code management UIs and domain telemetry for CX leaders is undifferentiated work best outsourced.52:14–58:51 · The hosts as informed peer 7/10 Multi-Turn Evals, User Simulations, and Hallucination Mitigations Alessio suggests using forward-looking, intentionally failing evals to track model progression over time. James enthusiastically adopts the idea and explains Brex's multi-turn eval techniques and prompt guardrails against phantom agent delegation.58:51–1:03:33 · The hosts as informed peer 5/10 AI Fluency Framework, Upskilling, and Agentic Coding Interviews James explains Brex's AI fluency framework and reveals that Brex instituted mandatory agentic coding re-interviews for all existing engineers and managers to force hands-on exposure and accelerate upskilling.1:03:33–1:07:24 · The hosts as informed peer 6/10 Headcount Planning, Engineering Leverage, and Industry Headwinds Swyx probes on whether AI productivity is causing layoffs. James takes a nuanced stance, pointing out that AI amplifies bad architecture alongside good output, and details a deep dive into Brex's graph of audit and review agents collaborating on expense fraud.2:36–5:10 · Guest teaching 5/10 Hiring Ex-Founders and Brex's 'Quitters Welcome' Philosophy Swyx and Alessio challenge the trendy practice of hiring ex-founders, questioning whether founder tenure is too short or an anti-signal. James acknowledges their concern but counters with Brex's deliberate 'Quitters Welcome' philosophy and how instant distribution attracts top talent.5:11–7:33 · Guest teaching 5/10 Engineering Organization and the Dedicated AI Center of Excellence Alessio asks about engineering organization structure and AI adoption disparities across teams. James explains their product domain divisions and the deliberate creation of a centralized 10-person AI center of excellence designed like an internal disruptor startup.7:34–11:40 · Guest teaching 6/10 Pod Composition and Managing Internal AI Perceptions Alessio raises the common cultural problem of non-AI engineers feeling alienated or unvalued. James reframes this, explaining that because Brex engineering heavily rewards revenue impact, core product engineers remain proud of driving direct business metrics over speculative AI projects.11:40–15:55 · Guest teaching 6/10 Brex Agent Platform Architecture, TypeScript, and Mastra Swyx presses on Brex's choice of Mastra over established tooling like LangChain. James defends the decision by noting ergonomics and historical limitations of early LangChain before explaining their greenfield TypeScript stack.15:55–20:31 · Guest teaching 7/10 Multi-Agent Architecture and the Brex Assistant James gives an in-depth breakdown of Brex's multi-agent architecture for the Brex Assistant, explaining how single-agent tool overloading failed and why encapsulated, inter-agent direct messaging modeled after human org charts worked better.20:31–22:44 · Guest teaching 6/10 Model Context Protocol vs. Multi-Turn Agent Conversations Alessio and James discuss whether Anthropic's Model Context Protocol (MCP) should govern agent-to-agent communication. James clarifies that MCP suits single-turn imperative tool calls, whereas multi-agent collaboration requires conversational, multi-turn dialogue.22:44–27:32 · Guest teaching 7/10 Deep Dive into Brex's Three AI Strategic Pillars James details Brex's three AI strategic pillars (corporate, operational, product). He explains how operational AI directly slashed costs in compliance and underwriting while transforming support staff into prompt and eval authors.27:32–31:23 · Guest teaching 6/10 ConductorOne Self-Service Tooling and Model Flexibility Swyx expresses enthusiasm for Brex's ConductorOne Slack-based model provisioning. James explains their philosophy of avoiding vendor lock-in by letting employees choose models dynamically and using usage metrics in enterprise renewals.31:23–33:58 · Guest teaching 5/10 Managing Code Slop, Review Rigor, and Shared Understanding Swyx challenges the idea that AI review tools alone can fix code slop, arguing human ownership must scale. James agrees with the downside of codebase drift and unrigorous reviews while Swyx pushes back on tool vendors claiming automated cures.33:59–37:10 · Guest teaching 6/10 The Evolution of Engineering Craft and Junior Workflows James shares findings from a college dinner where new grads used LLMs for design docs and architecture rather than blind code generation. Alessio and James discuss how junior vs senior workflows differ in design and schema specification.37:11–40:00 · Guest teaching 5/10 Code Consistency Rules, Greptile, and Semantic Review in CI Alessio brings up historical parallels like Danger Systems for semantic linting in CI. James validates this by detailing their adoption of Greptile and automated Claude Code review checks in GitHub Actions.40:01–45:06 · Guest teaching 8/10 Operational Lessons: From Reinforcement Learning to SOP Agents James shares a key technical failure and pivot: Brex invested heavily in reinforcement learning for credit underwriting, only to find simple web research agents and prompt-based SOPs vastly outperformed complex RL models.45:06–48:53 · Guest teaching 6/10 Ideal Customer Profiles and Retool-Driven Prompt Iteration Swyx presses James on whether Brex's expanded ICP represents a retreat back to the SMB segment. James clarifies the specific revenue and transaction thresholds defining their commercial segment, and highlights their Retool-based prompt management tooling.48:53–52:13 · Guest teaching 7/10 Knowledge Base Grounding and the Decision to Partner with Sierra Swyx questions why Brex bought Sierra instead of building customer support agents in-house. James explains that building low-code management UIs and domain telemetry for CX leaders is undifferentiated work best outsourced.52:14–58:51 · Guest teaching 6/10 Multi-Turn Evals, User Simulations, and Hallucination Mitigations Alessio suggests using forward-looking, intentionally failing evals to track model progression over time. James enthusiastically adopts the idea and explains Brex's multi-turn eval techniques and prompt guardrails against phantom agent delegation.58:51–1:03:33 · Guest teaching 7/10 AI Fluency Framework, Upskilling, and Agentic Coding Interviews James explains Brex's AI fluency framework and reveals that Brex instituted mandatory agentic coding re-interviews for all existing engineers and managers to force hands-on exposure and accelerate upskilling.1:03:33–1:07:24 · Guest teaching 7/10 Headcount Planning, Engineering Leverage, and Industry Headwinds Swyx probes on whether AI productivity is causing layoffs. James takes a nuanced stance, pointing out that AI amplifies bad architecture alongside good output, and details a deep dive into Brex's graph of audit and review agents collaborating on expense fraud.2:36–5:10 · Guest disagreement 2/10 Hiring Ex-Founders and Brex's 'Quitters Welcome' Philosophy Swyx and Alessio challenge the trendy practice of hiring ex-founders, questioning whether founder tenure is too short or an anti-signal. James acknowledges their concern but counters with Brex's deliberate 'Quitters Welcome' philosophy and how instant distribution attracts top talent.5:11–7:33 · Guest disagreement 1/10 Engineering Organization and the Dedicated AI Center of Excellence Alessio asks about engineering organization structure and AI adoption disparities across teams. James explains their product domain divisions and the deliberate creation of a centralized 10-person AI center of excellence designed like an internal disruptor startup.7:34–11:40 · Guest disagreement 2/10 Pod Composition and Managing Internal AI Perceptions Alessio raises the common cultural problem of non-AI engineers feeling alienated or unvalued. James reframes this, explaining that because Brex engineering heavily rewards revenue impact, core product engineers remain proud of driving direct business metrics over speculative AI projects.11:40–15:55 · Guest disagreement 3/10 Brex Agent Platform Architecture, TypeScript, and Mastra Swyx presses on Brex's choice of Mastra over established tooling like LangChain. James defends the decision by noting ergonomics and historical limitations of early LangChain before explaining their greenfield TypeScript stack.15:55–20:31 · Guest disagreement 1/10 Multi-Agent Architecture and the Brex Assistant James gives an in-depth breakdown of Brex's multi-agent architecture for the Brex Assistant, explaining how single-agent tool overloading failed and why encapsulated, inter-agent direct messaging modeled after human org charts worked better.20:31–22:44 · Guest disagreement 2/10 Model Context Protocol vs. Multi-Turn Agent Conversations Alessio and James discuss whether Anthropic's Model Context Protocol (MCP) should govern agent-to-agent communication. James clarifies that MCP suits single-turn imperative tool calls, whereas multi-agent collaboration requires conversational, multi-turn dialogue.22:44–27:32 · Guest disagreement 1/10 Deep Dive into Brex's Three AI Strategic Pillars James details Brex's three AI strategic pillars (corporate, operational, product). He explains how operational AI directly slashed costs in compliance and underwriting while transforming support staff into prompt and eval authors.27:32–31:23 · Guest disagreement 1/10 ConductorOne Self-Service Tooling and Model Flexibility Swyx expresses enthusiasm for Brex's ConductorOne Slack-based model provisioning. James explains their philosophy of avoiding vendor lock-in by letting employees choose models dynamically and using usage metrics in enterprise renewals.31:23–33:58 · Guest disagreement 2/10 Managing Code Slop, Review Rigor, and Shared Understanding Swyx challenges the idea that AI review tools alone can fix code slop, arguing human ownership must scale. James agrees with the downside of codebase drift and unrigorous reviews while Swyx pushes back on tool vendors claiming automated cures.33:59–37:10 · Guest disagreement 1/10 The Evolution of Engineering Craft and Junior Workflows James shares findings from a college dinner where new grads used LLMs for design docs and architecture rather than blind code generation. Alessio and James discuss how junior vs senior workflows differ in design and schema specification.37:11–40:00 · Guest disagreement 1/10 Code Consistency Rules, Greptile, and Semantic Review in CI Alessio brings up historical parallels like Danger Systems for semantic linting in CI. James validates this by detailing their adoption of Greptile and automated Claude Code review checks in GitHub Actions.40:01–45:06 · Guest disagreement 1/10 Operational Lessons: From Reinforcement Learning to SOP Agents James shares a key technical failure and pivot: Brex invested heavily in reinforcement learning for credit underwriting, only to find simple web research agents and prompt-based SOPs vastly outperformed complex RL models.45:06–48:53 · Guest disagreement 2/10 Ideal Customer Profiles and Retool-Driven Prompt Iteration Swyx presses James on whether Brex's expanded ICP represents a retreat back to the SMB segment. James clarifies the specific revenue and transaction thresholds defining their commercial segment, and highlights their Retool-based prompt management tooling.48:53–52:13 · Guest disagreement 2/10 Knowledge Base Grounding and the Decision to Partner with Sierra Swyx questions why Brex bought Sierra instead of building customer support agents in-house. James explains that building low-code management UIs and domain telemetry for CX leaders is undifferentiated work best outsourced.52:14–58:51 · Guest disagreement 1/10 Multi-Turn Evals, User Simulations, and Hallucination Mitigations Alessio suggests using forward-looking, intentionally failing evals to track model progression over time. James enthusiastically adopts the idea and explains Brex's multi-turn eval techniques and prompt guardrails against phantom agent delegation.58:51–1:03:33 · Guest disagreement 1/10 AI Fluency Framework, Upskilling, and Agentic Coding Interviews James explains Brex's AI fluency framework and reveals that Brex instituted mandatory agentic coding re-interviews for all existing engineers and managers to force hands-on exposure and accelerate upskilling.1:03:33–1:07:24 · Guest disagreement 2/10 Headcount Planning, Engineering Leverage, and Industry Headwinds Swyx probes on whether AI productivity is causing layoffs. James takes a nuanced stance, pointing out that AI amplifies bad architecture alongside good output, and details a deep dive into Brex's graph of audit and review agents collaborating on expense fraud.2:36–5:10 · The hosts pushing back 4/10 Hiring Ex-Founders and Brex's 'Quitters Welcome' Philosophy Swyx and Alessio challenge the trendy practice of hiring ex-founders, questioning whether founder tenure is too short or an anti-signal. James acknowledges their concern but counters with Brex's deliberate 'Quitters Welcome' philosophy and how instant distribution attracts top talent.5:11–7:33 · The hosts pushing back 2/10 Engineering Organization and the Dedicated AI Center of Excellence Alessio asks about engineering organization structure and AI adoption disparities across teams. James explains their product domain divisions and the deliberate creation of a centralized 10-person AI center of excellence designed like an internal disruptor startup.7:34–11:40 · The hosts pushing back 3/10 Pod Composition and Managing Internal AI Perceptions Alessio raises the common cultural problem of non-AI engineers feeling alienated or unvalued. James reframes this, explaining that because Brex engineering heavily rewards revenue impact, core product engineers remain proud of driving direct business metrics over speculative AI projects.11:40–15:55 · The hosts pushing back 3/10 Brex Agent Platform Architecture, TypeScript, and Mastra Swyx presses on Brex's choice of Mastra over established tooling like LangChain. James defends the decision by noting ergonomics and historical limitations of early LangChain before explaining their greenfield TypeScript stack.15:55–20:31 · The hosts pushing back 2/10 Multi-Agent Architecture and the Brex Assistant James gives an in-depth breakdown of Brex's multi-agent architecture for the Brex Assistant, explaining how single-agent tool overloading failed and why encapsulated, inter-agent direct messaging modeled after human org charts worked better.20:31–22:44 · The hosts pushing back 3/10 Model Context Protocol vs. Multi-Turn Agent Conversations Alessio and James discuss whether Anthropic's Model Context Protocol (MCP) should govern agent-to-agent communication. James clarifies that MCP suits single-turn imperative tool calls, whereas multi-agent collaboration requires conversational, multi-turn dialogue.22:44–27:32 · The hosts pushing back 1/10 Deep Dive into Brex's Three AI Strategic Pillars James details Brex's three AI strategic pillars (corporate, operational, product). He explains how operational AI directly slashed costs in compliance and underwriting while transforming support staff into prompt and eval authors.27:32–31:23 · The hosts pushing back 1/10 ConductorOne Self-Service Tooling and Model Flexibility Swyx expresses enthusiasm for Brex's ConductorOne Slack-based model provisioning. James explains their philosophy of avoiding vendor lock-in by letting employees choose models dynamically and using usage metrics in enterprise renewals.31:23–33:58 · The hosts pushing back 4/10 Managing Code Slop, Review Rigor, and Shared Understanding Swyx challenges the idea that AI review tools alone can fix code slop, arguing human ownership must scale. James agrees with the downside of codebase drift and unrigorous reviews while Swyx pushes back on tool vendors claiming automated cures.33:59–37:10 · The hosts pushing back 2/10 The Evolution of Engineering Craft and Junior Workflows James shares findings from a college dinner where new grads used LLMs for design docs and architecture rather than blind code generation. Alessio and James discuss how junior vs senior workflows differ in design and schema specification.37:11–40:00 · The hosts pushing back 2/10 Code Consistency Rules, Greptile, and Semantic Review in CI Alessio brings up historical parallels like Danger Systems for semantic linting in CI. James validates this by detailing their adoption of Greptile and automated Claude Code review checks in GitHub Actions.40:01–45:06 · The hosts pushing back 1/10 Operational Lessons: From Reinforcement Learning to SOP Agents James shares a key technical failure and pivot: Brex invested heavily in reinforcement learning for credit underwriting, only to find simple web research agents and prompt-based SOPs vastly outperformed complex RL models.45:06–48:53 · The hosts pushing back 3/10 Ideal Customer Profiles and Retool-Driven Prompt Iteration Swyx presses James on whether Brex's expanded ICP represents a retreat back to the SMB segment. James clarifies the specific revenue and transaction thresholds defining their commercial segment, and highlights their Retool-based prompt management tooling.48:53–52:13 · The hosts pushing back 3/10 Knowledge Base Grounding and the Decision to Partner with Sierra Swyx questions why Brex bought Sierra instead of building customer support agents in-house. James explains that building low-code management UIs and domain telemetry for CX leaders is undifferentiated work best outsourced.52:14–58:51 · The hosts pushing back 2/10 Multi-Turn Evals, User Simulations, and Hallucination Mitigations Alessio suggests using forward-looking, intentionally failing evals to track model progression over time. James enthusiastically adopts the idea and explains Brex's multi-turn eval techniques and prompt guardrails against phantom agent delegation.58:51–1:03:33 · The hosts pushing back 1/10 AI Fluency Framework, Upskilling, and Agentic Coding Interviews James explains Brex's AI fluency framework and reveals that Brex instituted mandatory agentic coding re-interviews for all existing engineers and managers to force hands-on exposure and accelerate upskilling.1:03:33–1:07:24 · The hosts pushing back 3/10 Headcount Planning, Engineering Leverage, and Industry Headwinds Swyx probes on whether AI productivity is causing layoffs. James takes a nuanced stance, pointing out that AI amplifies bad architecture alongside good output, and details a deep dive into Brex's graph of audit and review agents collaborating on expense fraud.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 44.5% · guest 55.5%0:00 · the hosts 44.5% · guest 55.5%3:00 · the hosts 15.5% · guest 84.5%3:00 · the hosts 15.5% · guest 84.5%6:00 · the hosts 27.7% · guest 72.3%6:00 · the hosts 27.7% · guest 72.3%9:00 · the hosts 22.5% · guest 77.5%9:00 · the hosts 22.5% · guest 77.5%12:00 · the hosts 2.1% · guest 97.9%12:00 · the hosts 2.1% · guest 97.9%15:00 · the hosts 4.3% · guest 95.7%15:00 · the hosts 4.3% · guest 95.7%18:00 · the hosts 2.3% · guest 97.7%18:00 · the hosts 2.3% · guest 97.7%21:00 · the hosts 20.5% · guest 79.5%21:00 · the hosts 20.5% · guest 79.5%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 6.1% · guest 93.9%27:00 · the hosts 6.1% · guest 93.9%30:00 · the hosts 11.7% · guest 88.3%30:00 · the hosts 11.7% · guest 88.3%33:00 · the hosts 26% · guest 74%33:00 · the hosts 26% · guest 74%36:00 · the hosts 60.1% · guest 39.9%36:00 · the hosts 60.1% · guest 39.9%39:00 · the hosts 14.2% · guest 85.8%39:00 · the hosts 14.2% · guest 85.8%42:00 · the hosts 6.6% · guest 93.4%42:00 · the hosts 6.6% · guest 93.4%45:00 · the hosts 4.7% · guest 95.3%45:00 · the hosts 4.7% · guest 95.3%48:00 · the hosts 22.2% · guest 77.8%48:00 · the hosts 22.2% · guest 77.8%51:00 · the hosts 6.7% · guest 93.3%51:00 · the hosts 6.7% · guest 93.3%54:00 · the hosts 19.7% · guest 80.3%54:00 · the hosts 19.7% · guest 80.3%57:00 · the hosts 26% · guest 74%57:00 · the hosts 26% · guest 74%1:00:00 · the hosts 0.3% · guest 99.7%1:00:00 · the hosts 0.3% · guest 99.7%1:03:00 · the hosts 9.5% · guest 90.5%1:03:00 · the hosts 9.5% · guest 90.5%1:06:00 · the hosts 27.9% · guest 72.1%1:06:00 · the hosts 27.9% · guest 72.1%1:09:00 · the hosts 11.6% · guest 88.4%1:09:00 · the hosts 11.6% · guest 88.4%1:12:00 · the hosts 15.9% · guest 84.1%1:12:00 · the hosts 15.9% · guest 84.1%
Sharpest disagreement ▶ 14:58 Defending against LangChain superiority

James pushes back on Swyx's defense of LangChain, maintaining that early framework shortcomings forced Brex to build proprietary tooling and switch to Mastra.

Hardest push from the hosts ▶ 32:53 Challenging AI code reviewer claims

Swyx firmly rejects the notion that companies can solve code slop by simply layering more automated AI reviewers onto AI-generated code.

Biggest teaching moment ▶ 40:26 RL failure versus simple SOP agentic workflows

James educates the hosts on how their expensive bet on reinforcement learning failed compared to simple web research agents executing granular SOP prompts.

The host holds their own ▶ 56:00 Suggesting saturation tracking through failing evals

Alessio leverages his experience with Baris AI to introduce a novel paradigm for tracking model capability progression, which the guest enthusiastically adopts.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Hiring Ex-Founders and Brex's 'Quitters Welcome' Philosophy 6524 Swyx and Alessio challenge the trendy practice of hiring ex-founders, questioning whether founder tenure is too short or an anti-signal. James acknowledges their concern but counters with Brex's deliberate 'Quitters Welcome' philosophy and how instant distribution attracts top talent.
Engineering Organization and the Dedicated AI Center of Excellence 5512 Alessio asks about engineering organization structure and AI adoption disparities across teams. James explains their product domain divisions and the deliberate creation of a centralized 10-person AI center of excellence designed like an internal disruptor startup.
Pod Composition and Managing Internal AI Perceptions 5623 Alessio raises the common cultural problem of non-AI engineers feeling alienated or unvalued. James reframes this, explaining that because Brex engineering heavily rewards revenue impact, core product engineers remain proud of driving direct business metrics over speculative AI projects.
Brex Agent Platform Architecture, TypeScript, and Mastra 6633 Swyx presses on Brex's choice of Mastra over established tooling like LangChain. James defends the decision by noting ergonomics and historical limitations of early LangChain before explaining their greenfield TypeScript stack.
Multi-Agent Architecture and the Brex Assistant 6712 James gives an in-depth breakdown of Brex's multi-agent architecture for the Brex Assistant, explaining how single-agent tool overloading failed and why encapsulated, inter-agent direct messaging modeled after human org charts worked better.
Model Context Protocol vs. Multi-Turn Agent Conversations 7623 Alessio and James discuss whether Anthropic's Model Context Protocol (MCP) should govern agent-to-agent communication. James clarifies that MCP suits single-turn imperative tool calls, whereas multi-agent collaboration requires conversational, multi-turn dialogue.
Deep Dive into Brex's Three AI Strategic Pillars 5711 James details Brex's three AI strategic pillars (corporate, operational, product). He explains how operational AI directly slashed costs in compliance and underwriting while transforming support staff into prompt and eval authors.
ConductorOne Self-Service Tooling and Model Flexibility 6611 Swyx expresses enthusiasm for Brex's ConductorOne Slack-based model provisioning. James explains their philosophy of avoiding vendor lock-in by letting employees choose models dynamically and using usage metrics in enterprise renewals.
Managing Code Slop, Review Rigor, and Shared Understanding 7524 Swyx challenges the idea that AI review tools alone can fix code slop, arguing human ownership must scale. James agrees with the downside of codebase drift and unrigorous reviews while Swyx pushes back on tool vendors claiming automated cures.
The Evolution of Engineering Craft and Junior Workflows 6612 James shares findings from a college dinner where new grads used LLMs for design docs and architecture rather than blind code generation. Alessio and James discuss how junior vs senior workflows differ in design and schema specification.
Code Consistency Rules, Greptile, and Semantic Review in CI 7512 Alessio brings up historical parallels like Danger Systems for semantic linting in CI. James validates this by detailing their adoption of Greptile and automated Claude Code review checks in GitHub Actions.
Operational Lessons: From Reinforcement Learning to SOP Agents 5811 James shares a key technical failure and pivot: Brex invested heavily in reinforcement learning for credit underwriting, only to find simple web research agents and prompt-based SOPs vastly outperformed complex RL models.
Ideal Customer Profiles and Retool-Driven Prompt Iteration 6623 Swyx presses James on whether Brex's expanded ICP represents a retreat back to the SMB segment. James clarifies the specific revenue and transaction thresholds defining their commercial segment, and highlights their Retool-based prompt management tooling.
Knowledge Base Grounding and the Decision to Partner with Sierra 6723 Swyx questions why Brex bought Sierra instead of building customer support agents in-house. James explains that building low-code management UIs and domain telemetry for CX leaders is undifferentiated work best outsourced.
Multi-Turn Evals, User Simulations, and Hallucination Mitigations 7612 Alessio suggests using forward-looking, intentionally failing evals to track model progression over time. James enthusiastically adopts the idea and explains Brex's multi-turn eval techniques and prompt guardrails against phantom agent delegation.
AI Fluency Framework, Upskilling, and Agentic Coding Interviews 5711 James explains Brex's AI fluency framework and reveals that Brex instituted mandatory agentic coding re-interviews for all existing engineers and managers to force hands-on exposure and accelerate upskilling.
Headcount Planning, Engineering Leverage, and Industry Headwinds 6723 Swyx probes on whether AI productivity is causing layoffs. James takes a nuanced stance, pointing out that AI amplifies bad architecture alongside good output, and details a deep dive into Brex's graph of audit and review agents collaborating on expense fraud.

Statements from this episode (39)

Disclosure
Reggio: Brex structures its AI strategy into corporate, operational, and product pillars
“We have like three pillars to our AI strategy. We have our corporate AI strategy, which is how are we going to adopt and like buy AI tooling across the business and basically every single function to be able to 10 X our workflows. And we have our operational …”
James Reggio Jan 17, 2026 ▶ 24:32
Disclosure
Brex launches 'Quitters Welcome' recruiting pitch encouraging future founders
“We actually launched sort of like a new recruiting and employee value proposition for Brex a couple months ago called Quitters Welcome, where we actually intentionally are leaning into this idea that we have a disproportionate number of folks who go on to beco…”
James Reggio Jan 17, 2026 ▶ 3:29
Assertion Partly supported
Reggio: Brex has roughly 40,000 customers from Fortune 100 to startups
“You can come into this business And build like financial AI applications and instantly have that deployed to roughly 40,000 customers across you know, the fortune 100 down to, you know, tens of thousands of startups.”
James Reggio Jan 17, 2026 ▶ 4:38
Disclosure
Brex formed a 10-person internal AI team to disrupt its product
“And AI is one of those areas where we have another team of Just roughly about 10 people who are focused primarily on LLM applications, and we wanted to create a bit of a separation there because the way that we were thinking about this, and this is actually so…”
James Reggio Jan 17, 2026 ▶ 6:07
Disclosure
Reggio: Cursor adoption is uniform across Brex engineering
“It's actually fairly, fairly uniform across the entire engineering department. It's actually kind of funny, like one of our largest cursor users is actually an engineering manager.”
James Reggio Jan 17, 2026 ▶ 7:01
Insight
Reggio: Deep Domain Experience Can Impede AI-First Problem Solving
“I think part of it is like sometimes the too much experience or too much knowledge of how to solve a problem and actually be an impediment to thinking differently about it and thinking about it from like an AI first lens.”
James Reggio Jan 17, 2026 ▶ 9:01
Assertion Not checkable as stated
Reggio: Brex Card Product Drives 60% of Direct Revenue
“The folks who say we're kind of a card product, they drive 60% of our direct revenue.”
James Reggio Jan 17, 2026 ▶ 10:59
Assertion Not checkable as stated
Brex replaced human KYC underwriting with automated web research agents
“We've we set up a completely automated Pipeline for evaluating customer applications to get them onboarded instantly to Brex, which is something that used to require human intervention either for underwriting or KYC. But now we basically have a series of agent…”
James Reggio Jan 17, 2026 ▶ 12:49
Insight
Reggio: Agentic coding has significantly shortened the half-life of code
“Because the half-life of code has declined so significantly with agentic coding. It's actually quite, ah, easy for us and for anyone else to kind of try on for size, a variety of different pieces of tech to figure out what is going to be most ergonomic for sol…”
James Reggio Jan 17, 2026 ▶ 13:51
Disclosure
Reggio: Half of Brex agent applications run on Mastra, half internally
“About half of the applications that we're building right now on the agent layer are running on Mastra and then the other half are actually still running on like yet another internally developed framework. Which is a framework that's focused more on networks of…”
James Reggio Jan 17, 2026 ▶ 15:26
Insight
Reggio: The Ideal Brex UX Is Just Swiping the Card
“Like the best UI UX for Brex is just the card. Like every single thing that you have to do in the software beyond just swiping the card is like an opportunity for AI to eliminate some work for you.”
James Reggio Jan 17, 2026 ▶ 17:07
Disclosure
Brex built a custom framework where sub-agents direct message each other
“We kind of just hit the eject button and built our own framework, which is one in which we have agents that are able to basically DM with other agents and have multi-term conversations amongst themselves to coordinate to complete a task to or like to complete …”
James Reggio Jan 17, 2026 ▶ 18:43
Insight
Reggio: Multi-turn agent interaction beats treating sub-agents as single tool calls
“There's actually a lot of value in having like multi-turn conversations from like the orchestrator or the assistant to like the sub agent, whereas like, you know, a tool call is basically just like one RPC.”
James Reggio Jan 17, 2026 ▶ 20:36
Insight
Reggio: MCP belongs on conventional systems, not inter-agent communication
“I think of like the MCP and tool usage as being like the interface to all of our conventional imperative systems, not at the AI space.”
James Reggio Jan 17, 2026 ▶ 21:40
Insight
Reggio: Multi-agent systems work best when structured like human org charts
“It's actually helpful for us to basically conceive of it as an org chart. And like, it's the agent org chart with you know, my EA is DMing other specialists and having brief conversations to support me as their client.”
James Reggio Jan 17, 2026 ▶ 22:30
Disclosure
Brex transitions operations staff from executing SOPs to building AI prompts
“She's been leaning really aggressively to help every member of the operations organization start rethinking their role as being not people who kind of execute against an SOP, but are people who are going to like build prompts, build evals, and like be, become …”
James Reggio Jan 17, 2026 ▶ 26:52
Disclosure
Reggio: Brex targets 80% automated application decisions within 60 seconds
“Like we, as soon as we put our mind to it, we said like, look, no, we want to hit 80% automated acceptance rate for all all startup and commercial businesses that apply for Brexit. Like we want a decision within 60 seconds that's fully touchless, no humans inv…”
James Reggio Jan 17, 2026 ▶ 28:05
Disclosure
Reggio: Brex does not pick single AI tool winners internally
“We're not going to try to pick winners in the horse race between the foundational model providers or the agentic coding tools, or like basically anywhere where there's an active horse race. What we do and instead of like trying to pick a single solution is we …”
James Reggio Jan 17, 2026 ▶ 28:42
Insight
Reggio: Rapid AI code generation increases incident response risks from knowledge drift
“As well as, like, maybe one of the other facets of being able to generate a lot more code more quickly is, like, the drift between team members as far as, like, understanding of the code that's in their services increases is like everybody's moving faster and …”
James Reggio Jan 17, 2026 ▶ 32:24
Prediction Not checkable as stated
Swix: AI productivity expectations will force engineers to own far more code
“Every engineer is just gonna own more code. Period. And be parachuted in and be expected to ramp up and be productive and also fix bugs. And if you're on, you know, pager duty or whatever to just because I mean, everyone's going to try to be more efficient and…”
Shawn Wang Jan 17, 2026 ▶ 33:13
Disclosure
Reggio: Took one-month leave to code full-time with Brex AI team
“Basically went on leave for a month and join the join the team that the AI team that we were building just to go and build alongside them.”
James Reggio Jan 17, 2026 ▶ 34:14
Assertion Not checkable as stated
Reggio: CS students use AI for design docs rather than raw code generation
“And I was surprised to hear the consensus was that most people there were using agents to collaborate on like building a design document and like, Collaborating on the architecture of the solution that they want to build, and then they'd be asking it to like e…”
James Reggio Jan 17, 2026 ▶ 35:26
Opinion
Reggio: Brex relies on Greptile for agentic code review
“We're big fans of creptile and we use them for basically all of sort of the Smarter than linting, ah, like, agentic code review, ah, that's been the one solution that we have aligned around that has served us extremely well.”
James Reggio Jan 17, 2026 ▶ 37:50
Disclosure
Reggio: Brex chose full TypeScript over Kotlin/Elixir for agent codebase
“When we started building this new agent code base, like, as we were saying, like, we were answering the question, what would you do if you built a, you know, a Brex disruptor today? And it's like, it wouldn't be to pick Kotlin and Elixir as the backend. And so…”
James Reggio Jan 17, 2026 ▶ 38:58
Assertion Not checkable as stated
Reggio: Simple web research agents outperformed RL for credit underwriting at Brex
“We made this big investment. We were working with some outside, like the, like a company that specializes in this and the performance we ended up getting was inferior to just building a, like a web research agent.”
James Reggio Jan 17, 2026 ▶ 40:58
Insight
Reggio: The real challenge in operational AI is codifying unwritten SOPs into prompts
“The challenge is really articulating and refining prompts to reflect. Reflect the execution of the SOP and like reflect all the sort of institutional knowledge that isn't written down so that agents can properly replace like the humans or the contractors that …”
James Reggio Jan 17, 2026 ▶ 42:15
Disclosure
Reggio: Tens of thousands of ROI-negative SMB customers almost threatened Brex
“It ended up being a huge burden for the business almost existential for us to have those tens of thousands of customers that all were ROI negative.”
James Reggio Jan 17, 2026 ▶ 46:07
Disclosure
Reggio: Brex sets minimum ICP at $1M revenue or $10K monthly card spend
“Right now our minimum threshold is, is like a million dollars a year in recurring in, in annual revenue or like 10,000 dollars or more per month in, in card transactions as kind of being like the low end of our ICP”
James Reggio Jan 17, 2026 ▶ 46:27
Disclosure
Reggio: Brex built its internal operational AI management platform in Retool
“Most of the UI UX for this platform is built in retool. And so like you can basically go into retool and there's like a prompt manager, a tool manager, an eval manager.”
James Reggio Jan 17, 2026 ▶ 47:25
Opinion
Reggio: Brex uses Sierra because CX agent administration is not differentiated
“That's like solving problems that are not differentiated enough for us. I think what, what's interesting about the Sierra that has been really helpful is that again, it's really easy for like the UI and UX of basically administering a Sierra agent is something…”
James Reggio Jan 17, 2026 ▶ 51:23
Disclosure
Brex evaluates multi-agent networks using simulated user agents and LLM judges
“And so what we do there is we try to adopt some of the state of the art for multi-turn evals where we will we'll basically have a, an agent embody the user and like, you know, have basically the end user agent is given an objective and then we basically have i…”
James Reggio Jan 17, 2026 ▶ 53:13
Disclosure
Reggio: Brex built LLM circuit breakers but does not use them
“Like I, we kind of built a couple of those circuit breakers or like the ability to put those circuit breakers in. And I don't believe we're using them for anything.”
James Reggio Jan 17, 2026 ▶ 58:22
Disclosure
Reggio: Brex gives spot bonuses and biweekly spotlights for AI workflows
“We'll do like spot bonuses for people who have like particularly novel uses of AI on, in their day to day. In our company, all hands, every two weeks, we'll do an AI spotlight. And it's very rarely somebody in EPD for the most part is folks in ETMs, ops financ…”
James Reggio Jan 17, 2026 ▶ 1:00:55
Disclosure
Reggio: Brex engineering interview revamp makes agentic coding mandatory
“We adapted our interview loop to be more AI sort of agentic coding native. So instead of we had like a coding and a system design question that we basically have revamped into a project where we'll give you like a brief, Before you come on site and then like a…”
James Reggio Jan 17, 2026 ▶ 1:01:43
Disclosure
Brex forced all existing engineers to pass its new AI interview
“As soon as we had the interview ready to ship, we started, we said everybody in engineering, including all the managers are going to have to go through this interview. And so we re-interviewed everybody internally.”
James Reggio Jan 17, 2026 ▶ 1:02:29
Insight
Reggio: Agentic development amplifies bad architecture, yielding less net capacity increase
“I view agentic development as being something that amplifies all the good, just as much as it amplifies all the bad. And it amplifies sloppiness, poor architectural thinking misunderstanding of the requirements. Like there are, for all of the acceleration of g…”
James Reggio Jan 17, 2026 ▶ 1:04:27
Disclosure
Brex will freeze engineering headcount at 300 while targeting doubled efficiency
“I like having 300 engineers. Like I would love, love to just, you know, a year from now have 300 engineers, but we're still, you know, a 30, 50, a hundred percent more efficient.”
James Reggio Jan 17, 2026 ▶ 1:05:47
Insight
Reggio: Constraining LLMs to deterministic DAGs undersells their planning power
“My intuition has been that trying to craft LLMs into deterministic workflows and DAGs is, is kind of underselling like the power that they have to actually plan and execute more in a more sophisticated, like fluid way.”
James Reggio Jan 17, 2026 ▶ 1:08:04
Disclosure
Brex built two-tier audit agents to flag and filter expense policy violations
“We built this audit agent that can, like, ingest your SOP and look, also ingest your. And what it does then is it's basically always looking for potential violations. And what it does is it is extremely zealous. Like it wants to have a minimum number of false …”
James Reggio Jan 17, 2026 ▶ 1:10:40
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.