Aug 7, 2026 · 58m · neon-show

The Billion Dollar AI Lab Founder Who Sees The Future First | Karan Goel, Founder & CEO of Cartesia

Karan Goel · 44m spoken Siddhartha Ahluwalia · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The Neon Show, Cartesia founder and CEO Karan Goel details the journey of spinning a frontier real-time AI lab out of Stanford University. He discusses the full-stack engineering required for ultra-low-latency multimodal models, Cartesia's lean product-led growth strategy, and the immense future market for enterprise voice agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

Siddhartha as informed peer 3.9 Guest teaching 4.5 Guest disagreement 0.4 Siddhartha pushing back 0.2
05100:0015:0030:0045:001:46–5:54 · Siddhartha as informed peer 3/10 Gaming Origins, Hardware Assembly, and AI Curiosity Karan reflects on assembling GPUs and gaming online in Delhi, which spurred his technical research journey. Siddharth facilitates the personal backstory with nostalgic rapport.5:55–9:03 · Siddhartha as informed peer 3/10 Cartesia's Mathematical Homage and Stanford Co-Founding Team Karan explains how Cartesia was named after René Descartes and describes spinning out of Chris Ré's Stanford AI lab with fellow graduate students. Siddharth inquires about the team dynamics.9:03–11:34 · Siddhartha as informed peer 4/10 Pre-Seed Fundraising and Pivoting to Real-Time Interactive AI Karan recounts shifting Cartesia from an exploratory pre-seed lab to targeting continuous real-time interactive intelligence. Siddharth draws a comparison to the interactive OS in the film 'Her'.11:34–14:14 · Siddhartha as informed peer 4/10 Prioritizing Voice as the Primary Synchronous Interface Karan outlines why voice is the primary modality for synchronous communication and clarifies that the founders approached speech from deep sequence modeling rather than traditional speech research.14:16–17:14 · Siddhartha as informed peer 3/10 Model Lab Infrastructure, Experimentation Culture, and Audio Evaluation Karan explains the complex infrastructure required to run a model lab and notes how subjective, unstandardized audio evaluation benchmarks remain a massive hurdle.17:18–20:04 · Siddhartha as informed peer 3/10 Full-Stack System Co-Design and Pareto Dominance Karan defines full-stack co-design and rejects the standard engineering trade-off between speed, cost, and quality in favor of Pareto dominance.20:05–23:21 · Siddhartha as informed peer 3/10 Cartesia's Key Milestones: Proprietary Inference Engine and Voice Agents Karan describes building Cartesia's proprietary inference engine specifically tailored for real-time interactive models rather than standard LLMs.23:22–26:21 · Siddhartha as informed peer 6/10 Overcoming Long-Horizon Context and Memory in Voice AI Siddharth raises the issue of maintaining context in long conversations. Karan reveals that current voice agents fake long conversations using engineered short turns, explaining that true multimodal reasoning remains unsolved.26:21–30:23 · Siddhartha as informed peer 5/10 Hyper-Personalization and Proactive Customer Engagement Siddharth illustrates a relationship-driven banking scenario. Karan expands on how low-cost AI agents enable proactive customer service during dead air, such as telephone hold times.30:28–33:55 · Siddhartha as informed peer 3/10 Platform Architecture, Production Scale, and Voice Research Karan outlines the architectural demands of production scale and describes Cartesia's specialized research team investigating how voice tone influences business metrics.33:56–36:05 · Siddhartha as informed peer 6/10 The Psychology of Latency: Voice Versus Text Interaction Siddharth offers an astute observation comparing the strict latency demands of voice against user tolerance for text delay. Karan agrees and explains how instant text can paradoxically degrade perceived intelligence.36:07–40:40 · Siddhartha as informed peer 6/10 Enterprise Adoption Across Regulated Industries and Internal Workflows Siddharth challenges Karan on whether foundation model labs are eating the lunch of application developers. Karan reframes the tension as natural information asymmetry rather than predatory competition.40:40–44:13 · Siddhartha as informed peer 3/10 Evaluating Product-Market Fit and Scaling to Billion-Minute Volumes Karan shares indicators of product-market fit, highlighting enterprise customers scaling toward one billion minutes annually across multiple languages.44:15–47:18 · Siddhartha as informed peer 3/10 Lean Go-to-Market Strategy and Product-Led Expansion Karan outlines Cartesia's lean go-to-market structure, asserting that high marketing budgets often conceal product mediocrity and emphasizing self-serve developer discovery.47:19–50:04 · Siddhartha as informed peer 4/10 Core Pillars of Cartesia's Growth: Model Quality and Research-Product Synergy Karan dismisses the conventional assumption that companies must choose between cutting-edge research and commercial product delivery, arguing both can be built together with disciplined alignment.50:05–55:00 · Siddhartha as informed peer 4/10 Founder Alignment, Radical Candor, and High-Trust Culture Karan describes the founding team's high-trust, ego-free alignment and proposes a framework for sizing the voice AI market by multiplying interaction minutes by per-minute business value.55:00–57:21 · Siddhartha as informed peer 3/10 Venture Financing Success and Long-Term Value Creation Karan attributes Cartesia's fundraising success with tier-one venture firms to razor-sharp focus on real-time interactive intelligence and attracting world-class research talent.1:46–5:54 · Guest teaching 2/10 Gaming Origins, Hardware Assembly, and AI Curiosity Karan reflects on assembling GPUs and gaming online in Delhi, which spurred his technical research journey. Siddharth facilitates the personal backstory with nostalgic rapport.5:55–9:03 · Guest teaching 2/10 Cartesia's Mathematical Homage and Stanford Co-Founding Team Karan explains how Cartesia was named after René Descartes and describes spinning out of Chris Ré's Stanford AI lab with fellow graduate students. Siddharth inquires about the team dynamics.9:03–11:34 · Guest teaching 4/10 Pre-Seed Fundraising and Pivoting to Real-Time Interactive AI Karan recounts shifting Cartesia from an exploratory pre-seed lab to targeting continuous real-time interactive intelligence. Siddharth draws a comparison to the interactive OS in the film 'Her'.11:34–14:14 · Guest teaching 5/10 Prioritizing Voice as the Primary Synchronous Interface Karan outlines why voice is the primary modality for synchronous communication and clarifies that the founders approached speech from deep sequence modeling rather than traditional speech research.14:16–17:14 · Guest teaching 6/10 Model Lab Infrastructure, Experimentation Culture, and Audio Evaluation Karan explains the complex infrastructure required to run a model lab and notes how subjective, unstandardized audio evaluation benchmarks remain a massive hurdle.17:18–20:04 · Guest teaching 6/10 Full-Stack System Co-Design and Pareto Dominance Karan defines full-stack co-design and rejects the standard engineering trade-off between speed, cost, and quality in favor of Pareto dominance.20:05–23:21 · Guest teaching 5/10 Cartesia's Key Milestones: Proprietary Inference Engine and Voice Agents Karan describes building Cartesia's proprietary inference engine specifically tailored for real-time interactive models rather than standard LLMs.23:22–26:21 · Guest teaching 6/10 Overcoming Long-Horizon Context and Memory in Voice AI Siddharth raises the issue of maintaining context in long conversations. Karan reveals that current voice agents fake long conversations using engineered short turns, explaining that true multimodal reasoning remains unsolved.26:21–30:23 · Guest teaching 5/10 Hyper-Personalization and Proactive Customer Engagement Siddharth illustrates a relationship-driven banking scenario. Karan expands on how low-cost AI agents enable proactive customer service during dead air, such as telephone hold times.30:28–33:55 · Guest teaching 5/10 Platform Architecture, Production Scale, and Voice Research Karan outlines the architectural demands of production scale and describes Cartesia's specialized research team investigating how voice tone influences business metrics.33:56–36:05 · Guest teaching 5/10 The Psychology of Latency: Voice Versus Text Interaction Siddharth offers an astute observation comparing the strict latency demands of voice against user tolerance for text delay. Karan agrees and explains how instant text can paradoxically degrade perceived intelligence.36:07–40:40 · Guest teaching 5/10 Enterprise Adoption Across Regulated Industries and Internal Workflows Siddharth challenges Karan on whether foundation model labs are eating the lunch of application developers. Karan reframes the tension as natural information asymmetry rather than predatory competition.40:40–44:13 · Guest teaching 4/10 Evaluating Product-Market Fit and Scaling to Billion-Minute Volumes Karan shares indicators of product-market fit, highlighting enterprise customers scaling toward one billion minutes annually across multiple languages.44:15–47:18 · Guest teaching 4/10 Lean Go-to-Market Strategy and Product-Led Expansion Karan outlines Cartesia's lean go-to-market structure, asserting that high marketing budgets often conceal product mediocrity and emphasizing self-serve developer discovery.47:19–50:04 · Guest teaching 5/10 Core Pillars of Cartesia's Growth: Model Quality and Research-Product Synergy Karan dismisses the conventional assumption that companies must choose between cutting-edge research and commercial product delivery, arguing both can be built together with disciplined alignment.50:05–55:00 · Guest teaching 5/10 Founder Alignment, Radical Candor, and High-Trust Culture Karan describes the founding team's high-trust, ego-free alignment and proposes a framework for sizing the voice AI market by multiplying interaction minutes by per-minute business value.55:00–57:21 · Guest teaching 3/10 Venture Financing Success and Long-Term Value Creation Karan attributes Cartesia's fundraising success with tier-one venture firms to razor-sharp focus on real-time interactive intelligence and attracting world-class research talent.1:46–5:54 · Guest disagreement 0/10 Gaming Origins, Hardware Assembly, and AI Curiosity Karan reflects on assembling GPUs and gaming online in Delhi, which spurred his technical research journey. Siddharth facilitates the personal backstory with nostalgic rapport.5:55–9:03 · Guest disagreement 0/10 Cartesia's Mathematical Homage and Stanford Co-Founding Team Karan explains how Cartesia was named after René Descartes and describes spinning out of Chris Ré's Stanford AI lab with fellow graduate students. Siddharth inquires about the team dynamics.9:03–11:34 · Guest disagreement 0/10 Pre-Seed Fundraising and Pivoting to Real-Time Interactive AI Karan recounts shifting Cartesia from an exploratory pre-seed lab to targeting continuous real-time interactive intelligence. Siddharth draws a comparison to the interactive OS in the film 'Her'.11:34–14:14 · Guest disagreement 1/10 Prioritizing Voice as the Primary Synchronous Interface Karan outlines why voice is the primary modality for synchronous communication and clarifies that the founders approached speech from deep sequence modeling rather than traditional speech research.14:16–17:14 · Guest disagreement 1/10 Model Lab Infrastructure, Experimentation Culture, and Audio Evaluation Karan explains the complex infrastructure required to run a model lab and notes how subjective, unstandardized audio evaluation benchmarks remain a massive hurdle.17:18–20:04 · Guest disagreement 1/10 Full-Stack System Co-Design and Pareto Dominance Karan defines full-stack co-design and rejects the standard engineering trade-off between speed, cost, and quality in favor of Pareto dominance.20:05–23:21 · Guest disagreement 0/10 Cartesia's Key Milestones: Proprietary Inference Engine and Voice Agents Karan describes building Cartesia's proprietary inference engine specifically tailored for real-time interactive models rather than standard LLMs.23:22–26:21 · Guest disagreement 1/10 Overcoming Long-Horizon Context and Memory in Voice AI Siddharth raises the issue of maintaining context in long conversations. Karan reveals that current voice agents fake long conversations using engineered short turns, explaining that true multimodal reasoning remains unsolved.26:21–30:23 · Guest disagreement 0/10 Hyper-Personalization and Proactive Customer Engagement Siddharth illustrates a relationship-driven banking scenario. Karan expands on how low-cost AI agents enable proactive customer service during dead air, such as telephone hold times.30:28–33:55 · Guest disagreement 0/10 Platform Architecture, Production Scale, and Voice Research Karan outlines the architectural demands of production scale and describes Cartesia's specialized research team investigating how voice tone influences business metrics.33:56–36:05 · Guest disagreement 0/10 The Psychology of Latency: Voice Versus Text Interaction Siddharth offers an astute observation comparing the strict latency demands of voice against user tolerance for text delay. Karan agrees and explains how instant text can paradoxically degrade perceived intelligence.36:07–40:40 · Guest disagreement 1/10 Enterprise Adoption Across Regulated Industries and Internal Workflows Siddharth challenges Karan on whether foundation model labs are eating the lunch of application developers. Karan reframes the tension as natural information asymmetry rather than predatory competition.40:40–44:13 · Guest disagreement 0/10 Evaluating Product-Market Fit and Scaling to Billion-Minute Volumes Karan shares indicators of product-market fit, highlighting enterprise customers scaling toward one billion minutes annually across multiple languages.44:15–47:18 · Guest disagreement 0/10 Lean Go-to-Market Strategy and Product-Led Expansion Karan outlines Cartesia's lean go-to-market structure, asserting that high marketing budgets often conceal product mediocrity and emphasizing self-serve developer discovery.47:19–50:04 · Guest disagreement 1/10 Core Pillars of Cartesia's Growth: Model Quality and Research-Product Synergy Karan dismisses the conventional assumption that companies must choose between cutting-edge research and commercial product delivery, arguing both can be built together with disciplined alignment.50:05–55:00 · Guest disagreement 1/10 Founder Alignment, Radical Candor, and High-Trust Culture Karan describes the founding team's high-trust, ego-free alignment and proposes a framework for sizing the voice AI market by multiplying interaction minutes by per-minute business value.55:00–57:21 · Guest disagreement 0/10 Venture Financing Success and Long-Term Value Creation Karan attributes Cartesia's fundraising success with tier-one venture firms to razor-sharp focus on real-time interactive intelligence and attracting world-class research talent.1:46–5:54 · Siddhartha pushing back 0/10 Gaming Origins, Hardware Assembly, and AI Curiosity Karan reflects on assembling GPUs and gaming online in Delhi, which spurred his technical research journey. Siddharth facilitates the personal backstory with nostalgic rapport.5:55–9:03 · Siddhartha pushing back 0/10 Cartesia's Mathematical Homage and Stanford Co-Founding Team Karan explains how Cartesia was named after René Descartes and describes spinning out of Chris Ré's Stanford AI lab with fellow graduate students. Siddharth inquires about the team dynamics.9:03–11:34 · Siddhartha pushing back 0/10 Pre-Seed Fundraising and Pivoting to Real-Time Interactive AI Karan recounts shifting Cartesia from an exploratory pre-seed lab to targeting continuous real-time interactive intelligence. Siddharth draws a comparison to the interactive OS in the film 'Her'.11:34–14:14 · Siddhartha pushing back 0/10 Prioritizing Voice as the Primary Synchronous Interface Karan outlines why voice is the primary modality for synchronous communication and clarifies that the founders approached speech from deep sequence modeling rather than traditional speech research.14:16–17:14 · Siddhartha pushing back 0/10 Model Lab Infrastructure, Experimentation Culture, and Audio Evaluation Karan explains the complex infrastructure required to run a model lab and notes how subjective, unstandardized audio evaluation benchmarks remain a massive hurdle.17:18–20:04 · Siddhartha pushing back 1/10 Full-Stack System Co-Design and Pareto Dominance Karan defines full-stack co-design and rejects the standard engineering trade-off between speed, cost, and quality in favor of Pareto dominance.20:05–23:21 · Siddhartha pushing back 0/10 Cartesia's Key Milestones: Proprietary Inference Engine and Voice Agents Karan describes building Cartesia's proprietary inference engine specifically tailored for real-time interactive models rather than standard LLMs.23:22–26:21 · Siddhartha pushing back 1/10 Overcoming Long-Horizon Context and Memory in Voice AI Siddharth raises the issue of maintaining context in long conversations. Karan reveals that current voice agents fake long conversations using engineered short turns, explaining that true multimodal reasoning remains unsolved.26:21–30:23 · Siddhartha pushing back 0/10 Hyper-Personalization and Proactive Customer Engagement Siddharth illustrates a relationship-driven banking scenario. Karan expands on how low-cost AI agents enable proactive customer service during dead air, such as telephone hold times.30:28–33:55 · Siddhartha pushing back 0/10 Platform Architecture, Production Scale, and Voice Research Karan outlines the architectural demands of production scale and describes Cartesia's specialized research team investigating how voice tone influences business metrics.33:56–36:05 · Siddhartha pushing back 0/10 The Psychology of Latency: Voice Versus Text Interaction Siddharth offers an astute observation comparing the strict latency demands of voice against user tolerance for text delay. Karan agrees and explains how instant text can paradoxically degrade perceived intelligence.36:07–40:40 · Siddhartha pushing back 2/10 Enterprise Adoption Across Regulated Industries and Internal Workflows Siddharth challenges Karan on whether foundation model labs are eating the lunch of application developers. Karan reframes the tension as natural information asymmetry rather than predatory competition.40:40–44:13 · Siddhartha pushing back 0/10 Evaluating Product-Market Fit and Scaling to Billion-Minute Volumes Karan shares indicators of product-market fit, highlighting enterprise customers scaling toward one billion minutes annually across multiple languages.44:15–47:18 · Siddhartha pushing back 0/10 Lean Go-to-Market Strategy and Product-Led Expansion Karan outlines Cartesia's lean go-to-market structure, asserting that high marketing budgets often conceal product mediocrity and emphasizing self-serve developer discovery.47:19–50:04 · Siddhartha pushing back 0/10 Core Pillars of Cartesia's Growth: Model Quality and Research-Product Synergy Karan dismisses the conventional assumption that companies must choose between cutting-edge research and commercial product delivery, arguing both can be built together with disciplined alignment.50:05–55:00 · Siddhartha pushing back 0/10 Founder Alignment, Radical Candor, and High-Trust Culture Karan describes the founding team's high-trust, ego-free alignment and proposes a framework for sizing the voice AI market by multiplying interaction minutes by per-minute business value.55:00–57:21 · Siddhartha pushing back 0/10 Venture Financing Success and Long-Term Value Creation Karan attributes Cartesia's fundraising success with tier-one venture firms to razor-sharp focus on real-time interactive intelligence and attracting world-class research talent.

speaking balance: gold is Siddhartha, purple is the guest (3 minute bins)

0:00 · Siddhartha 0% · guest 100%0:00 · Siddhartha 0% · guest 100%3:00 · Siddhartha 0% · guest 100%3:00 · Siddhartha 0% · guest 100%6:00 · Siddhartha 0% · guest 100%6:00 · Siddhartha 0% · guest 100%9:00 · Siddhartha 0% · guest 100%9:00 · Siddhartha 0% · guest 100%12:00 · Siddhartha 0% · guest 100%12:00 · Siddhartha 0% · guest 100%15:00 · Siddhartha 0% · guest 100%15:00 · Siddhartha 0% · guest 100%18:00 · Siddhartha 0% · guest 100%18:00 · Siddhartha 0% · guest 100%21:00 · Siddhartha 0% · guest 100%21:00 · Siddhartha 0% · guest 100%24:00 · Siddhartha 0% · guest 100%24:00 · Siddhartha 0% · guest 100%27:00 · Siddhartha 0% · guest 100%27:00 · Siddhartha 0% · guest 100%30:00 · Siddhartha 0% · guest 100%30:00 · Siddhartha 0% · guest 100%33:00 · Siddhartha 0% · guest 100%33:00 · Siddhartha 0% · guest 100%36:00 · Siddhartha 0% · guest 100%36:00 · Siddhartha 0% · guest 100%39:00 · Siddhartha 0% · guest 100%39:00 · Siddhartha 0% · guest 100%42:00 · Siddhartha 0% · guest 100%42:00 · Siddhartha 0% · guest 100%45:00 · Siddhartha 0% · guest 100%45:00 · Siddhartha 0% · guest 100%48:00 · Siddhartha 0% · guest 100%48:00 · Siddhartha 0% · guest 100%51:00 · Siddhartha 0% · guest 100%51:00 · Siddhartha 0% · guest 100%54:00 · Siddhartha 0% · guest 100%54:00 · Siddhartha 0% · guest 100%57:00 · Siddhartha 0% · guest 100%57:00 · Siddhartha 0% · guest 100%
Sharpest disagreement ▶ 49:12 Rejecting the research versus product dichotomy

Karan firmly challenges conventional industry consensus, arguing that treating fundamental AI research and commercial product execution as mutually exclusive is a false assumption.

Hardest push from Siddhartha ▶ 37:58 Pushing on model labs eating application developers' lunch

Siddharth directly challenges Karan on whether foundation model providers risk cannibalizing and alienating the application layer by building end-user software.

Biggest teaching moment ▶ 24:35 Demystifying 30-minute voice agent memory

Karan educates the host on the reality of current voice agents, demonstrating that 30-minute conversations are engineered hacks of sequential short turns rather than continuous intelligence.

Siddhartha holds their own ▶ 33:56 Psychology of latency across text vs voice modalities

Siddharth demonstrates deep domain understanding by contrasting human cognitive expectations for immediate conversational voice against tolerance for delayed text research.

the scores for every segment, with the reasoning behind each
ChapterTopicSiddhartha as informed peerGuest teachingGuest disagreementSiddhartha pushing backWhy
Gaming Origins, Hardware Assembly, and AI Curiosity 3200 Karan reflects on assembling GPUs and gaming online in Delhi, which spurred his technical research journey. Siddharth facilitates the personal backstory with nostalgic rapport.
Cartesia's Mathematical Homage and Stanford Co-Founding Team 3200 Karan explains how Cartesia was named after René Descartes and describes spinning out of Chris Ré's Stanford AI lab with fellow graduate students. Siddharth inquires about the team dynamics.
Pre-Seed Fundraising and Pivoting to Real-Time Interactive AI 4400 Karan recounts shifting Cartesia from an exploratory pre-seed lab to targeting continuous real-time interactive intelligence. Siddharth draws a comparison to the interactive OS in the film 'Her'.
Prioritizing Voice as the Primary Synchronous Interface 4510 Karan outlines why voice is the primary modality for synchronous communication and clarifies that the founders approached speech from deep sequence modeling rather than traditional speech research.
Model Lab Infrastructure, Experimentation Culture, and Audio Evaluation 3610 Karan explains the complex infrastructure required to run a model lab and notes how subjective, unstandardized audio evaluation benchmarks remain a massive hurdle.
Full-Stack System Co-Design and Pareto Dominance 3611 Karan defines full-stack co-design and rejects the standard engineering trade-off between speed, cost, and quality in favor of Pareto dominance.
Cartesia's Key Milestones: Proprietary Inference Engine and Voice Agents 3500 Karan describes building Cartesia's proprietary inference engine specifically tailored for real-time interactive models rather than standard LLMs.
Overcoming Long-Horizon Context and Memory in Voice AI 6611 Siddharth raises the issue of maintaining context in long conversations. Karan reveals that current voice agents fake long conversations using engineered short turns, explaining that true multimodal reasoning remains unsolved.
Hyper-Personalization and Proactive Customer Engagement 5500 Siddharth illustrates a relationship-driven banking scenario. Karan expands on how low-cost AI agents enable proactive customer service during dead air, such as telephone hold times.
Platform Architecture, Production Scale, and Voice Research 3500 Karan outlines the architectural demands of production scale and describes Cartesia's specialized research team investigating how voice tone influences business metrics.
The Psychology of Latency: Voice Versus Text Interaction 6500 Siddharth offers an astute observation comparing the strict latency demands of voice against user tolerance for text delay. Karan agrees and explains how instant text can paradoxically degrade perceived intelligence.
Enterprise Adoption Across Regulated Industries and Internal Workflows 6512 Siddharth challenges Karan on whether foundation model labs are eating the lunch of application developers. Karan reframes the tension as natural information asymmetry rather than predatory competition.
Evaluating Product-Market Fit and Scaling to Billion-Minute Volumes 3400 Karan shares indicators of product-market fit, highlighting enterprise customers scaling toward one billion minutes annually across multiple languages.
Lean Go-to-Market Strategy and Product-Led Expansion 3400 Karan outlines Cartesia's lean go-to-market structure, asserting that high marketing budgets often conceal product mediocrity and emphasizing self-serve developer discovery.
Core Pillars of Cartesia's Growth: Model Quality and Research-Product Synergy 4510 Karan dismisses the conventional assumption that companies must choose between cutting-edge research and commercial product delivery, arguing both can be built together with disciplined alignment.
Founder Alignment, Radical Candor, and High-Trust Culture 4510 Karan describes the founding team's high-trust, ego-free alignment and proposes a framework for sizing the voice AI market by multiplying interaction minutes by per-minute business value.
Venture Financing Success and Long-Term Value Creation 3300 Karan attributes Cartesia's fundraising success with tier-one venture firms to razor-sharp focus on real-time interactive intelligence and attracting world-class research talent.

Statements from this episode (24)

Disclosure
Goel: Incremental Academic Papers Felt Less Valuable Than Practical Commercialization
“Eventually I think what we realized is we made enough progress where it didn't feel as valuable to do the N plus one-eth academic paper. And it actually felt a lot more important to try to figure out how to use the technology work that we've done, and obviousl…”
Karan Goel Aug 7, 2026 ▶ 4:32
Assertion Supported
Goel: Cartesia spun out of Chris Ré's Stanford lab with four students
“We incubated in Chris's lab, right, and Chris is one of our co-founders as well as, like you know, obviously we're his four, like, grad students, so, so we spun out of the lab.”
Karan Goel Aug 7, 2026 ▶ 6:38
Prediction Not checkable as stated
Goel: Always-On Synchronous AI Will Become a Primary Future Interface
“We think that AI that's always on, always running, that's able to communicate synchronously with you, is going to be one of the future interfaces for how you consume intelligence.”
Karan Goel Aug 7, 2026 ▶ 10:20
Prediction Not checkable as stated
Goel: Single Multimodal Models Will Dominate Low-Latency Interactive AI
“Again, it was clear that the future is going to be single, single models that are multimodal can process multiple streams of data. Can basically do reasoning over that and operate at pretty low latencies.”
Karan Goel Aug 7, 2026 ▶ 12:15
Assertion Supported
Goel: Cartesia Founders Entered Voice AI Without Speech Research Expertise
“I had done one project in speech in my PhD, and that was for three months. So we didn't, we weren't experts in speech, but we knew a lot about sequence modeling. We'd spent a lot of time in, in the deep learning side.”
Karan Goel Aug 7, 2026 ▶ 12:49
Opinion
Goel: Best Interactive AI Models Require New Architectures
“And we believe that to build them, you need to build them on new architectures.”
Karan Goel Aug 7, 2026 ▶ 13:24
Insight
Goel: Building Multimodal Models Lacks Any Established Recipe or Published Papers
“In a lot of these new areas, like multimodal models, like, there's no, Known recipe, right? Like you can't just go, go to the internet and say like, hey, this is how we're going to build the model. Here's, you know, here's the recipe we can follow. Here's a pa…”
Karan Goel Aug 7, 2026 ▶ 15:56
Opinion
Goel: Model Evaluation Is Undervalued; Audio AI Benchmarks Remain Inadequate
“The most undervalued part of this is evaluating your models. Because, like, I think, especially in things like audio. Yeah. There, when we started the benchmarks were pretty non-existent. Even today, I would say the benchmarks are not great.”
Karan Goel Aug 7, 2026 ▶ 16:32
Insight
Goel: Building Pareto-dominant AI models requires full-stack co-design
“The only way you can build models like that is you have to go basically co-design everything. Otherwise you're going to sacrifice something for something else.”
Karan Goel Aug 7, 2026 ▶ 19:31
Disclosure
Goel: Cartesia Built Custom Inference Engine for Real-Time Models
“So as an example, we've built our own inference engine. It's pretty important because we think that these interactive real-time models are going to need to run in a pretty different way. So having the ability to have your own engine that is not designed for LL…”
Karan Goel Aug 7, 2026 ▶ 21:08
Disclosure
Goel: Cartesia Will Unify Listening, Thinking, and Responding Into One Model
“So ultimately our ambition is, again, how do we bring together listening, thinking, and response into a single model and just do that continuously. And that's the next phase.”
Karan Goel Aug 7, 2026 ▶ 23:06
Insight
Goel: Current Voice Agents Fake Long Conversations Through Engineered Short Turns
“Today when you build a voice agent, it is actually not running a 30 minute conversation, truly. It is actually running a series of 10:02 conversations, because it is basically splitting the 30 minutes into small turns, And then every turn is treated as one thi…”
Karan Goel Aug 7, 2026 ▶ 24:22
Insight
Goel: Multimodal AI Models Degrade in Intelligence Versus Text-Only Models
“They tend to degrade in terms of their intelligence when you try to build models that can do multimodal reasoning. Like text only models tend to be pretty smart, but when you do multimodal, they tend to get a little bit less smart.”
Karan Goel Aug 7, 2026 ▶ 25:35
Insight
Goel: AI agents make high-touch customer experiences economically feasible
“I actually think that's the kind of thing that's super exciting, which is that it opens up the opportunity to provide an experience that was otherwise infeasible.”
Karan Goel Aug 7, 2026 ▶ 28:01
Insight
Goel: Changing Voice Acoustic Profiles Drives Enormous Shifts in Business Metrics
“How do metrics shift because you changed how a voice sounds? It has enormous impact in use cases and different use cases need different experiences, right? If you're doing a nursing use case versus you're doing a banking use case versus you're actually doing a…”
Karan Goel Aug 7, 2026 ▶ 33:07
Insight
Goel: Multi-Second AI Chat Delays Are a Feature to Signal Thinking
“When you're chatting to AI, if it responds too quickly, it actually feels like it hasn't done any thinking at all. So, it is actually not a, it's a feature, not a bug, that it can take a few seconds to respond to you.”
Karan Goel Aug 7, 2026 ▶ 34:44
Insight
Goel: Minimizing latency without sacrificing intelligence is voice AI's biggest challenge
“A lot of what we see when people are building these systems is that they're struggling to understand how to bring latency down while keeping the overall intelligence of the system as high as possible. That is the biggest challenge a lot of these folks face.”
Karan Goel Aug 7, 2026 ▶ 35:20
Opinion
Goel: AI labs build apps out of capability excitement, not competition
“I don't think it comes from a place of trying to compete. I think a lot of times the stories around how these products got built is that people inside were amazed by what it was able to do. They built a prototype, and the prototype just happened to be somethin…”
Karan Goel Aug 7, 2026 ▶ 39:46
Insight
Goel: Information Asymmetry Gives AI Labs an Unavoidable Application Head Start
“I think the thing that, that is hard is the information asymmetry that exists, which is that model developers see where the future is before the application developers do. So that asymmetry is impossible to sort of, like, actually remove.”
Karan Goel Aug 7, 2026 ▶ 40:13
Assertion Not checkable as stated
Goel: Cartesia's Enterprise Clients Are Approaching One Billion Audio Minutes Annually
“We're seeing customers that are starting to go towards pretty insane volumes now on our, on top of us, like anywhere To like order one billion minutes, right? Getting to that scale a year for one customer.”
Karan Goel Aug 7, 2026 ▶ 42:03
Insight
Goel: AI research labs and product companies are not mutually exclusive
“This is something that a lot of people think is mutually exclusive, right? People think either you can be a research lab or you're a product company. Yeah. We, I think what we did was we said, no, no, no, that's not, we don't think those two things are not pos…”
Karan Goel Aug 7, 2026 ▶ 49:13
Prediction Not checkable as stated
Goel: Open Offices Will Disappear as Workers Constantly Talk to AI
“Imagine a workplace of the future where actually the way you work is you constantly have this collaborator and you actually redesign the workplace to Help people be able to work alongside the AI collaborator. So you can't have open office spaces anymore, becau…”
Karan Goel Aug 7, 2026 ▶ 52:25
Insight
Goel: Voice AI market scale hinges on interaction minutes and per-minute value
“I think there's two variables that matter for those companies. One is how many minutes of interaction are you driving? The other is what is the value of every minute? Because in different applications, every minute has different value.”
Karan Goel Aug 7, 2026 ▶ 53:24
Disclosure
Goel: Cartesia Focuses Strictly on Real-Time Voice Over Other Audio Use Cases
“For us, the mission was always about building interactive AI and closing the gap between AI and humans in terms of how smart and how interactive these models are. So we are always super focused on real time interaction. And that means also, for example, we nev…”
Karan Goel Aug 7, 2026 ▶ 55:30
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.