Jun 27, 2024 · 34m · no-priors

No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu

Karan Goel · 15m spoken Albert Gu · 11m spoken Sarah Guo · 2m spoken Elad Gil · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Cartesia co-founders Karan Goel and Albert Gu discuss their pioneering work on State Space Models like S4 and Mamba, their mission to challenge Transformer dominance, and the launch of Sonic, their ultra-low-latency voice synthesis engine.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 13.9% of the talking time here. How this is scored →

The hosts as informed peer 4.9 Guest teaching 4.1 Guest disagreement 1.7 The hosts pushing back 1.9
05100:0010:0020:0030:001:32–5:05 · The hosts as informed peer 3/10 Academic Roots at Stanford and the Birth of S4 Sarah and Elad invite the founders to recount their academic history at Stanford under Chris Re. The co-founders share lighthearted stories of working together on S4 and filling up Google Cloud disk space.5:05–9:06 · The hosts as informed peer 4/10 State Space Models versus Transformer Architectures Albert explains state-space models as fuzzy compressors versus transformers. He educates listeners on how transformers struggle with raw waveforms and continuous signals without heavy tokenization.9:06–11:50 · The hosts as informed peer 5/10 Linear Scaling, Exact Retrieval, and Hybrid Architectures Sarah asks about efficiency versus quality across data types. Albert describes linear scaling advantages and why hybrid architectures with a roughly 10-to-1 SSM-to-attention ratio outperform either architecture alone.11:50–17:28 · The hosts as informed peer 7/10 Exploring Domain Applications: From Text to Genomics When Albert mentions applying Mamba to DNA modeling, Elad intervenes using his biology background to question what the specific problem formulation is. Albert concedes he is not a biologist and Karan steers the discussion toward audio and edge inference.17:28–20:33 · The hosts as informed peer 6/10 The Industry Shift Toward Efficient Local AI Hardware Elad and Sarah discuss Apple's 3B on-device models and compute economics. Karan counters that 3B models remain underpowered, arguing SSMs enable true high-performance intelligence on local commodity chips.20:33–25:00 · The hosts as informed peer 5/10 The Nuance of Speech: Why Text-to-Speech Is Unsolved Sarah pushes back on whether text-to-speech is already solved. Karan forcefully disagrees, outlining the lack of true conversational engagement, intonation nuance, and social role modeling in existing TTS systems.25:01–29:29 · The hosts as informed peer 6/10 Unified Multimodal Models versus Pipeline Orchestration Elad highlights the severe latency bottlenecks caused by orchestrating multi-model speech-to-text-to-speech pipelines, calling it inelegant. Karan agrees and details Cartesia's strategy to build a unified native multimodal model.29:30–32:46 · The hosts as informed peer 5/10 Research Aesthetics and 'Proofs from The Book' Albert explains his aesthetic research philosophy, prompting Elad to bring up Erdos's 'Proofs from The Book'. Sarah pushes for a live, un-cooked demo to test latency on a randomized quote.32:47–33:47 · The hosts as informed peer 3/10 Team Growth, Intern Culture, and Open Roles Sarah jokingly ribs Karan about his large intern ratio. The founders talk through hiring needs for their modeling team and their ongoing mission to overthrow the transformer empire.1:32–5:05 · Guest teaching 2/10 Academic Roots at Stanford and the Birth of S4 Sarah and Elad invite the founders to recount their academic history at Stanford under Chris Re. The co-founders share lighthearted stories of working together on S4 and filling up Google Cloud disk space.5:05–9:06 · Guest teaching 7/10 State Space Models versus Transformer Architectures Albert explains state-space models as fuzzy compressors versus transformers. He educates listeners on how transformers struggle with raw waveforms and continuous signals without heavy tokenization.9:06–11:50 · Guest teaching 6/10 Linear Scaling, Exact Retrieval, and Hybrid Architectures Sarah asks about efficiency versus quality across data types. Albert describes linear scaling advantages and why hybrid architectures with a roughly 10-to-1 SSM-to-attention ratio outperform either architecture alone.11:50–17:28 · Guest teaching 4/10 Exploring Domain Applications: From Text to Genomics When Albert mentions applying Mamba to DNA modeling, Elad intervenes using his biology background to question what the specific problem formulation is. Albert concedes he is not a biologist and Karan steers the discussion toward audio and edge inference.17:28–20:33 · Guest teaching 4/10 The Industry Shift Toward Efficient Local AI Hardware Elad and Sarah discuss Apple's 3B on-device models and compute economics. Karan counters that 3B models remain underpowered, arguing SSMs enable true high-performance intelligence on local commodity chips.20:33–25:00 · Guest teaching 6/10 The Nuance of Speech: Why Text-to-Speech Is Unsolved Sarah pushes back on whether text-to-speech is already solved. Karan forcefully disagrees, outlining the lack of true conversational engagement, intonation nuance, and social role modeling in existing TTS systems.25:01–29:29 · Guest teaching 4/10 Unified Multimodal Models versus Pipeline Orchestration Elad highlights the severe latency bottlenecks caused by orchestrating multi-model speech-to-text-to-speech pipelines, calling it inelegant. Karan agrees and details Cartesia's strategy to build a unified native multimodal model.29:30–32:46 · Guest teaching 3/10 Research Aesthetics and 'Proofs from The Book' Albert explains his aesthetic research philosophy, prompting Elad to bring up Erdos's 'Proofs from The Book'. Sarah pushes for a live, un-cooked demo to test latency on a randomized quote.32:47–33:47 · Guest teaching 1/10 Team Growth, Intern Culture, and Open Roles Sarah jokingly ribs Karan about his large intern ratio. The founders talk through hiring needs for their modeling team and their ongoing mission to overthrow the transformer empire.1:32–5:05 · Guest disagreement 1/10 Academic Roots at Stanford and the Birth of S4 Sarah and Elad invite the founders to recount their academic history at Stanford under Chris Re. The co-founders share lighthearted stories of working together on S4 and filling up Google Cloud disk space.5:05–9:06 · Guest disagreement 2/10 State Space Models versus Transformer Architectures Albert explains state-space models as fuzzy compressors versus transformers. He educates listeners on how transformers struggle with raw waveforms and continuous signals without heavy tokenization.9:06–11:50 · Guest disagreement 1/10 Linear Scaling, Exact Retrieval, and Hybrid Architectures Sarah asks about efficiency versus quality across data types. Albert describes linear scaling advantages and why hybrid architectures with a roughly 10-to-1 SSM-to-attention ratio outperform either architecture alone.11:50–17:28 · Guest disagreement 2/10 Exploring Domain Applications: From Text to Genomics When Albert mentions applying Mamba to DNA modeling, Elad intervenes using his biology background to question what the specific problem formulation is. Albert concedes he is not a biologist and Karan steers the discussion toward audio and edge inference.17:28–20:33 · Guest disagreement 2/10 The Industry Shift Toward Efficient Local AI Hardware Elad and Sarah discuss Apple's 3B on-device models and compute economics. Karan counters that 3B models remain underpowered, arguing SSMs enable true high-performance intelligence on local commodity chips.20:33–25:00 · Guest disagreement 4/10 The Nuance of Speech: Why Text-to-Speech Is Unsolved Sarah pushes back on whether text-to-speech is already solved. Karan forcefully disagrees, outlining the lack of true conversational engagement, intonation nuance, and social role modeling in existing TTS systems.25:01–29:29 · Guest disagreement 1/10 Unified Multimodal Models versus Pipeline Orchestration Elad highlights the severe latency bottlenecks caused by orchestrating multi-model speech-to-text-to-speech pipelines, calling it inelegant. Karan agrees and details Cartesia's strategy to build a unified native multimodal model.29:30–32:46 · Guest disagreement 1/10 Research Aesthetics and 'Proofs from The Book' Albert explains his aesthetic research philosophy, prompting Elad to bring up Erdos's 'Proofs from The Book'. Sarah pushes for a live, un-cooked demo to test latency on a randomized quote.32:47–33:47 · Guest disagreement 1/10 Team Growth, Intern Culture, and Open Roles Sarah jokingly ribs Karan about his large intern ratio. The founders talk through hiring needs for their modeling team and their ongoing mission to overthrow the transformer empire.1:32–5:05 · The hosts pushing back 0/10 Academic Roots at Stanford and the Birth of S4 Sarah and Elad invite the founders to recount their academic history at Stanford under Chris Re. The co-founders share lighthearted stories of working together on S4 and filling up Google Cloud disk space.5:05–9:06 · The hosts pushing back 1/10 State Space Models versus Transformer Architectures Albert explains state-space models as fuzzy compressors versus transformers. He educates listeners on how transformers struggle with raw waveforms and continuous signals without heavy tokenization.9:06–11:50 · The hosts pushing back 1/10 Linear Scaling, Exact Retrieval, and Hybrid Architectures Sarah asks about efficiency versus quality across data types. Albert describes linear scaling advantages and why hybrid architectures with a roughly 10-to-1 SSM-to-attention ratio outperform either architecture alone.11:50–17:28 · The hosts pushing back 5/10 Exploring Domain Applications: From Text to Genomics When Albert mentions applying Mamba to DNA modeling, Elad intervenes using his biology background to question what the specific problem formulation is. Albert concedes he is not a biologist and Karan steers the discussion toward audio and edge inference.17:28–20:33 · The hosts pushing back 1/10 The Industry Shift Toward Efficient Local AI Hardware Elad and Sarah discuss Apple's 3B on-device models and compute economics. Karan counters that 3B models remain underpowered, arguing SSMs enable true high-performance intelligence on local commodity chips.20:33–25:00 · The hosts pushing back 4/10 The Nuance of Speech: Why Text-to-Speech Is Unsolved Sarah pushes back on whether text-to-speech is already solved. Karan forcefully disagrees, outlining the lack of true conversational engagement, intonation nuance, and social role modeling in existing TTS systems.25:01–29:29 · The hosts pushing back 1/10 Unified Multimodal Models versus Pipeline Orchestration Elad highlights the severe latency bottlenecks caused by orchestrating multi-model speech-to-text-to-speech pipelines, calling it inelegant. Karan agrees and details Cartesia's strategy to build a unified native multimodal model.29:30–32:46 · The hosts pushing back 2/10 Research Aesthetics and 'Proofs from The Book' Albert explains his aesthetic research philosophy, prompting Elad to bring up Erdos's 'Proofs from The Book'. Sarah pushes for a live, un-cooked demo to test latency on a randomized quote.32:47–33:47 · The hosts pushing back 2/10 Team Growth, Intern Culture, and Open Roles Sarah jokingly ribs Karan about his large intern ratio. The founders talk through hiring needs for their modeling team and their ongoing mission to overthrow the transformer empire.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 18.4% · guest 81.6%0:00 · the hosts 18.4% · guest 81.6%3:00 · the hosts 6.1% · guest 93.9%3:00 · the hosts 6.1% · guest 93.9%6:00 · the hosts 4.7% · guest 95.3%6:00 · the hosts 4.7% · guest 95.3%9:00 · the hosts 12% · guest 88%9:00 · the hosts 12% · guest 88%12:00 · the hosts 18% · guest 82%12:00 · the hosts 18% · guest 82%15:00 · the hosts 4.8% · guest 95.2%15:00 · the hosts 4.8% · guest 95.2%18:00 · the hosts 27.4% · guest 72.6%18:00 · the hosts 27.4% · guest 72.6%21:00 · the hosts 9.6% · guest 90.4%21:00 · the hosts 9.6% · guest 90.4%24:00 · the hosts 13.2% · guest 86.8%24:00 · the hosts 13.2% · guest 86.8%27:00 · the hosts 15.3% · guest 84.7%27:00 · the hosts 15.3% · guest 84.7%30:00 · the hosts 17.2% · guest 82.8%30:00 · the hosts 17.2% · guest 82.8%33:00 · the hosts 31.9% · guest 68.1%33:00 · the hosts 31.9% · guest 68.1%
Sharpest disagreement ▶ 22:14 Karan refutes that speech generation is solved

Karan directly rejects Sarah's premise that audio generation was solved over the past year, asserting existing systems fail basic engagement tests.

Hardest push from the hosts ▶ 12:14 Elad interrogates the DNA modeling claim

Elad refuses to accept the vague premise of DNA foundation models without understanding whether it targets protein translation or folding.

Biggest teaching moment ▶ 7:55 Albert details why transformers fail on continuous waveforms

Albert explains why throwing transformers at raw continuous signals fails without artificial tokenization, demonstrating the theoretical necessity of SSMs.

The host holds their own ▶ 12:30 Elad demonstrates biology domain depth

Elad invokes his former career as a biologist to drill down on molecular mechanisms, leaving Albert to admit his own limitations in biology.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Academic Roots at Stanford and the Birth of S4 3210 Sarah and Elad invite the founders to recount their academic history at Stanford under Chris Re. The co-founders share lighthearted stories of working together on S4 and filling up Google Cloud disk space.
State Space Models versus Transformer Architectures 4721 Albert explains state-space models as fuzzy compressors versus transformers. He educates listeners on how transformers struggle with raw waveforms and continuous signals without heavy tokenization.
Linear Scaling, Exact Retrieval, and Hybrid Architectures 5611 Sarah asks about efficiency versus quality across data types. Albert describes linear scaling advantages and why hybrid architectures with a roughly 10-to-1 SSM-to-attention ratio outperform either architecture alone.
Exploring Domain Applications: From Text to Genomics 7425 When Albert mentions applying Mamba to DNA modeling, Elad intervenes using his biology background to question what the specific problem formulation is. Albert concedes he is not a biologist and Karan steers the discussion toward audio and edge inference.
The Industry Shift Toward Efficient Local AI Hardware 6421 Elad and Sarah discuss Apple's 3B on-device models and compute economics. Karan counters that 3B models remain underpowered, arguing SSMs enable true high-performance intelligence on local commodity chips.
The Nuance of Speech: Why Text-to-Speech Is Unsolved 5644 Sarah pushes back on whether text-to-speech is already solved. Karan forcefully disagrees, outlining the lack of true conversational engagement, intonation nuance, and social role modeling in existing TTS systems.
Unified Multimodal Models versus Pipeline Orchestration 6411 Elad highlights the severe latency bottlenecks caused by orchestrating multi-model speech-to-text-to-speech pipelines, calling it inelegant. Karan agrees and details Cartesia's strategy to build a unified native multimodal model.
Research Aesthetics and 'Proofs from The Book' 5312 Albert explains his aesthetic research philosophy, prompting Elad to bring up Erdos's 'Proofs from The Book'. Sarah pushes for a live, un-cooked demo to test latency on a randomized quote.
Team Growth, Intern Culture, and Open Roles 3112 Sarah jokingly ribs Karan about his large intern ratio. The founders talk through hiring needs for their modeling team and their ongoing mission to overthrow the transformer empire.

Statements from this episode (16)

Assertion Supported
Goel: Cartesia's Sonic TTS engine reduces typical voice latency by 150ms
“Even with what we've done with Sonic, we're already kind of shaving off, like, a 150 milliseconds off of you know, what they typically use.”
Karan Goel Jun 27, 2024 ▶ 1:16
Disclosure
Goel: Cartesia aims to cut another 600ms of voice latency this year
“And so, you know, the roadmap is let's, let's get to the next 600 milliseconds and try to shape those off in the, over the course of the year.”
Karan Goel Jun 27, 2024 ▶ 1:23
Assertion Supported
Gu: Mamba successfully applied state-space models to language modeling
“Recently proposed a model called Mamba which was kind of brought these to language modeling and showed really good results there.”
Albert Gu Jun 27, 2024 ▶ 2:24
Insight
Gu: State Space Models can be applied to almost all data types
“So it really can be applied to pretty much everything. So just like kind of Transformers, these are applied to everything. So can these sort of models over the course of research over a few years, we kind of realized that there are different advantages for dif…”
Albert Gu Jun 27, 2024 ▶ 6:51
Assertion Supported
Gu: Early SSMs excelled at raw signals but lagged Transformers on text
“The first types of models we were looking at were really good actually at modeling kind of these raw waveforms raw pixels, things like that, but not as good at modeling text, and transformers are way better there.”
Albert Gu Jun 27, 2024 ▶ 7:49
Opinion
Gu: Transformers struggle significantly on raw pixel or audio waveform data
“People think that like you can throw a transformer at like anything and it just works. Actually it doesn't really like if you try to throw it at like the raw pixel level or the raw sample level and in audio waveforms I think it doesn't work nearly as well.”
Albert Gu Jun 27, 2024 ▶ 8:16
Assertion Supported
Gu: Optimal hybrid models use a 10:1 ratio of SSM to attention
“People have found that the optimal ratio tends to be mostly SSM layers with a little bit of attention. So maybe a ratio of like 10 to one, I know of at least Probably like five groups that have independently verified that this is kind of the optimal ratio of t…”
Albert Gu Jun 27, 2024 ▶ 11:30
Assertion Supported
Gu: Researchers are applying Mamba-based foundation models to DNA sequences
“So I actually just heard from some collaborators today that they applied a mama based model on DNA modeling. They're basically bringing this idea of foundation models To DNA, which is kind of this new idea.”
Albert Gu Jun 27, 2024 ▶ 12:02
Prediction Not checkable as stated
Goel: AI models generating real-time data will replace traditional graphics engines
“In the future, like you will be essentially replacing, you know, graphics and rendering with essentially models that are outputting streams of data. In real time.”
Karan Goel Jun 27, 2024 ▶ 15:10
Disclosure
Goel: Cartesia aims to build edge-oriented infrastructure for SSMs
“I think in general for what we're trying to build is, is sort of the infrastructure to be able to train these models, make them run fast, and then bring them closer and closer to kind of be, you know, very edge oriented rather than cloud oriented.”
Karan Goel Jun 27, 2024 ▶ 17:17
Opinion
Goel: 3B parameter models are still too small and lack strong capabilities
“So the three V models are interesting, I think, but they aren't like, they're still small and not very capable.”
Karan Goel Jun 27, 2024 ▶ 18:23
Opinion
Goel: Current AI speech models fail to capture profession-specific vocal nuances
“And that's sort of the nuance that I don't think any models really capture well, which is like, you know, if you're a nurse, you need to talk in a different way than if you're a lawyer, or if you're a judge, or if you're a venture capitalist, you know, very di…”
Karan Goel Jun 27, 2024 ▶ 23:42
Insight
Gu: High-quality speech synthesis requires multimodal foundation models
“And so actually to really get, like, perfect even just TTS or, like, speech-to-speech you actually really need to have, like, a model that has, More understanding, like at least of the language, but kind of like, it's not really an isolated component anymore. …”
Albert Gu Jun 27, 2024 ▶ 24:27
Disclosure
Goel: Cartesia intends to train an on-device multimodal model
“Yes. But, you know, we have our own sort of set of techniques that we're developing in order to be able to do that effectively. I think that I will maybe leave for another podcast. But I think, yeah, I think that is the intention at the end of the day is build…”
Karan Goel Jun 27, 2024 ▶ 27:54
Insight
Gu: Aesthetic elegance was the primary driver behind inventing State Space Models
“People ask me, like, how do I treat my research problems, and my, I can't explain. My answer is just aesthetic. It's just like, there's something that I find elegant, and we're aesthetically pleasing about things, and to me, that's almost the most important th…”
Albert Gu Jun 27, 2024 ▶ 29:30
Assertion Supported
Goel: Cartesia runs Sonic TTS model locally on a standard Mac
“Yeah, I have a you know, our model running on our standard issue Mac here. Basically, this is you know, our text-to-speech model. Sonic on our playground is running in the cloud, and so you know, part of what I talked about earlier was how do you kind of bring…”
Karan Goel Jun 27, 2024 ▶ 31:29
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.