Dec 30, 2025 · 28m · latent-space

[State of AI Startups] Memory/Learning, RL Envs & DBT-Fivetran — Sarah Catanzaro, Amplify

Sarah Catanzaro · 19m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Amplify Partners General Partner Sarah Catanzaro explores the intersection of data infrastructure and artificial intelligence, offering critical insights on venture funding realities, emerging architectural bottlenecks, and defensible startup strategies.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.2 Guest teaching 3.8 Guest disagreement 3.3 The hosts pushing back 2.7
05100:0010:0020:000:11–3:55 · The hosts as informed peer 5/10 Welcome and Exploring the Intersect of Data and AI Sarah reframes the host's question about the dbt-Fivetran merger signaling the end of the modern data stack, clarifying that the combined entity is targeting an elevated $600M+ IPO threshold rather than struggling. The host engages knowledgeably on revenue bars and customer profiles.3:56–8:16 · The hosts as informed peer 5/10 Comparing Analytical Workloads with Frontier AI Data Prep The exchange is highly reflective, with Sarah openly admitting she was wrong about data catalogs becoming a standalone category. The host asks probing questions about why discoverability failed to stick compared to embedded catalog features in Snowflake and dbt.8:17–17:02 · The hosts as informed peer 6/10 GPU Data Loading Bottlenecks and Infrastructure Scalability Both host and guest vent frustration over founders raising $100M seed rounds without near-term milestones. Sarah forcefully explains why private valuations are entirely made-up numbers, while the host cites specific market examples like Antithesis.17:04–23:25 · The hosts as informed peer 6/10 Deconstructing the Hype Surrounding World Models Sarah expresses skepticism towards world models as undefined while highlighting memory and continual learning. The host matches her depth by analyzing Cursor rules limitations, user churn, and the systems-level challenge of stateful inference weights.23:25–25:45 · The hosts as informed peer 5/10 Assessing Reinforcement Learning Environments: Substance vs. Fad Sarah takes a strong contrarian stance calling RL environments a fad. When the host pushes back with eight-figure lab contracts, Sarah dismisses lab spending habits as equivalent to their past overpayment for low-quality data labeling.25:46–28:03 · The hosts as informed peer 4/10 The Archetype of Research-Backed AI Application Startups Sarah wraps up by outlining her thesis on AI application startups that hire in-house research talent to solve specific bottlenecks like rule-following (Sierra) and RAG (Harvey). The dynamic is collaborative and analytical.0:11–3:55 · Guest teaching 5/10 Welcome and Exploring the Intersect of Data and AI Sarah reframes the host's question about the dbt-Fivetran merger signaling the end of the modern data stack, clarifying that the combined entity is targeting an elevated $600M+ IPO threshold rather than struggling. The host engages knowledgeably on revenue bars and customer profiles.3:56–8:16 · Guest teaching 4/10 Comparing Analytical Workloads with Frontier AI Data Prep The exchange is highly reflective, with Sarah openly admitting she was wrong about data catalogs becoming a standalone category. The host asks probing questions about why discoverability failed to stick compared to embedded catalog features in Snowflake and dbt.8:17–17:02 · Guest teaching 3/10 GPU Data Loading Bottlenecks and Infrastructure Scalability Both host and guest vent frustration over founders raising $100M seed rounds without near-term milestones. Sarah forcefully explains why private valuations are entirely made-up numbers, while the host cites specific market examples like Antithesis.17:04–23:25 · Guest teaching 3/10 Deconstructing the Hype Surrounding World Models Sarah expresses skepticism towards world models as undefined while highlighting memory and continual learning. The host matches her depth by analyzing Cursor rules limitations, user churn, and the systems-level challenge of stateful inference weights.23:25–25:45 · Guest teaching 6/10 Assessing Reinforcement Learning Environments: Substance vs. Fad Sarah takes a strong contrarian stance calling RL environments a fad. When the host pushes back with eight-figure lab contracts, Sarah dismisses lab spending habits as equivalent to their past overpayment for low-quality data labeling.25:46–28:03 · Guest teaching 2/10 The Archetype of Research-Backed AI Application Startups Sarah wraps up by outlining her thesis on AI application startups that hire in-house research talent to solve specific bottlenecks like rule-following (Sierra) and RAG (Harvey). The dynamic is collaborative and analytical.0:11–3:55 · Guest disagreement 4/10 Welcome and Exploring the Intersect of Data and AI Sarah reframes the host's question about the dbt-Fivetran merger signaling the end of the modern data stack, clarifying that the combined entity is targeting an elevated $600M+ IPO threshold rather than struggling. The host engages knowledgeably on revenue bars and customer profiles.3:56–8:16 · Guest disagreement 1/10 Comparing Analytical Workloads with Frontier AI Data Prep The exchange is highly reflective, with Sarah openly admitting she was wrong about data catalogs becoming a standalone category. The host asks probing questions about why discoverability failed to stick compared to embedded catalog features in Snowflake and dbt.8:17–17:02 · Guest disagreement 4/10 GPU Data Loading Bottlenecks and Infrastructure Scalability Both host and guest vent frustration over founders raising $100M seed rounds without near-term milestones. Sarah forcefully explains why private valuations are entirely made-up numbers, while the host cites specific market examples like Antithesis.17:04–23:25 · Guest disagreement 3/10 Deconstructing the Hype Surrounding World Models Sarah expresses skepticism towards world models as undefined while highlighting memory and continual learning. The host matches her depth by analyzing Cursor rules limitations, user churn, and the systems-level challenge of stateful inference weights.23:25–25:45 · Guest disagreement 7/10 Assessing Reinforcement Learning Environments: Substance vs. Fad Sarah takes a strong contrarian stance calling RL environments a fad. When the host pushes back with eight-figure lab contracts, Sarah dismisses lab spending habits as equivalent to their past overpayment for low-quality data labeling.25:46–28:03 · Guest disagreement 1/10 The Archetype of Research-Backed AI Application Startups Sarah wraps up by outlining her thesis on AI application startups that hire in-house research talent to solve specific bottlenecks like rule-following (Sierra) and RAG (Harvey). The dynamic is collaborative and analytical.0:11–3:55 · The hosts pushing back 2/10 Welcome and Exploring the Intersect of Data and AI Sarah reframes the host's question about the dbt-Fivetran merger signaling the end of the modern data stack, clarifying that the combined entity is targeting an elevated $600M+ IPO threshold rather than struggling. The host engages knowledgeably on revenue bars and customer profiles.3:56–8:16 · The hosts pushing back 2/10 Comparing Analytical Workloads with Frontier AI Data Prep The exchange is highly reflective, with Sarah openly admitting she was wrong about data catalogs becoming a standalone category. The host asks probing questions about why discoverability failed to stick compared to embedded catalog features in Snowflake and dbt.8:17–17:02 · The hosts pushing back 3/10 GPU Data Loading Bottlenecks and Infrastructure Scalability Both host and guest vent frustration over founders raising $100M seed rounds without near-term milestones. Sarah forcefully explains why private valuations are entirely made-up numbers, while the host cites specific market examples like Antithesis.17:04–23:25 · The hosts pushing back 3/10 Deconstructing the Hype Surrounding World Models Sarah expresses skepticism towards world models as undefined while highlighting memory and continual learning. The host matches her depth by analyzing Cursor rules limitations, user churn, and the systems-level challenge of stateful inference weights.23:25–25:45 · The hosts pushing back 5/10 Assessing Reinforcement Learning Environments: Substance vs. Fad Sarah takes a strong contrarian stance calling RL environments a fad. When the host pushes back with eight-figure lab contracts, Sarah dismisses lab spending habits as equivalent to their past overpayment for low-quality data labeling.25:46–28:03 · The hosts pushing back 1/10 The Archetype of Research-Backed AI Application Startups Sarah wraps up by outlining her thesis on AI application startups that hire in-house research talent to solve specific bottlenecks like rule-following (Sierra) and RAG (Harvey). The dynamic is collaborative and analytical.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 23:41 Declaring RL environments a fad

Sarah boldly dismisses one of the hottest AI infrastructure trends as a fad, brushing aside eight-figure contracts by comparing them to labs historically overpaying for poor data labeling.

Hardest push from the hosts ▶ 24:00 Challenging the dismissal of RL environments

The host refuses to accept that RL environments are fake or useless, pressing Sarah with the fact that top frontier labs are spending seven to eight figures on external environments instead of building in-house.

Biggest teaching moment ▶ 1:08 Correcting the death of data stack narrative

Sarah decisively corrects the host's premise that the dbt-Fivetran merger marks the end of the modern data stack, pointing out their healthy growth and the realistic $600M+ IPO revenue bar.

The host holds their own ▶ 22:58 Explaining stateful inference systems bottlenecks

The host demonstrates deep technical fluency by identifying how personalized continual learning breaks stateless inference paradigms, forcing weight updates, caching, and GPU loading overhead.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Welcome and Exploring the Intersect of Data and AI 5542 Sarah reframes the host's question about the dbt-Fivetran merger signaling the end of the modern data stack, clarifying that the combined entity is targeting an elevated $600M+ IPO threshold rather than struggling. The host engages knowledgeably on revenue bars and customer profiles.
Comparing Analytical Workloads with Frontier AI Data Prep 5412 The exchange is highly reflective, with Sarah openly admitting she was wrong about data catalogs becoming a standalone category. The host asks probing questions about why discoverability failed to stick compared to embedded catalog features in Snowflake and dbt.
GPU Data Loading Bottlenecks and Infrastructure Scalability 6343 Both host and guest vent frustration over founders raising $100M seed rounds without near-term milestones. Sarah forcefully explains why private valuations are entirely made-up numbers, while the host cites specific market examples like Antithesis.
Deconstructing the Hype Surrounding World Models 6333 Sarah expresses skepticism towards world models as undefined while highlighting memory and continual learning. The host matches her depth by analyzing Cursor rules limitations, user churn, and the systems-level challenge of stateful inference weights.
Assessing Reinforcement Learning Environments: Substance vs. Fad 5675 Sarah takes a strong contrarian stance calling RL environments a fad. When the host pushes back with eight-figure lab contracts, Sarah dismisses lab spending habits as equivalent to their past overpayment for low-quality data labeling.
The Archetype of Research-Backed AI Application Startups 4211 Sarah wraps up by outlining her thesis on AI application startups that hire in-house research talent to solve specific bottlenecks like rule-following (Sierra) and RAG (Harvey). The dynamic is collaborative and analytical.

Statements from this episode (19)

Assertion Not checkable as stated
Catanzaro: Both dbt and Fivetran Were Beating Revenue Targets Before Merger
“Both of the companies were actually like beating their revenue targets.”
Sarah Catanzaro Dec 30, 2025 ▶ 1:30
Assertion Supported
Catanzaro: Combined dbt and Fivetran Revenue Will Be Near $600M
“I believe that they'll actually be close to 600. I don't have the exact number.”
Sarah Catanzaro Dec 30, 2025 ▶ 1:54
Insight
Catanzaro: Companies Need Moderately Sized Data Teams, Not Armies of Engineers
“And I've become actually convinced that like, well, every company does need analytics engineers and does need data scientists. They probably don't need armies of them. And probably having like a moderately sized data and analytics team is a good thing.”
Sarah Catanzaro Dec 30, 2025 ▶ 3:40
Insight
Catanzaro: AI Data Prep Workloads Are Far Less Predictable Than BI Analytics
“I think one of the things that we saw with analytics that, you know, was surprising to some of the people in the data infrastructure space was that, like, the workloads were actually quite predictable. They were quite predictable because, like, many of them we…”
Sarah Catanzaro Dec 30, 2025 ▶ 4:16
Opinion
Catanzaro: Data Catalogs Failed to Become a Standalone Software Category
“They all have struggled a bit as a category. Many of them have been, you know, acquired subsequently which suggests that like, this was not, you know, perhaps a standalone category as a data scientist.”
Sarah Catanzaro Dec 30, 2025 ▶ 5:56
Insight
Catanzaro: Built-In Catalog Features in Snowflake and dbt Were Good Enough
“Many of these products offered kind of like data cataloging capabilities as a feature. And I think for humans, that was good enough. Like the data catalog that you had available in Snowflake was good enough. The data cataloging capabilities available in dbt, l…”
Sarah Catanzaro Dec 30, 2025 ▶ 6:50
Insight
Catanzaro: Data Catalog Startups Wrongly Prioritized Discoverability Over Governance
“I think a lot of data cataloging companies ended up focusing on, like, discoverability when perhaps, like, the real market opportunity was in governance.”
Sarah Catanzaro Dec 30, 2025 ▶ 8:08
Insight
Catanzaro: Data Infrastructure Has Scaled Surprisingly Well for AI Use Cases
“One of the things that has surprised me though is actually that, like, So much data infrastructure has actually scaled quite elegantly to meet the AI use case.”
Sarah Catanzaro Dec 30, 2025 ▶ 9:24
Prediction Not checkable as stated
Catanzaro: AI Agent Transactions Could Explode and Rival Ad-Tech Scale
“I think that could change, you know, like, as agents actually become kind of, like, more prevalent and are interfacing with each other, and therefore, like, perhaps, like, the number of transactions explodes.”
Sarah Catanzaro Dec 30, 2025 ▶ 9:45
Assertion Not checkable as stated
Catanzaro: $100M+ AI seed rounds without roadmaps happen frequently
“Like upwards of a hundred million dollars in a seed round where you have a long-term vision, but not a near-term roadmap. This is something that I'm seeing happening not just occasionally, but quite Frequently.”
Sarah Catanzaro Dec 30, 2025 ▶ 10:41
Opinion
Catanzaro: The AI Community Has Not Yet Defined World Models
“My, like, take on world models is that, like, we have not yet defined, like, what a world model is.”
Sarah Catanzaro Dec 30, 2025 ▶ 17:46
Insight
Catanzaro: World Models Designed for Specific Use Cases May Not Generalize
“I think one challenge that people have seen is that, like, world models perhaps designed for one specific use case might not generalize to others.”
Sarah Catanzaro Dec 30, 2025 ▶ 18:21
Assertion Not checkable as stated
Catanzaro: Fast-Growing AI Startups Suffer from Low Retention and High Churn
“I think what we're seeing is that like a lot of AI application companies, they're growing really quickly, but they suffer from, you know, relatively low retention, relatively high churn.”
Sarah Catanzaro Dec 30, 2025 ▶ 19:30
Assertion Not checkable as stated
Catanzaro: Continual Learning Requires Making Stateless Inference Stateful
“If you must update weights, then, like, you know, weights become stateful, and today, like, inference is not stateful.”
Sarah Catanzaro Dec 30, 2025 ▶ 23:00
Opinion
Catanzaro: RL Environments Are Just a Temporary Industry Fad
“So, I know I'm going on record on this, and like, I'm actually okay to be wrong, but I think RL Environments is just a fad.”
Sarah Catanzaro Dec 30, 2025 ▶ 23:42
Assertion Not checkable as stated
Catanzaro: Cursor Uses Real User Activity to Improve Agents and Tab
“And this is what cursor does. Like they actually do use you know, real user activity on their platform to, you know, significantly, like, improve both their coding agents as well as Tab.”
Sarah Catanzaro Dec 30, 2025 ▶ 24:59
Insight
Catanzaro: Application Startups Benefited Most from RAG Breakthroughs
“I think there were a lot of advances in RAG, and the biggest beneficiaries of these advances were the application companies for whom, you know, retrieval was a critical unlock.”
Sarah Catanzaro Dec 30, 2025 ▶ 26:16
Opinion
Catanzaro: Sierra's Success Stems from Solving Rule-Following in Customer Support
“Rule following is like a hard research problem, but if you solve rule following, then you unlock, you know, better customer support. And I think a lot of Sierra's success can be attributed to, like, their focus on this.”
Sarah Catanzaro Dec 30, 2025 ▶ 26:53
Opinion
Catanzaro: Runway Only Built Foundation Models Out of Necessity
“I don't think they would have built models if they didn't have to.”
Sarah Catanzaro Dec 30, 2025 ▶ 27:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.