May 31, 2023 · 30m · mad

Long Term Memory for AI with Pinecone Founder & CEO, Edo Liberty

Edo Liberty · 19m spoken Matt Turck · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Data Driven NYC on The MAD Podcast, host Matt Turck interviews Pinecone Founder and CEO Edo Liberty about vector databases serving as long-term memory for generative AI. Liberty explains vector embeddings, Retrieval-Augmented Generation (RAG), Pinecone's developer-first growth model, and the future of autonomous AI agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 21.8% of the talking time here. How this is scored →

Matt as informed peer 3.2 Guest teaching 5.6 Guest disagreement 1.2 Matt pushing back 0.2
05100:0010:0020:0030:000:39–3:44 · Matt as informed peer 2/10 Defining Vector Embeddings and High-Dimensional Space Matt opens by citing Pinecone's $138M funding round and asking for foundational definitions. Edo provides a comprehensive explanation of how neural networks process vector embeddings as brain-like activations.3:44–6:01 · Matt as informed peer 1/10 Multi-Modal Data Transformation and External AI Memory When Matt asks about technical mechanics for converting video into numbers, Edo gently dismisses the premise as less interesting. He reframes the topic around external memory retrieval using a medical school analogy.6:01–8:11 · Matt as informed peer 4/10 What Is a Vector Database? Matt shows MLOps domain knowledge by asking whether feature stores overlap with vector databases. Edo clarifies that feature stores serve real-time state changes while vector databases provide long-term semantic memory.8:11–10:17 · Matt as informed peer 1/10 Semantic Search vs. Traditional Keyword Search Matt asks Edo to explain semantic search. Edo provides historical context on keyword indexing dating back to ancient print, contrasting it with searching by conceptual meaning.10:17–13:45 · Matt as informed peer 5/10 Pinecone's Founding Story and Edo Liberty's Background Matt demonstrates deep background knowledge on Edo's career, interjecting details about his Yale PhD and the current scale of AWS SageMaker. Edo details his journey from academia through Yahoo and AWS to founding Pinecone in 2019.13:45–16:31 · Matt as informed peer 3/10 The Generative AI Tech Stack and Autonomous Agents Matt lists key components of the generative AI stack, prompting Edo to elaborate. Edo explains how autonomous agents act as recursive software layers planning multi-step tasks.16:31–20:00 · Matt as informed peer 4/10 Enterprise Pinecone Use Cases and Reducing Hallucinations Matt asks about enterprise use cases and correctly suggests vector databases act as long-term memory to prevent hallucinations. Edo confirms this hypothesis with internal measurement metrics.20:00–23:30 · Matt as informed peer 5/10 Implementing Retrieval-Augmented Generation (RAG) Matt outlines a realistic architecture scenario for querying GPT-4 with enterprise data and asks about go-to-market strategies. Edo outlines the end-to-end prompt embedding pipeline and Pinecone's product-led growth model.23:30–26:19 · Matt as informed peer 4/10 Architectural Differentiation and System Scale Matt asks about market competition and playfully encourages Edo to badmouth competitors. Edo politely declines, focusing instead on Pinecone's custom storage architecture and customer obsession.26:19–29:59 · Matt as informed peer 3/10 Audience Q&A: Hyperscalers, Recommendation Systems, and Prompt Engineers During audience Q&A, an attendee asks about the longevity of prompt engineering roles. Edo strongly rejects the premise that prompt engineering is a permanent profession, calling it a temporary workaround for crude early technology.0:39–3:44 · Guest teaching 6/10 Defining Vector Embeddings and High-Dimensional Space Matt opens by citing Pinecone's $138M funding round and asking for foundational definitions. Edo provides a comprehensive explanation of how neural networks process vector embeddings as brain-like activations.3:44–6:01 · Guest teaching 7/10 Multi-Modal Data Transformation and External AI Memory When Matt asks about technical mechanics for converting video into numbers, Edo gently dismisses the premise as less interesting. He reframes the topic around external memory retrieval using a medical school analogy.6:01–8:11 · Guest teaching 6/10 What Is a Vector Database? Matt shows MLOps domain knowledge by asking whether feature stores overlap with vector databases. Edo clarifies that feature stores serve real-time state changes while vector databases provide long-term semantic memory.8:11–10:17 · Guest teaching 6/10 Semantic Search vs. Traditional Keyword Search Matt asks Edo to explain semantic search. Edo provides historical context on keyword indexing dating back to ancient print, contrasting it with searching by conceptual meaning.10:17–13:45 · Guest teaching 4/10 Pinecone's Founding Story and Edo Liberty's Background Matt demonstrates deep background knowledge on Edo's career, interjecting details about his Yale PhD and the current scale of AWS SageMaker. Edo details his journey from academia through Yahoo and AWS to founding Pinecone in 2019.13:45–16:31 · Guest teaching 6/10 The Generative AI Tech Stack and Autonomous Agents Matt lists key components of the generative AI stack, prompting Edo to elaborate. Edo explains how autonomous agents act as recursive software layers planning multi-step tasks.16:31–20:00 · Guest teaching 5/10 Enterprise Pinecone Use Cases and Reducing Hallucinations Matt asks about enterprise use cases and correctly suggests vector databases act as long-term memory to prevent hallucinations. Edo confirms this hypothesis with internal measurement metrics.20:00–23:30 · Guest teaching 5/10 Implementing Retrieval-Augmented Generation (RAG) Matt outlines a realistic architecture scenario for querying GPT-4 with enterprise data and asks about go-to-market strategies. Edo outlines the end-to-end prompt embedding pipeline and Pinecone's product-led growth model.23:30–26:19 · Guest teaching 5/10 Architectural Differentiation and System Scale Matt asks about market competition and playfully encourages Edo to badmouth competitors. Edo politely declines, focusing instead on Pinecone's custom storage architecture and customer obsession.26:19–29:59 · Guest teaching 6/10 Audience Q&A: Hyperscalers, Recommendation Systems, and Prompt Engineers During audience Q&A, an attendee asks about the longevity of prompt engineering roles. Edo strongly rejects the premise that prompt engineering is a permanent profession, calling it a temporary workaround for crude early technology.0:39–3:44 · Guest disagreement 1/10 Defining Vector Embeddings and High-Dimensional Space Matt opens by citing Pinecone's $138M funding round and asking for foundational definitions. Edo provides a comprehensive explanation of how neural networks process vector embeddings as brain-like activations.3:44–6:01 · Guest disagreement 3/10 Multi-Modal Data Transformation and External AI Memory When Matt asks about technical mechanics for converting video into numbers, Edo gently dismisses the premise as less interesting. He reframes the topic around external memory retrieval using a medical school analogy.6:01–8:11 · Guest disagreement 1/10 What Is a Vector Database? Matt shows MLOps domain knowledge by asking whether feature stores overlap with vector databases. Edo clarifies that feature stores serve real-time state changes while vector databases provide long-term semantic memory.8:11–10:17 · Guest disagreement 0/10 Semantic Search vs. Traditional Keyword Search Matt asks Edo to explain semantic search. Edo provides historical context on keyword indexing dating back to ancient print, contrasting it with searching by conceptual meaning.10:17–13:45 · Guest disagreement 0/10 Pinecone's Founding Story and Edo Liberty's Background Matt demonstrates deep background knowledge on Edo's career, interjecting details about his Yale PhD and the current scale of AWS SageMaker. Edo details his journey from academia through Yahoo and AWS to founding Pinecone in 2019.13:45–16:31 · Guest disagreement 1/10 The Generative AI Tech Stack and Autonomous Agents Matt lists key components of the generative AI stack, prompting Edo to elaborate. Edo explains how autonomous agents act as recursive software layers planning multi-step tasks.16:31–20:00 · Guest disagreement 0/10 Enterprise Pinecone Use Cases and Reducing Hallucinations Matt asks about enterprise use cases and correctly suggests vector databases act as long-term memory to prevent hallucinations. Edo confirms this hypothesis with internal measurement metrics.20:00–23:30 · Guest disagreement 0/10 Implementing Retrieval-Augmented Generation (RAG) Matt outlines a realistic architecture scenario for querying GPT-4 with enterprise data and asks about go-to-market strategies. Edo outlines the end-to-end prompt embedding pipeline and Pinecone's product-led growth model.23:30–26:19 · Guest disagreement 2/10 Architectural Differentiation and System Scale Matt asks about market competition and playfully encourages Edo to badmouth competitors. Edo politely declines, focusing instead on Pinecone's custom storage architecture and customer obsession.26:19–29:59 · Guest disagreement 4/10 Audience Q&A: Hyperscalers, Recommendation Systems, and Prompt Engineers During audience Q&A, an attendee asks about the longevity of prompt engineering roles. Edo strongly rejects the premise that prompt engineering is a permanent profession, calling it a temporary workaround for crude early technology.0:39–3:44 · Matt pushing back 0/10 Defining Vector Embeddings and High-Dimensional Space Matt opens by citing Pinecone's $138M funding round and asking for foundational definitions. Edo provides a comprehensive explanation of how neural networks process vector embeddings as brain-like activations.3:44–6:01 · Matt pushing back 0/10 Multi-Modal Data Transformation and External AI Memory When Matt asks about technical mechanics for converting video into numbers, Edo gently dismisses the premise as less interesting. He reframes the topic around external memory retrieval using a medical school analogy.6:01–8:11 · Matt pushing back 1/10 What Is a Vector Database? Matt shows MLOps domain knowledge by asking whether feature stores overlap with vector databases. Edo clarifies that feature stores serve real-time state changes while vector databases provide long-term semantic memory.8:11–10:17 · Matt pushing back 0/10 Semantic Search vs. Traditional Keyword Search Matt asks Edo to explain semantic search. Edo provides historical context on keyword indexing dating back to ancient print, contrasting it with searching by conceptual meaning.10:17–13:45 · Matt pushing back 0/10 Pinecone's Founding Story and Edo Liberty's Background Matt demonstrates deep background knowledge on Edo's career, interjecting details about his Yale PhD and the current scale of AWS SageMaker. Edo details his journey from academia through Yahoo and AWS to founding Pinecone in 2019.13:45–16:31 · Matt pushing back 0/10 The Generative AI Tech Stack and Autonomous Agents Matt lists key components of the generative AI stack, prompting Edo to elaborate. Edo explains how autonomous agents act as recursive software layers planning multi-step tasks.16:31–20:00 · Matt pushing back 0/10 Enterprise Pinecone Use Cases and Reducing Hallucinations Matt asks about enterprise use cases and correctly suggests vector databases act as long-term memory to prevent hallucinations. Edo confirms this hypothesis with internal measurement metrics.20:00–23:30 · Matt pushing back 0/10 Implementing Retrieval-Augmented Generation (RAG) Matt outlines a realistic architecture scenario for querying GPT-4 with enterprise data and asks about go-to-market strategies. Edo outlines the end-to-end prompt embedding pipeline and Pinecone's product-led growth model.23:30–26:19 · Matt pushing back 1/10 Architectural Differentiation and System Scale Matt asks about market competition and playfully encourages Edo to badmouth competitors. Edo politely declines, focusing instead on Pinecone's custom storage architecture and customer obsession.26:19–29:59 · Matt pushing back 0/10 Audience Q&A: Hyperscalers, Recommendation Systems, and Prompt Engineers During audience Q&A, an attendee asks about the longevity of prompt engineering roles. Edo strongly rejects the premise that prompt engineering is a permanent profession, calling it a temporary workaround for crude early technology.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 41.9% · guest 58.1%0:00 · Matt 41.9% · guest 58.1%3:00 · Matt 17.5% · guest 82.5%3:00 · Matt 17.5% · guest 82.5%6:00 · Matt 21.9% · guest 78.1%6:00 · Matt 21.9% · guest 78.1%9:00 · Matt 12.5% · guest 87.5%9:00 · Matt 12.5% · guest 87.5%12:00 · Matt 23.2% · guest 76.8%12:00 · Matt 23.2% · guest 76.8%15:00 · Matt 13.1% · guest 86.9%15:00 · Matt 13.1% · guest 86.9%18:00 · Matt 30% · guest 70%18:00 · Matt 30% · guest 70%21:00 · Matt 27.7% · guest 72.3%21:00 · Matt 27.7% · guest 72.3%24:00 · Matt 16% · guest 84%24:00 · Matt 16% · guest 84%27:00 · Matt 0.2% · guest 99.8%27:00 · Matt 0.2% · guest 99.8%30:00 · Matt 99% · guest 1%30:00 · Matt 99% · guest 1%
Sharpest disagreement ▶ 29:01 Rejecting prompt engineering as a sustainable profession

Edo directly rejects the popular premise that prompt engineering is a vital long-term role, arguing it is merely a temporary limitation of early crude AI models.

Hardest push from Matt ▶ 24:16 Host playfully encouraging competitor trash-talk

Matt nudges Edo to criticize vector database competitors by reminding him the podcast is recorded and granting permission to speak freely.

Biggest teaching moment ▶ 4:07 Reframing model ingestion into external memory access

Edo reframes Matt's question on data transformation mechanics, explaining that modern AI architecture gives models access to external memory during inference rather than forcing data into models.

Matt holds his own ▶ 7:00 Connecting MLOps feature stores to vector databases

Matt demonstrates his tech stack knowledge by bringing up feature stores from MLOps history to probe whether vector databases belong to the same architectural category.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining Vector Embeddings and High-Dimensional Space 2610 Matt opens by citing Pinecone's $138M funding round and asking for foundational definitions. Edo provides a comprehensive explanation of how neural networks process vector embeddings as brain-like activations.
Multi-Modal Data Transformation and External AI Memory 1730 When Matt asks about technical mechanics for converting video into numbers, Edo gently dismisses the premise as less interesting. He reframes the topic around external memory retrieval using a medical school analogy.
What Is a Vector Database? 4611 Matt shows MLOps domain knowledge by asking whether feature stores overlap with vector databases. Edo clarifies that feature stores serve real-time state changes while vector databases provide long-term semantic memory.
Semantic Search vs. Traditional Keyword Search 1600 Matt asks Edo to explain semantic search. Edo provides historical context on keyword indexing dating back to ancient print, contrasting it with searching by conceptual meaning.
Pinecone's Founding Story and Edo Liberty's Background 5400 Matt demonstrates deep background knowledge on Edo's career, interjecting details about his Yale PhD and the current scale of AWS SageMaker. Edo details his journey from academia through Yahoo and AWS to founding Pinecone in 2019.
The Generative AI Tech Stack and Autonomous Agents 3610 Matt lists key components of the generative AI stack, prompting Edo to elaborate. Edo explains how autonomous agents act as recursive software layers planning multi-step tasks.
Enterprise Pinecone Use Cases and Reducing Hallucinations 4500 Matt asks about enterprise use cases and correctly suggests vector databases act as long-term memory to prevent hallucinations. Edo confirms this hypothesis with internal measurement metrics.
Implementing Retrieval-Augmented Generation (RAG) 5500 Matt outlines a realistic architecture scenario for querying GPT-4 with enterprise data and asks about go-to-market strategies. Edo outlines the end-to-end prompt embedding pipeline and Pinecone's product-led growth model.
Architectural Differentiation and System Scale 4521 Matt asks about market competition and playfully encourages Edo to badmouth competitors. Edo politely declines, focusing instead on Pinecone's custom storage architecture and customer obsession.
Audience Q&A: Hyperscalers, Recommendation Systems, and Prompt Engineers 3640 During audience Q&A, an attendee asks about the longevity of prompt engineering roles. Edo strongly rejects the premise that prompt engineering is a permanent profession, calling it a temporary workaround for crude early technology.

Statements from this episode (10)

Insight
Liberty: AI architecture is shifting toward external memory during inference
“I think there is a significant shift in the way that people think about these building these smart AI systems. And They don't try to fit the data into the model as much as give the model access to the right data when it's doing the inference, right?”
Edo Liberty May 31, 2023 ▶ 4:22
Insight
Edo Liberty: Vector databases serve as long-term memory for AI
“A vector database is the infrastructure that, that supports that long-term memory, right?”
Edo Liberty May 31, 2023 ▶ 6:05
Insight
Liberty: Keyword search dominated due to scaling efficiency, not search quality
“Traditional search is, is, is keyword based. Not because keyword search is that great, but just because we have great mechanism to run it at scale, and it's very efficient.”
Edo Liberty May 31, 2023 ▶ 8:30
Insight
Edo Liberty: Semantic search mirrors human memory better than keyword lookup
“With semantic search, the whole idea is that You know what it means. You can, you should be able to search by the meaning in free text, and the match would not be because, oh, like three words matched, and they're next to each other, but rather the meaning is …”
Edo Liberty May 31, 2023 ▶ 9:50
Disclosure
Liberty: Pinecone did not foresee the massive ChatGPT-driven AI surge
“We didn't, by the way, foresee any of this ChatGPT thing happening. We knew it was, it would keep growing, but that, I think, completely took everybody by surprise, including us.”
Edo Liberty May 31, 2023 ▶ 13:28
Assertion Not checkable as stated
Liberty: A developer used Pinecone to build facial recognition for cows
“One of my favorite applications is somebody built a face detection for cows”
Edo Liberty May 31, 2023 ▶ 18:03
Disclosure
Liberty: Pinecone measures hallucination reduction as a core product metric
“In fact, I was looking at experiments today from one of our teams, and we literally measure reduction in hallucination as one of the core metrics that, that, that we try to drive.”
Edo Liberty May 31, 2023 ▶ 19:00
Insight
Liberty: Even naive RAG implementation significantly reduces AI hallucinations
“Like, you can play with it in a million different ways, but even if you do it relatively naively, that already gives you a huge bump in, in inaccuracy or reduction in hallucination, depending on how you want to measure it.”
Edo Liberty May 31, 2023 ▶ 21:21
Assertion Supported
Liberty: Cloud hyperscalers are actively building competing vector database capabilities
“I know for a fact they're looking at it. I know for a fact they will have something.”
Edo Liberty May 31, 2023 ▶ 27:08
Opinion
Liberty: Prompt engineering is a fad, not a long-term profession
“A lot of people believe that prompt engineering is incredibly crucial. I don't necessarily buy into that. I don't know if you remember, but people, like, thought that, like you know query building for Google would be a profession too, right? And that didn't qu…”
Edo Liberty May 31, 2023 ▶ 29:02
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.