Feb 22, 2024 · 31m · no-priors

No Priors Ep. 52 | With Pinecone CEO Edo Liberty

Edo Liberty · 22m spoken Sarah Guo · 2m spoken Elad Gil · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Pinecone founder and CEO Edo Liberty joins hosts Sarah Guo and Elad Gil on No Priors to discuss vector database architecture, enterprise retrieval-augmented generation (RAG), and why specialized infrastructure outperforms legacy databases and massive context windows. Liberty details the launch of Pinecone Serverless and shares his architectural vision of decoupling AI reasoning from external knowledge storage.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.4% of the talking time here. How this is scored →

The hosts as informed peer 5.6 Guest teaching 5.9 Guest disagreement 1.9 The hosts pushing back 1.2
05100:0010:0020:0030:000:30–3:06 · The hosts as informed peer 5/10 Origins of Pinecone and Vector Database Category Sarah and Elad ask foundational questions about vector databases and the timing of Pinecone's founding in 2019 before the generative AI explosion. Edo explains how models represent data as numeric embeddings rather than raw text/pixels.3:07–6:01 · The hosts as informed peer 6/10 Implementing RAG Architecture and Benchmarking Accuracy Sarah frames the trade-offs in context length and reliability that lead developers to RAG. Edo explains Pinecone's Common Crawl experiment showing RAG cuts hallucinations across major models by up to 50%.6:01–10:12 · The hosts as informed peer 5/10 Enterprise Production Use Cases with Notion and Gong Edo gently corrects Elad's question by clarifying that Canopy is an open-source framework whereas Serverless is Pinecone's core database architecture. He then explains how production customers like Notion and Gong scale to billions of vectors.10:13–13:54 · The hosts as informed peer 6/10 Simplifying RAG Pipelines with Canopy Framework Sarah probes developer pain points and asks why traditional solutions like Postgres/PGVector or Elastic cannot suffice. Edo details why retrofitted vector search on relational engines breaks down under scale and cost constraints.13:54–16:50 · The hosts as informed peer 6/10 Mechanics of Hybrid Search and Vector Representations Sarah presses Edo on whether hybrid search combining keywords and embeddings is merely a temporary bridge. Edo explains the mathematical equivalence of keywords to sparse vectors and predicts explicit keyword matching will fade.16:51–20:22 · The hosts as informed peer 7/10 Pinecone's Managed SaaS Architecture Versus Open Source Sarah questions Pinecone's proprietary closed-source model relative to open-source database norms. Elad reinforces this with an analogy from Databricks CEO Ali Ghodsi comparing open-source commercialization to hitting a grand slam with a baseball bat after a golf hole-in-one.20:24–22:36 · The hosts as informed peer 4/10 Long Context Windows Versus Vector Database Retrieval Elad asks about the rise of massive context windows and infinite context. Edo pushes back forcefully against marketing hype, noting that model providers profit off token billing and that context stuffing degrades performance while proving economically unfeasible.22:36–25:27 · The hosts as informed peer 6/10 Enterprise Data Privacy Personalization and GDPR Compliance Elad raises concerns about data leakage across multi-tenant enterprise models. Edo explains how decoupling model inference from data storage in a vector DB preserves GDPR compliance and instantaneous data deletion without complex retraining.25:29–27:34 · The hosts as informed peer 6/10 Choosing Between Prompt Engineering Fine-Tuning and RAG Elad breaks down the three architectural options developers weigh: prompt engineering, fine-tuning, and RAG. Edo contrasts the scientific promise of fine-tuning with the commercial reality that poor execution frequently degrades model accuracy.27:35–31:06 · The hosts as informed peer 5/10 Pinecone Infrastructure Roadmap and Classical IR Challenges Elad asks for Pinecone's roadmap and broader AI predictions. Edo passionately critiques the inefficiency of current monolithic architectures that cram the internet into GPU memory, calling for clear separation between reasoning and knowledge engines.0:30–3:06 · Guest teaching 5/10 Origins of Pinecone and Vector Database Category Sarah and Elad ask foundational questions about vector databases and the timing of Pinecone's founding in 2019 before the generative AI explosion. Edo explains how models represent data as numeric embeddings rather than raw text/pixels.3:07–6:01 · Guest teaching 6/10 Implementing RAG Architecture and Benchmarking Accuracy Sarah frames the trade-offs in context length and reliability that lead developers to RAG. Edo explains Pinecone's Common Crawl experiment showing RAG cuts hallucinations across major models by up to 50%.6:01–10:12 · Guest teaching 6/10 Enterprise Production Use Cases with Notion and Gong Edo gently corrects Elad's question by clarifying that Canopy is an open-source framework whereas Serverless is Pinecone's core database architecture. He then explains how production customers like Notion and Gong scale to billions of vectors.10:13–13:54 · Guest teaching 6/10 Simplifying RAG Pipelines with Canopy Framework Sarah probes developer pain points and asks why traditional solutions like Postgres/PGVector or Elastic cannot suffice. Edo details why retrofitted vector search on relational engines breaks down under scale and cost constraints.13:54–16:50 · Guest teaching 6/10 Mechanics of Hybrid Search and Vector Representations Sarah presses Edo on whether hybrid search combining keywords and embeddings is merely a temporary bridge. Edo explains the mathematical equivalence of keywords to sparse vectors and predicts explicit keyword matching will fade.16:51–20:22 · Guest teaching 5/10 Pinecone's Managed SaaS Architecture Versus Open Source Sarah questions Pinecone's proprietary closed-source model relative to open-source database norms. Elad reinforces this with an analogy from Databricks CEO Ali Ghodsi comparing open-source commercialization to hitting a grand slam with a baseball bat after a golf hole-in-one.20:24–22:36 · Guest teaching 7/10 Long Context Windows Versus Vector Database Retrieval Elad asks about the rise of massive context windows and infinite context. Edo pushes back forcefully against marketing hype, noting that model providers profit off token billing and that context stuffing degrades performance while proving economically unfeasible.22:36–25:27 · Guest teaching 6/10 Enterprise Data Privacy Personalization and GDPR Compliance Elad raises concerns about data leakage across multi-tenant enterprise models. Edo explains how decoupling model inference from data storage in a vector DB preserves GDPR compliance and instantaneous data deletion without complex retraining.25:29–27:34 · Guest teaching 6/10 Choosing Between Prompt Engineering Fine-Tuning and RAG Elad breaks down the three architectural options developers weigh: prompt engineering, fine-tuning, and RAG. Edo contrasts the scientific promise of fine-tuning with the commercial reality that poor execution frequently degrades model accuracy.27:35–31:06 · Guest teaching 6/10 Pinecone Infrastructure Roadmap and Classical IR Challenges Elad asks for Pinecone's roadmap and broader AI predictions. Edo passionately critiques the inefficiency of current monolithic architectures that cram the internet into GPU memory, calling for clear separation between reasoning and knowledge engines.0:30–3:06 · Guest disagreement 1/10 Origins of Pinecone and Vector Database Category Sarah and Elad ask foundational questions about vector databases and the timing of Pinecone's founding in 2019 before the generative AI explosion. Edo explains how models represent data as numeric embeddings rather than raw text/pixels.3:07–6:01 · Guest disagreement 1/10 Implementing RAG Architecture and Benchmarking Accuracy Sarah frames the trade-offs in context length and reliability that lead developers to RAG. Edo explains Pinecone's Common Crawl experiment showing RAG cuts hallucinations across major models by up to 50%.6:01–10:12 · Guest disagreement 2/10 Enterprise Production Use Cases with Notion and Gong Edo gently corrects Elad's question by clarifying that Canopy is an open-source framework whereas Serverless is Pinecone's core database architecture. He then explains how production customers like Notion and Gong scale to billions of vectors.10:13–13:54 · Guest disagreement 2/10 Simplifying RAG Pipelines with Canopy Framework Sarah probes developer pain points and asks why traditional solutions like Postgres/PGVector or Elastic cannot suffice. Edo details why retrofitted vector search on relational engines breaks down under scale and cost constraints.13:54–16:50 · Guest disagreement 2/10 Mechanics of Hybrid Search and Vector Representations Sarah presses Edo on whether hybrid search combining keywords and embeddings is merely a temporary bridge. Edo explains the mathematical equivalence of keywords to sparse vectors and predicts explicit keyword matching will fade.16:51–20:22 · Guest disagreement 2/10 Pinecone's Managed SaaS Architecture Versus Open Source Sarah questions Pinecone's proprietary closed-source model relative to open-source database norms. Elad reinforces this with an analogy from Databricks CEO Ali Ghodsi comparing open-source commercialization to hitting a grand slam with a baseball bat after a golf hole-in-one.20:24–22:36 · Guest disagreement 4/10 Long Context Windows Versus Vector Database Retrieval Elad asks about the rise of massive context windows and infinite context. Edo pushes back forcefully against marketing hype, noting that model providers profit off token billing and that context stuffing degrades performance while proving economically unfeasible.22:36–25:27 · Guest disagreement 1/10 Enterprise Data Privacy Personalization and GDPR Compliance Elad raises concerns about data leakage across multi-tenant enterprise models. Edo explains how decoupling model inference from data storage in a vector DB preserves GDPR compliance and instantaneous data deletion without complex retraining.25:29–27:34 · Guest disagreement 1/10 Choosing Between Prompt Engineering Fine-Tuning and RAG Elad breaks down the three architectural options developers weigh: prompt engineering, fine-tuning, and RAG. Edo contrasts the scientific promise of fine-tuning with the commercial reality that poor execution frequently degrades model accuracy.27:35–31:06 · Guest disagreement 3/10 Pinecone Infrastructure Roadmap and Classical IR Challenges Elad asks for Pinecone's roadmap and broader AI predictions. Edo passionately critiques the inefficiency of current monolithic architectures that cram the internet into GPU memory, calling for clear separation between reasoning and knowledge engines.0:30–3:06 · The hosts pushing back 1/10 Origins of Pinecone and Vector Database Category Sarah and Elad ask foundational questions about vector databases and the timing of Pinecone's founding in 2019 before the generative AI explosion. Edo explains how models represent data as numeric embeddings rather than raw text/pixels.3:07–6:01 · The hosts pushing back 1/10 Implementing RAG Architecture and Benchmarking Accuracy Sarah frames the trade-offs in context length and reliability that lead developers to RAG. Edo explains Pinecone's Common Crawl experiment showing RAG cuts hallucinations across major models by up to 50%.6:01–10:12 · The hosts pushing back 1/10 Enterprise Production Use Cases with Notion and Gong Edo gently corrects Elad's question by clarifying that Canopy is an open-source framework whereas Serverless is Pinecone's core database architecture. He then explains how production customers like Notion and Gong scale to billions of vectors.10:13–13:54 · The hosts pushing back 1/10 Simplifying RAG Pipelines with Canopy Framework Sarah probes developer pain points and asks why traditional solutions like Postgres/PGVector or Elastic cannot suffice. Edo details why retrofitted vector search on relational engines breaks down under scale and cost constraints.13:54–16:50 · The hosts pushing back 2/10 Mechanics of Hybrid Search and Vector Representations Sarah presses Edo on whether hybrid search combining keywords and embeddings is merely a temporary bridge. Edo explains the mathematical equivalence of keywords to sparse vectors and predicts explicit keyword matching will fade.16:51–20:22 · The hosts pushing back 1/10 Pinecone's Managed SaaS Architecture Versus Open Source Sarah questions Pinecone's proprietary closed-source model relative to open-source database norms. Elad reinforces this with an analogy from Databricks CEO Ali Ghodsi comparing open-source commercialization to hitting a grand slam with a baseball bat after a golf hole-in-one.20:24–22:36 · The hosts pushing back 2/10 Long Context Windows Versus Vector Database Retrieval Elad asks about the rise of massive context windows and infinite context. Edo pushes back forcefully against marketing hype, noting that model providers profit off token billing and that context stuffing degrades performance while proving economically unfeasible.22:36–25:27 · The hosts pushing back 1/10 Enterprise Data Privacy Personalization and GDPR Compliance Elad raises concerns about data leakage across multi-tenant enterprise models. Edo explains how decoupling model inference from data storage in a vector DB preserves GDPR compliance and instantaneous data deletion without complex retraining.25:29–27:34 · The hosts pushing back 1/10 Choosing Between Prompt Engineering Fine-Tuning and RAG Elad breaks down the three architectural options developers weigh: prompt engineering, fine-tuning, and RAG. Edo contrasts the scientific promise of fine-tuning with the commercial reality that poor execution frequently degrades model accuracy.27:35–31:06 · The hosts pushing back 1/10 Pinecone Infrastructure Roadmap and Classical IR Challenges Elad asks for Pinecone's roadmap and broader AI predictions. Edo passionately critiques the inefficiency of current monolithic architectures that cram the internet into GPU memory, calling for clear separation between reasoning and knowledge engines.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 35.5% · guest 64.5%0:00 · the hosts 35.5% · guest 64.5%3:00 · the hosts 15.8% · guest 84.2%3:00 · the hosts 15.8% · guest 84.2%6:00 · the hosts 18.1% · guest 81.9%6:00 · the hosts 18.1% · guest 81.9%9:00 · the hosts 14% · guest 86%9:00 · the hosts 14% · guest 86%12:00 · the hosts 21.8% · guest 78.2%12:00 · the hosts 21.8% · guest 78.2%15:00 · the hosts 15.1% · guest 84.9%15:00 · the hosts 15.1% · guest 84.9%18:00 · the hosts 32% · guest 68%18:00 · the hosts 32% · guest 68%21:00 · the hosts 24.2% · guest 75.8%21:00 · the hosts 24.2% · guest 75.8%24:00 · the hosts 19.4% · guest 80.6%24:00 · the hosts 19.4% · guest 80.6%27:00 · the hosts 6.8% · guest 93.2%27:00 · the hosts 6.8% · guest 93.2%30:00 · the hosts 24.6% · guest 75.4%30:00 · the hosts 24.6% · guest 75.4%
Sharpest disagreement ▶ 20:42 Edo calls out infinite context claims as token sales marketing

Edo rejects the premise of infinite context windows, pointing out that vendors sell by the token and that stuffing entire corpora into context is practically absurd.

Hardest push from the hosts ▶ 15:29 Sarah challenges Edo on whether hybrid search is merely temporary

Sarah directly questions Edo's prediction on hybrid search, asking him to clarify if combining keywords with embeddings is just a transient stopgap.

Biggest teaching moment ▶ 20:57 Edo dismantles the long-context window replacement narrative

Edo explains why context stuffing degrades accuracy and scales poorly, using the analogy that sending an entire query with the internet is as impractical as replacing Google with full-corpus context.

The host holds their own ▶ 19:20 Elad shares Databricks insight on open-source business models

Elad demonstrates industry depth by citing Databricks founder Ali Ghodsi's sports metaphor regarding the difficulty of building a viable enterprise business on top of open-source software.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins of Pinecone and Vector Database Category 5511 Sarah and Elad ask foundational questions about vector databases and the timing of Pinecone's founding in 2019 before the generative AI explosion. Edo explains how models represent data as numeric embeddings rather than raw text/pixels.
Implementing RAG Architecture and Benchmarking Accuracy 6611 Sarah frames the trade-offs in context length and reliability that lead developers to RAG. Edo explains Pinecone's Common Crawl experiment showing RAG cuts hallucinations across major models by up to 50%.
Enterprise Production Use Cases with Notion and Gong 5621 Edo gently corrects Elad's question by clarifying that Canopy is an open-source framework whereas Serverless is Pinecone's core database architecture. He then explains how production customers like Notion and Gong scale to billions of vectors.
Simplifying RAG Pipelines with Canopy Framework 6621 Sarah probes developer pain points and asks why traditional solutions like Postgres/PGVector or Elastic cannot suffice. Edo details why retrofitted vector search on relational engines breaks down under scale and cost constraints.
Mechanics of Hybrid Search and Vector Representations 6622 Sarah presses Edo on whether hybrid search combining keywords and embeddings is merely a temporary bridge. Edo explains the mathematical equivalence of keywords to sparse vectors and predicts explicit keyword matching will fade.
Pinecone's Managed SaaS Architecture Versus Open Source 7521 Sarah questions Pinecone's proprietary closed-source model relative to open-source database norms. Elad reinforces this with an analogy from Databricks CEO Ali Ghodsi comparing open-source commercialization to hitting a grand slam with a baseball bat after a golf hole-in-one.
Long Context Windows Versus Vector Database Retrieval 4742 Elad asks about the rise of massive context windows and infinite context. Edo pushes back forcefully against marketing hype, noting that model providers profit off token billing and that context stuffing degrades performance while proving economically unfeasible.
Enterprise Data Privacy Personalization and GDPR Compliance 6611 Elad raises concerns about data leakage across multi-tenant enterprise models. Edo explains how decoupling model inference from data storage in a vector DB preserves GDPR compliance and instantaneous data deletion without complex retraining.
Choosing Between Prompt Engineering Fine-Tuning and RAG 6611 Elad breaks down the three architectural options developers weigh: prompt engineering, fine-tuning, and RAG. Edo contrasts the scientific promise of fine-tuning with the commercial reality that poor execution frequently degrades model accuracy.
Pinecone Infrastructure Roadmap and Classical IR Challenges 5631 Elad asks for Pinecone's roadmap and broader AI predictions. Edo passionately critiques the inefficiency of current monolithic architectures that cram the internet into GPU memory, calling for clear separation between reasoning and knowledge engines.

Statements from this episode (17)

Assertion Not checkable as stated
Liberty: Mainstream Engineers Were Already Adopting BERT by 2019
“In 2019, the earthquake had already happened. Deep learning models and so on have already been grappled with. Large language models and transformer models like BERT and others started being used by the more mainstream engineering cohorts.”
Edo Liberty Feb 22, 2024 ▶ 2:28
Assertion Partly supported
Liberty: RAG Over Internet Data Reduces LLM Hallucinations by 50%
“And you could see that if you augment all of them with RAG on, even on the internet, which is data that they were trained on, you can reduce hallucinations significantly up to 50% sometimes.”
Edo Liberty Feb 22, 2024 ▶ 5:17
Insight
Liberty: RAG Levels the Playing Field Across Different Foundation Models
“Interestingly enough, many of them actually start behaving quite similarly in terms of level of accuracy, even though without RAG, they actually have quite different behaviors. So it's sort of both like a uniform improvement and a little bit of leveling the pl…”
Edo Liberty Feb 22, 2024 ▶ 5:27
Assertion Supported
Liberty: Notion Q&A runs AI question answering on Pinecone
“Notion Q&A now runs on, on Pinecone, and they serve essentially question answering with AI to tens of thousands and probably hundreds of thousands of their own customers.”
Edo Liberty Feb 22, 2024 ▶ 7:03
Assertion Supported
Liberty: Gong uses Pinecone for all customer sales call search
“Gong does the same thing with sales calls. Again, serves all of their use cases for all of their customers, and so on.”
Edo Liberty Feb 22, 2024 ▶ 7:16
Assertion Not checkable as stated
Liberty: Pinecone Serverless Tested with Tens of Billions of Vectors
“We've tested it with tens and tens of billions with live customers and live traffic.”
Edo Liberty Feb 22, 2024 ▶ 9:47
Opinion
Liberty: Retrofitted Vector Indexes Like pgvector Fail at Production Scale
“Those other products don't work. They don't work either because they don't scale in terms of the efficiency scale, cost, the trade-offs that they can offer, because they're not designed to do this. They're designed to do something else. They kind of thought ab…”
Edo Liberty Feb 22, 2024 ▶ 12:47
Insight
Liberty: Keyword search is a deeply flawed retrieval method for AI
“With other search technologies, this is again, this is the wrong search mode. If you're searching with keywords and just not finding The relevant information, because the embeddings, the contextual space in which these pieces of text, documents, or images live…”
Edo Liberty Feb 22, 2024 ▶ 13:18
Insight
Liberty: Proper embedding retrieval rarely requires keywords alongside embeddings
“Our research actually shows that when you do this well, we, you very rarely need keywords alongside embeddings, but getting embeddings to perform perfectly is, is actually, it could be quite intricate.”
Edo Liberty Feb 22, 2024 ▶ 14:19
Assertion Partly supported
Liberty: High 90s Percentage of OSS Code Comes From Vendor Employees
“And in fact, if you look at statistics, even companies that are open source, 99% of the contributions are actually from the company itself. Not 99, but high nineties.”
Edo Liberty Feb 22, 2024 ▶ 18:18
Opinion
Liberty: OSS Vector Database Competitors Are Already Struggling With Commercialization
“And in fact, we already see, even though new players in the vector database space that, that, that basically started to try to take us down, all took the open source angle. We already see them, even young as they might be, they are already struggling, struggli…”
Edo Liberty Feb 22, 2024 ▶ 19:57
Assertion Not checkable as stated
Liberty: Pinecone Serverless is the fourth near-complete rewrite of its database
“Serverless is the fourth complete, almost complete rewrite of the entire database at Pinecon.”
Edo Liberty Feb 22, 2024 ▶ 20:17
Assertion Not checkable as stated
Liberty: Context Window Stuffing Increases LLM Costs Without Improving Results
“There's plenty of evidence that increasing the context size doesn't actually improve results unless, you know, you do this very carefully, right? So just what's called constant stuffing is not helping. You just pay more and don't actually get much for it.”
Edo Liberty Feb 22, 2024 ▶ 21:13
Disclosure
Liberty: RAG Architectures Pair Small Models With Trillion-Parameter Vector Databases
“Already today, we have users who use not even very large models, you know, maybe a few billion parameters, and the vector database next to the model contains trillions of parameters. And they get, you know, much better performance that way.”
Edo Liberty Feb 22, 2024 ▶ 22:07
Insight
Liberty: Vector databases enable GDPR compliance without complex model unlearning
“And the added benefit to that is, by the way, that you can be GDPR compliant. You can actually delete data. So if, you know, so, you know, if you're a company like a legal company and somebody deletes a document, you can just delete it from the vector database…”
Edo Liberty Feb 22, 2024 ▶ 24:51
Insight
Liberty: Fine-tuning without expert teams often worsens model performance
“Unless you have the research team and the AI experts that know how to fine tune, you might actually make things significantly worse. Okay. So there is, there's nothing that says that more data is going to make your model do better. In fact, it oftentimes gets …”
Edo Liberty Feb 22, 2024 ▶ 26:22
Insight
Liberty: Foundation Models Fundamentally Flawed by Combining Reasoning With Knowledge
“Foundational models get it fundamentally wrong. When we learn how to Build the subsystems of AI correctly, and for each one of them to do their roles optimally. Either we're going to do, be able to do the, to achieve the same tasks much cheaper, faster, better…”
Edo Liberty Feb 22, 2024 ▶ 29:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.