May 1, 2023 · 34m · mad

LLM Powered Search | Vectara Founder & CEO, Amr Awadallah

Amr Awadallah · 26m spoken Matt Turck · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At a Data Driven NYC live event, Vectara Founder and CEO Amr Awadallah sits down with Matt Turck to discuss how Retrieval-Augmented Generation (RAG) and LLM-powered search are replacing traditional keyword engines for enterprise data. He outlines Vectara's commercial strategy, platform architecture, and developer positioning, while offering insights into AI ethics, regulatory needs, and the future impact of artificial intelligence on jobs.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 9.4% of the talking time here. How this is scored →

Matt as informed peer 2.3 Guest teaching 5.2 Guest disagreement 1.1 Matt pushing back 0.3
05100:0010:0020:0030:000:13–2:55 · Matt as informed peer 1/10 Defining LLM Search as ChatGPT for Enterprise Data Matt asks a basic opening question regarding LLM search versus keyword search. Amr reframes enterprise search using a multilingual intern analogy to explain ChatGPT for internal data.2:55–6:11 · Matt as informed peer 2/10 Keyword Search versus Neural Answer Engines Matt asks how traditional indexing differs from neural search. Amr explains semantic language-to-meaning space mapping and the shift from keyword search engines to answer engines.6:11–10:35 · Matt as informed peer 2/10 Grounded Generation and RAG Architecture Explained Matt introduces grounded generation as a key term. Amr educates the audience on RAG using an intern highlighter analogy before detailing the three-part technical pipeline.10:35–14:41 · Matt as informed peer 3/10 RAG versus Fine-Tuning and Future Action Engines Matt asks probing questions about fine-tuning trade-offs and whether RAG is experimental. Amr breaks down why fine-tuning is cost-prohibitive and introduces future action engines.14:41–19:36 · Matt as informed peer 3/10 Vectara Business Strategy, Pricing, and Primary Use Cases Matt demonstrates industry knowledge by correctly naming Snowflake as the prime proprietary SaaS model. Amr explains why open-core fails against cloud providers and details Vectara's primary use cases.19:36–22:31 · Matt as informed peer 2/10 Platform Building Blocks and Developer Target Personas Matt prompts Amr on platform building blocks. Amr reframes software development targeting into Home Depot (descriptive) versus Ikea (prescriptive) developer personas.22:31–26:23 · Matt as informed peer 4/10 Physical Demonstration of Vectors and Vector Database Space Matt demonstrates domain context by placing vector DBs like Pinecone and orchestrators like LangChain in context. Amr dismisses raw vector DB tinkering for enterprise scale and conducts a physical vector demonstration using Matt's arm.26:23–28:40 · Matt as informed peer 3/10 Addressing AI Ethics, Regulations, and Research Pause Proposals Matt asks about proposed AI research pauses. Amr jokingly criticizes Elon Musk's motives while outlining a balanced stance favoring safety regulation over halting development.28:40–30:46 · Matt as informed peer 1/10 Audience Questions on Hybrid Search and Embedding Evaluation An audience member asks about sparse versus dense hybrid search evaluation. Amr explains Vectara's hybrid implementation and academic research collaborations.0:13–2:55 · Guest teaching 5/10 Defining LLM Search as ChatGPT for Enterprise Data Matt asks a basic opening question regarding LLM search versus keyword search. Amr reframes enterprise search using a multilingual intern analogy to explain ChatGPT for internal data.2:55–6:11 · Guest teaching 5/10 Keyword Search versus Neural Answer Engines Matt asks how traditional indexing differs from neural search. Amr explains semantic language-to-meaning space mapping and the shift from keyword search engines to answer engines.6:11–10:35 · Guest teaching 6/10 Grounded Generation and RAG Architecture Explained Matt introduces grounded generation as a key term. Amr educates the audience on RAG using an intern highlighter analogy before detailing the three-part technical pipeline.10:35–14:41 · Guest teaching 6/10 RAG versus Fine-Tuning and Future Action Engines Matt asks probing questions about fine-tuning trade-offs and whether RAG is experimental. Amr breaks down why fine-tuning is cost-prohibitive and introduces future action engines.14:41–19:36 · Guest teaching 5/10 Vectara Business Strategy, Pricing, and Primary Use Cases Matt demonstrates industry knowledge by correctly naming Snowflake as the prime proprietary SaaS model. Amr explains why open-core fails against cloud providers and details Vectara's primary use cases.19:36–22:31 · Guest teaching 6/10 Platform Building Blocks and Developer Target Personas Matt prompts Amr on platform building blocks. Amr reframes software development targeting into Home Depot (descriptive) versus Ikea (prescriptive) developer personas.22:31–26:23 · Guest teaching 5/10 Physical Demonstration of Vectors and Vector Database Space Matt demonstrates domain context by placing vector DBs like Pinecone and orchestrators like LangChain in context. Amr dismisses raw vector DB tinkering for enterprise scale and conducts a physical vector demonstration using Matt's arm.26:23–28:40 · Guest teaching 4/10 Addressing AI Ethics, Regulations, and Research Pause Proposals Matt asks about proposed AI research pauses. Amr jokingly criticizes Elon Musk's motives while outlining a balanced stance favoring safety regulation over halting development.28:40–30:46 · Guest teaching 5/10 Audience Questions on Hybrid Search and Embedding Evaluation An audience member asks about sparse versus dense hybrid search evaluation. Amr explains Vectara's hybrid implementation and academic research collaborations.0:13–2:55 · Guest disagreement 1/10 Defining LLM Search as ChatGPT for Enterprise Data Matt asks a basic opening question regarding LLM search versus keyword search. Amr reframes enterprise search using a multilingual intern analogy to explain ChatGPT for internal data.2:55–6:11 · Guest disagreement 1/10 Keyword Search versus Neural Answer Engines Matt asks how traditional indexing differs from neural search. Amr explains semantic language-to-meaning space mapping and the shift from keyword search engines to answer engines.6:11–10:35 · Guest disagreement 0/10 Grounded Generation and RAG Architecture Explained Matt introduces grounded generation as a key term. Amr educates the audience on RAG using an intern highlighter analogy before detailing the three-part technical pipeline.10:35–14:41 · Guest disagreement 1/10 RAG versus Fine-Tuning and Future Action Engines Matt asks probing questions about fine-tuning trade-offs and whether RAG is experimental. Amr breaks down why fine-tuning is cost-prohibitive and introduces future action engines.14:41–19:36 · Guest disagreement 1/10 Vectara Business Strategy, Pricing, and Primary Use Cases Matt demonstrates industry knowledge by correctly naming Snowflake as the prime proprietary SaaS model. Amr explains why open-core fails against cloud providers and details Vectara's primary use cases.19:36–22:31 · Guest disagreement 1/10 Platform Building Blocks and Developer Target Personas Matt prompts Amr on platform building blocks. Amr reframes software development targeting into Home Depot (descriptive) versus Ikea (prescriptive) developer personas.22:31–26:23 · Guest disagreement 2/10 Physical Demonstration of Vectors and Vector Database Space Matt demonstrates domain context by placing vector DBs like Pinecone and orchestrators like LangChain in context. Amr dismisses raw vector DB tinkering for enterprise scale and conducts a physical vector demonstration using Matt's arm.26:23–28:40 · Guest disagreement 2/10 Addressing AI Ethics, Regulations, and Research Pause Proposals Matt asks about proposed AI research pauses. Amr jokingly criticizes Elon Musk's motives while outlining a balanced stance favoring safety regulation over halting development.28:40–30:46 · Guest disagreement 1/10 Audience Questions on Hybrid Search and Embedding Evaluation An audience member asks about sparse versus dense hybrid search evaluation. Amr explains Vectara's hybrid implementation and academic research collaborations.0:13–2:55 · Matt pushing back 0/10 Defining LLM Search as ChatGPT for Enterprise Data Matt asks a basic opening question regarding LLM search versus keyword search. Amr reframes enterprise search using a multilingual intern analogy to explain ChatGPT for internal data.2:55–6:11 · Matt pushing back 0/10 Keyword Search versus Neural Answer Engines Matt asks how traditional indexing differs from neural search. Amr explains semantic language-to-meaning space mapping and the shift from keyword search engines to answer engines.6:11–10:35 · Matt pushing back 0/10 Grounded Generation and RAG Architecture Explained Matt introduces grounded generation as a key term. Amr educates the audience on RAG using an intern highlighter analogy before detailing the three-part technical pipeline.10:35–14:41 · Matt pushing back 1/10 RAG versus Fine-Tuning and Future Action Engines Matt asks probing questions about fine-tuning trade-offs and whether RAG is experimental. Amr breaks down why fine-tuning is cost-prohibitive and introduces future action engines.14:41–19:36 · Matt pushing back 1/10 Vectara Business Strategy, Pricing, and Primary Use Cases Matt demonstrates industry knowledge by correctly naming Snowflake as the prime proprietary SaaS model. Amr explains why open-core fails against cloud providers and details Vectara's primary use cases.19:36–22:31 · Matt pushing back 0/10 Platform Building Blocks and Developer Target Personas Matt prompts Amr on platform building blocks. Amr reframes software development targeting into Home Depot (descriptive) versus Ikea (prescriptive) developer personas.22:31–26:23 · Matt pushing back 1/10 Physical Demonstration of Vectors and Vector Database Space Matt demonstrates domain context by placing vector DBs like Pinecone and orchestrators like LangChain in context. Amr dismisses raw vector DB tinkering for enterprise scale and conducts a physical vector demonstration using Matt's arm.26:23–28:40 · Matt pushing back 0/10 Addressing AI Ethics, Regulations, and Research Pause Proposals Matt asks about proposed AI research pauses. Amr jokingly criticizes Elon Musk's motives while outlining a balanced stance favoring safety regulation over halting development.28:40–30:46 · Matt pushing back 0/10 Audience Questions on Hybrid Search and Embedding Evaluation An audience member asks about sparse versus dense hybrid search evaluation. Amr explains Vectara's hybrid implementation and academic research collaborations.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 9.5% · guest 90.5%0:00 · Matt 9.5% · guest 90.5%3:00 · Matt 12.1% · guest 87.9%3:00 · Matt 12.1% · guest 87.9%6:00 · Matt 6.8% · guest 93.2%6:00 · Matt 6.8% · guest 93.2%9:00 · Matt 11.6% · guest 88.4%9:00 · Matt 11.6% · guest 88.4%12:00 · Matt 8% · guest 92%12:00 · Matt 8% · guest 92%15:00 · Matt 0.2% · guest 99.8%15:00 · Matt 0.2% · guest 99.8%18:00 · Matt 10% · guest 90%18:00 · Matt 10% · guest 90%21:00 · Matt 15.6% · guest 84.4%21:00 · Matt 15.6% · guest 84.4%24:00 · Matt 16.3% · guest 83.7%24:00 · Matt 16.3% · guest 83.7%27:00 · Matt 4% · guest 96%27:00 · Matt 4% · guest 96%30:00 · Matt 5.7% · guest 94.3%30:00 · Matt 5.7% · guest 94.3%33:00 · Matt 18.9% · guest 81.1%33:00 · Matt 18.9% · guest 81.1%
Sharpest disagreement ▶ 23:15 Dismissing raw vector DB power users

Amr directly rejects developers who insist on low-level Hugging Face model tweaks, telling them that Vectara is not for them and that they should work elsewhere.

Hardest push from Matt ▶ 11:34 Challenging RAG readiness

Matt pushes back on the claim that grounded generation is already solved, questioning whether it is currently experimental versus proven in production.

Biggest teaching moment ▶ 6:23 Intern highlighter analogy for grounded generation

Amr educates the host and audience on RAG architecture by contrasting memorizing entire books with an intern highlighting relevant sentences.

Matt holds his own ▶ 15:24 Identifying Snowflake business model

Matt instantly names Snowflake when Amr prompts for the most successful closed-source software IPO, demonstrating sharp business model comprehension.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining LLM Search as ChatGPT for Enterprise Data 1510 Matt asks a basic opening question regarding LLM search versus keyword search. Amr reframes enterprise search using a multilingual intern analogy to explain ChatGPT for internal data.
Keyword Search versus Neural Answer Engines 2510 Matt asks how traditional indexing differs from neural search. Amr explains semantic language-to-meaning space mapping and the shift from keyword search engines to answer engines.
Grounded Generation and RAG Architecture Explained 2600 Matt introduces grounded generation as a key term. Amr educates the audience on RAG using an intern highlighter analogy before detailing the three-part technical pipeline.
RAG versus Fine-Tuning and Future Action Engines 3611 Matt asks probing questions about fine-tuning trade-offs and whether RAG is experimental. Amr breaks down why fine-tuning is cost-prohibitive and introduces future action engines.
Vectara Business Strategy, Pricing, and Primary Use Cases 3511 Matt demonstrates industry knowledge by correctly naming Snowflake as the prime proprietary SaaS model. Amr explains why open-core fails against cloud providers and details Vectara's primary use cases.
Platform Building Blocks and Developer Target Personas 2610 Matt prompts Amr on platform building blocks. Amr reframes software development targeting into Home Depot (descriptive) versus Ikea (prescriptive) developer personas.
Physical Demonstration of Vectors and Vector Database Space 4521 Matt demonstrates domain context by placing vector DBs like Pinecone and orchestrators like LangChain in context. Amr dismisses raw vector DB tinkering for enterprise scale and conducts a physical vector demonstration using Matt's arm.
Addressing AI Ethics, Regulations, and Research Pause Proposals 3420 Matt asks about proposed AI research pauses. Amr jokingly criticizes Elon Musk's motives while outlining a balanced stance favoring safety regulation over halting development.
Audience Questions on Hybrid Search and Embedding Evaluation 1510 An audience member asks about sparse versus dense hybrid search evaluation. Amr explains Vectara's hybrid implementation and academic research collaborations.

Statements from this episode (12)

Disclosure
Awadallah: Vectara provides ChatGPT for proprietary enterprise data
“So what we do is ChatGPT for your own data.”
Amr Awadallah May 1, 2023 ▶ 0:34
Assertion Contradicted
Amr Awadallah claims Vectara has solved the LLM hallucination problem
“I mean, you do have a human in the loop, so you can still qualify it, but we want it to be minimized to zero, and that's the problem that we solved.”
Amr Awadallah May 1, 2023 ▶ 2:47
Insight
Search engines are legacy and being replaced by answer engines
“So we're moving away from legacy, which I'm calling search engines. That's legacy to what we have today is answer engines that just give you the answer itself.”
Amr Awadallah May 1, 2023 ▶ 5:18
Disclosure
Vectara's sentence encoder processed Vietnamese despite it being excluded from training
“So we just trained a new model, a fresher version that we're working on right now on about 30 languages and Vietnamese was not one of them. And it's still figured out how to go from Vietnamese too.”
Amr Awadallah May 1, 2023 ▶ 9:11
Assertion Not checkable as stated
Retrieval-augmented generation is 100 times cheaper than fine-tuning, says Amr Awadallah
“The grounded generation approach is real time. A new fact comes in is showing up in the answers right away. There is no hallucination. I mean, you minimize it significantly and the cost is a hundred times cheaper.”
Amr Awadallah May 1, 2023 ▶ 11:14
Prediction Not checkable as stated
Secondary AI fact-checking models will reduce hallucinations to zero, predicts Awadallah
“You're going to have models that check the accuracy of other models, and that will help us bring down hallucination to zero.”
Amr Awadallah May 1, 2023 ▶ 13:07
Insight
Open-core business models fail against AWS due to Amazon's manageability advantages
“Open core does not work against Amazon because open core, you're building the core open and then manageability, security, reliability. That's your pro pro Amazon is really good at that shit. Like that's the, they know how to do security. So, so it becomes very…”
Amr Awadallah May 1, 2023 ▶ 15:06
Disclosure
Vectara is avoiding enterprise search sales during the economic downturn
“That's number 11 on the CIO list of things to spend on. I'm not going to go after that during an economic downturn.”
Amr Awadallah May 1, 2023 ▶ 18:43
Insight
Global developers outside Silicon Valley prefer turnkey APIs over assembling components
“The rest of the world, the entire world, Indonesia, the U S Ohio, like different states, they are prescriptive developers. They prefer the Ikea model. The Ikea model is here's the recipe for how to put the best together. Put these pieces, put these screws over…”
Amr Awadallah May 1, 2023 ▶ 20:50
Insight
DIY modular vector stacks face scale and latency issues in production
“And actually for prototyping, it's great to use that. Now get that into production at scale. And that's when you're going to start running into hiccups in terms of scalability cost wise, scalability performance wise on the ingest side, and then on the query ru…”
Amr Awadallah May 1, 2023 ▶ 23:13
Insight
Semantic vector models perform poorly for exact lookups like VIN numbers
“They work very well for long form queries that have a meaning in them, where you can really extract the meaning. They don't work very well for lookups a VIN number, an ISBN number, somebody's name. They don't actually work very well for that.”
Amr Awadallah May 1, 2023 ▶ 29:43
Prediction Not checkable as stated
Top developers, lawyers, and doctors will become 1,000x more productive using AI
“The 10 X developers, the 10 Xer developers, they are not going to be a hundred X. They're going to be a thousand X. The 10 X lawyers, they're going to be a thousand X. The 10 X VCs, they're going to be a thousand X. The 10 X doctors, they're going to be a thou…”
Amr Awadallah May 1, 2023 ▶ 32:52
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.