Disclosure certainty 4/5 debate potential 2/5

LlamaIndex is not building a vector database, integrates with 12-20 existing ones

Jerry Liu · Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI · Jul 12, 2023 · at 9:20

Jerry Liu, co-founder and CEO of LlamaIndex, discusses the framework's architecture and storage strategy relative to existing vector databases.

0:00 / 0:09exact quote · 9.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“We're not building our own vector database, but we have a rich set of integrations with, you know like 12 to like 20 different vector databases out there these days.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Jerry Liu

Disclosure
LlamaIndex focuses on data infrastructure while LangChain builds broader application frameworks
“Blind train is a great application framework for you to just like get us out of building blocks for a lot of different components, for instance, from like LL modules to prompts to some basic like retrieval and vector database abstractions to like also agent fr…”
Jerry Liu Jul 12, 2023 ▶ 12:23 Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Insight
Liu: LLM data pipelines differ from traditional ETL stacks due to unstructured content comprehension
“If we're building this new age of L-empowered applications, The kind of requirements for the type of data that like you want to load as well as how you want to extract information from that data will be a little bit different than the existing ETL stack. The r…”
Jerry Liu Jul 12, 2023 ▶ 18:57 Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Assertion Contradicted
Liu: Uber and Instabase use LlamaIndex for enterprise data applications
“We've seen people build these workflows at different settings from, for instance, like hacks on projects at startups building, for instance, like track GPT, like, plugin over, like, your Slack or your Notion all the way to kind of, like, bigger companies, for …”
Jerry Liu Jul 12, 2023 ▶ 26:58 Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Insight
Liu: Generating long-form content over custom data remains a hard problem
“Generating something that's like a paragraph is pretty easy for ChatGPT to do. Generating like an entire blog post or essay, especially over your data is a pretty challenging problem.”
Jerry Liu Jul 12, 2023 ▶ 28:37 Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Insight
Jerry Liu: Injecting metadata into text chunks improves LLM retrieval performance
“Second is being able to inject metadata actually is quite important to actually improve retrieval performance of like the downstream application, because like, you know, let's say you're splitting up like a sec, 10 K filing into a bunch of chunks within a sing…”
Jerry Liu Jul 12, 2023 ▶ 8:23 Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Assertion Supported
Liu: LlamaHub features over 100 community-contributed data loaders
“And so these days, like Lama Hub is a very rich repository of, like, you know, the hundred plus, like, different data loaders from all different services and formats, and it's growing, like, every day.”
Jerry Liu Jul 12, 2023 ▶ 15:39 Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.