“And so these days, like Lama Hub is a very rich repository of, like, you know, the hundred plus, like, different data loaders from all different services and formats, and it's growing, like, every day.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Jerry Liu
Disclosure
LlamaIndex is not building a vector database, integrates with 12-20 existing ones
“We're not building our own vector database, but we have a rich set of integrations with, you know like 12 to like 20 different vector databases out there these days.”
Jerry LiuJul 12, 2023▶ 9:20Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Disclosure
LlamaIndex focuses on data infrastructure while LangChain builds broader application frameworks
“Blind train is a great application framework for you to just like get us out of building blocks for a lot of different components, for instance, from like LL modules to prompts to some basic like retrieval and vector database abstractions to like also agent fr…”
Jerry LiuJul 12, 2023▶ 12:23Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Insight
Liu: LLM data pipelines differ from traditional ETL stacks due to unstructured content comprehension
“If we're building this new age of L-empowered applications, The kind of requirements for the type of data that like you want to load as well as how you want to extract information from that data will be a little bit different than the existing ETL stack. The r…”
Jerry LiuJul 12, 2023▶ 18:57Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
AssertionContradicted
Liu: Uber and Instabase use LlamaIndex for enterprise data applications
“We've seen people build these workflows at different settings from, for instance, like hacks on projects at startups building, for instance, like track GPT, like, plugin over, like, your Slack or your Notion all the way to kind of, like, bigger companies, for …”
Jerry LiuJul 12, 2023▶ 26:58Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Insight
Liu: Generating long-form content over custom data remains a hard problem
“Generating something that's like a paragraph is pretty easy for ChatGPT to do. Generating like an entire blog post or essay, especially over your data is a pretty challenging problem.”
Jerry LiuJul 12, 2023▶ 28:37Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Insight
Jerry Liu: Injecting metadata into text chunks improves LLM retrieval performance
“Second is being able to inject metadata actually is quite important to actually improve retrieval performance of like the downstream application, because like, you know, let's say you're splitting up like a sec, 10 K filing into a bunch of chunks within a sing…”
Jerry LiuJul 12, 2023▶ 8:23Building LlamaIndex: Jerry Liu on Scaling Retrieval-Augmented AI
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.