Insight certainty 4/5 debate potential 2/5

Shreya Shankar: DocETL builds semantic unstructured layers, not point-lookup RAG systems

Shreya Shankar · [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar · Nov 29, 2024 · at 47:13

Shreya Shankar clarifies the architectural boundary of DocETL, differentiating full-corpus unstructured document analytics from traditional Retrieval-Augmented Generation (RAG).

0:00 / 0:54exact quote · 54.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“This is very different from traditional rag or Q&A or document processing for a chatbot. Like, the kinds of queries that people are, people want to use .etl for can be expressed as etl style sweep and harvest, kind of, I want to look at my entire dataset. I would hire an analyst to look at my entire dataset and then just tell me things generate me reports that I can then learn from. So for example, this officer and misconduct thing is much more digestible than, say, looking at all these millions of PDFs. So in that sense, I mean, you can certainly apply chunking or like maybe a strategy used for processing. One of the documents could be used for rag systems, but the focus is not rag. It's like building a kind of semantic layer on top of your unstructured data is the way that I like to think about it.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Shreya Shankar

Insight
Shankar: AI evals differ from MLOps due to data scarcity
“The other thing is I think that AI engineering Evaluation or evals here is actually different from MLOps or ML evaluation for traditional ML models. We were in a much more, you know, data rich setting in MLOps. So we were taught to come up with loss metrics or…”
Shreya Shankar Mar 13, 2025 ▶ 3:20 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shankar: Decompose LLM pipelines into unit tasks with standalone intermediate assertions
“The idea is to have each node in your graph kind of be a standalone, do a standalone thing that you can have standalone assertions for. And if you think about, you know, infinitely many inputs flowing through your pipeline, there's going to be some fraction of…”
Shreya Shankar Sep 28, 2024 ▶ 57:35 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Insight
Shankar: LLM failure modes and evaluation techniques have largely stabilized
“I think techniques have stabilized. I think the kinds of failure modes of LLMs, I mean, they're still there, but it's not like changing every single day. We know that LLMs are bad at certain things. We know a little bit more about say limitations of the transf…”
Shreya Shankar Mar 13, 2025 ▶ 11:21 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shankar: Grounded synthetic data generation beats slow human annotation
“People don't know how to do anything other than Plan A, which is to go out and try to collect as much real world data as possible and take months because we're going to employ, like, human annotator teams to do this. Or Plan B, which is I'm going to ask an LLM…”
Shreya Shankar Mar 13, 2025 ▶ 13:48 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shankar: Evals are necessary to train AI reasoning models
“You need evals to train your reasoning models.”
Shreya Shankar Mar 13, 2025 ▶ 27:13 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Insight
Shreya Shankar: Chunking efficacy is task-specific, requiring automated pipeline optimization
“Sometimes it's beneficial to chunk and sometimes you should not chunk. And we have observed this in a number of workloads and the insight that we've gained is that we will never know what it's all task specific and data specific. And we are so further convince…”
Shreya Shankar Nov 29, 2024 ▶ 19:31 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.