DocETL

product on 1 show · 8 statements across 1 episodes · said 13 times in 2 episodes since 2024

Latent Space 13

Mentions by year, every show

tap a year for its mentions
008115120242025episodesmentions
01120242025episodes it came up in
007.50.515120242025episodesmentions per episode

Latent Space 13

every mention on every show, scene by scene, with the transcript →

8 statements about DocETL, every show

Shreya Shankar: Practical DocETL users never have pre-annotated ground truth datasets
“So we're redoing our evaluation to be on data sets where we actually have ground truth from human annotators, but that's just not a practical setting. Like, nobody's coming to doc ETL with the ground truth.”
Shreya Shankar Nov 29, 2024 ▶ 49:36 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Shreya Shankar: DocETL builds semantic unstructured layers, not point-lookup RAG systems
“This is very different from traditional rag or Q&A or document processing for a chatbot. Like, the kinds of queries that people are, people want to use .etl for can be expressed as etl style sweep and harvest, kind of, I want to look at my entire dataset. I wo…”
Shreya Shankar Nov 29, 2024 ▶ 47:13 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Shreya Shankar: GPT-4o Mini reduces DocETL optimization cost by 90%
“The reason it was a hundred dollars, if I ran the optimizer with GPT-Foro mini as the LLMs, it would be 10 dollars. But we use GPT four. Oh, just because I think we did this at a time where many hadn't come out yet.”
Shreya Shankar Nov 29, 2024 ▶ 45:58 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
LATENT SPACE Assertion Supported
Eugene Yan: DocETL pipeline optimizer costs roughly $100 and 30 minutes
“Running the optimizer, right, they used the optimizer. I don't know how many plans the optimizer generated, but it cost approximately a hundred dollars and less than half an hour. Just go get lunch and you come back and you get your optimized pipeline. And the…”
Eugene Yan Nov 29, 2024 ▶ 37:11 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Eugene Yan: LLM pipeline validation still requires seed human-labeled data
“I'm of a slightly different take. I feel like we do need some set of seed human labeled data.”
Eugene Yan Nov 29, 2024 ▶ 33:20 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Shreya Shankar: Chunking efficacy is task-specific, requiring automated pipeline optimization
“Sometimes it's beneficial to chunk and sometimes you should not chunk. And we have observed this in a number of workloads and the insight that we've gained is that we will never know what it's all task specific and data specific. And we are so further convince…”
Shreya Shankar Nov 29, 2024 ▶ 19:31 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
LATENT SPACE Assertion Supported
Eugene Yan: DocETL applies database concepts to unstructured document processing
“Essentially, what this is trying to do is it's trying to take database, database concepts and pandas data frame concepts and try to apply them to shapeless documents.”
Eugene Yan Nov 29, 2024 ▶ 2:28 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
LATENT SPACE Assertion Not checkable as stated
Eugene Yan: LLMs are fairly inaccurate on complex document processing tasks
“But the problem is, is that for fairly complex tasks and data, LLM outputs for what we wanted to do is fairly inaccurate.”
Eugene Yan Nov 29, 2024 ▶ 0:46 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.