Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 1/5

Eugene Yan: DocETL pipeline optimizer costs roughly $100 and 30 minutes

Eugene Yan · [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar · Nov 29, 2024 · at 37:11

Eugene Yan reviews experimental results from the DocETL paper regarding the compute cost of automated LLM query plan optimization.

0:00 / 0:15exact quote · 15.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Running the optimizer, right, they used the optimizer. I don't know how many plans the optimizer generated, but it cost approximately a hundred dollars and less than half an hour. Just go get lunch and you come back and you get your optimized pipeline. And the optimization cost was just a hundred dollars.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Eugene Yan

Opinion
Eugene Yan: LLM ranking should use pairwise comparisons for stability
“I'm actually strongly convinced that it should be pairwise and we can debate that and see how it works. And I also think that... I think it's just more reliable and stable that way.”
Eugene Yan Nov 29, 2024 ▶ 12:42 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Opinion
Eugene Yan: Verifying LLM outputs is often harder than generating them
“And Shreya also has an interesting point, that it's much easier to verify the output and generate it. I actually observe the opposite. Or maybe it depends on the task. Like, for classification tasks, yes, it's easy. For, like, factuality, or comprehensiveness,…”
Eugene Yan Nov 29, 2024 ▶ 26:57 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Opinion
Eugene Yan: LLM pipeline validation still requires seed human-labeled data
“I'm of a slightly different take. I feel like we do need some set of seed human labeled data.”
Eugene Yan Nov 29, 2024 ▶ 33:20 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Insight
Eugene Yan: LLM-as-a-Judge works reliably when reduced to binary classification
“I think when we simplify it to binary classification metrics, I think it can work. And I think a lot of things can be simplified, like Shreya mentioned, I think a lot of things can be simplified to binary classification metrics. And I've seen evidence of it wo…”
Eugene Yan Nov 29, 2024 ▶ 41:04 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Opinion
Yan: Using LLMs as evaluators is the only way to scale
“I know that we have to use an LLM as an evaluator. There's no way around it. If we want to scale, I think that's the only way.”
Eugene Yan Sep 28, 2024 ▶ 44:00 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Insight
Yan: Pairwise evaluation fails for objective metrics like factuality
“The reason why pairwise preferences cannot work is that if you give two things that are both factual or if you give two things that are both non-factual, you would say that one is better than the other, but it still doesn't meet the bar of being factual enough…”
Eugene Yan Sep 28, 2024 ▶ 49:36 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.