Eugene Yan

9 statements across 2 episodes · 4 bullish · 2 bearish · 1 people on the record · first statement Sep 28, 2024 by Eugene Yan · said 7 times in 3 episodes since 2024 · across every show →

On the record as a speaker too: Eugene Yan's record, appearances and statements → this page counts the times other people say the name.

Mentions by year

brought up most by Shawn Wang (6)

tap a year for its mentions
003151202420252026episodesmentions
011202420252026episodes it came up in
002.50.551202420252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Eugene Yan, oldest first

Sep 28, 2024 neutral
Opinion
Yan: Optimizing LLM evaluation prompts requires 100 to 400 labeled examples
“I actually think the right number should be maybe a hundred to 400 if you want to be optimizing based on this.”
Eugene Yan Sep 28, 2024 ▶ 34:39 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Sep 28, 2024 negative
Insight
Yan: Pairwise evaluation fails for objective metrics like factuality
“The reason why pairwise preferences cannot work is that if you give two things that are both factual or if you give two things that are both non-factual, you would say that one is better than the other, but it still doesn't meet the bar of being factual enough…”
Eugene Yan Sep 28, 2024 ▶ 49:36 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Sep 28, 2024 positive
Insight
Yan: Annotators must update past grades when evaluation criteria drift
“If you find that your criteria has drifted, instead of trying to maintain the same criteria, And aligning to the previous grades. Instead, what we should do is we should revisit those previous grades and fix it because it's an iterative process.”
Eugene Yan Sep 28, 2024 ▶ 25:27 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Sep 28, 2024 positive
Opinion
Yan: Using LLMs as evaluators is the only way to scale
“I know that we have to use an LLM as an evaluator. There's no way around it. If we want to scale, I think that's the only way.”
Eugene Yan Sep 28, 2024 ▶ 44:00 [Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Nov 29, 2024 neutral
Opinion
Eugene Yan: LLM pipeline validation still requires seed human-labeled data
“I'm of a slightly different take. I feel like we do need some set of seed human labeled data.”
Eugene Yan Nov 29, 2024 ▶ 33:20 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Nov 29, 2024 negative
Opinion
Eugene Yan: Verifying LLM outputs is often harder than generating them
“And Shreya also has an interesting point, that it's much easier to verify the output and generate it. I actually observe the opposite. Or maybe it depends on the task. Like, for classification tasks, yes, it's easy. For, like, factuality, or comprehensiveness,…”
Eugene Yan Nov 29, 2024 ▶ 26:57 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Nov 29, 2024 positive
Assertion Supported
Eugene Yan: DocETL pipeline optimizer costs roughly $100 and 30 minutes
“Running the optimizer, right, they used the optimizer. I don't know how many plans the optimizer generated, but it cost approximately a hundred dollars and less than half an hour. Just go get lunch and you come back and you get your optimized pipeline. And the…”
Eugene Yan Nov 29, 2024 ▶ 37:11 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Nov 29, 2024 bullish
Prediction Not checkable as stated
Eugene Yan predicts LLM data pipelines will become reliable within two years
“It's gonna be a bit lossy, it's gonna be a bit stochastic, but I think we will figure it out in the next one to two years to get it to a more reliable state.”
Eugene Yan Nov 29, 2024 ▶ 17:24 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Nov 29, 2024
Disclosure
Eugene Yan: Removing document chunking boosted pipeline downstream metrics by 50%
“I was asked to help with a pipeline, and I was able to improve downstream metrics significantly by 20 to 50% by removing chunking.”
Eugene Yan Nov 29, 2024 ▶ 16:15 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.