Assertion Supported AI assessment confidence: 90% certainty 4/5 debate potential 3/5

Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned

Josh Albrecht · State of the Art: Training 70B LLMs on 10,000 H100 clusters · Jun 25, 2024 · at 1:01:43

Josh Albrecht, CTO of Imbue, explains their research findings after auditing common LLM benchmarks such as ANLI, RACE, and BoolQ.

0:00 / 0:32exact quote · 32.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NLI or race or pool cue or something like what you're really talking about is like performance on questions that make no sense. Like it's just like, did it guess the answer in this like really weird scenario? Like those are the ones that are left. Like when you look at the performance on the ones that actually make sense to everyone, all the models agree. We agree. Like everyone's on the same page”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Josh Albrecht

Insight
Albrecht: LLM emergence is an artifact of non-linear evaluation metrics
“This emergent behavior that you're seeing, Is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
Josh Albrecht Jun 25, 2024 ▶ 55:56 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Opinion
Albrecht: Training on AWS prevents diagnosing low-level hardware errors
“And if we're just using, you know, AWS or some other cloud provider, These errors are still going to be there, and you're gonna have no way to know and no way to debug this and no way to diagnose what's going wrong.”
Josh Albrecht Jun 25, 2024 ▶ 19:49 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Albrecht: AgentBench paper's appendix examples are actually incorrect solutions
“Like we were looking at the agent bench paper, I think just last week for our paper club. And one of the things that we noticed is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not…”
Josh Albrecht Jun 25, 2024 ▶ 1:06:05 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Albrecht: Optimizing for competitive coding benchmarks does not create useful programmers
“Like, we do a lot of code generation, but we don't really do a lot on, like, code competition problems for the very, very hard ones, so that you can go very far down that route and make something like really good at those problems, but not actually that useful…”
Josh Albrecht Jun 25, 2024 ▶ 1:13:00 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Albrecht: Imbue avoids Kubernetes to keep cluster infrastructure simple to debug
“Less layers of infrastructure, less layers of abstraction, make it a lot easier to work with. Like we don't use Kubernetes, for example, I would just directly launch these things and it's just been much easier to debug this way.”
Josh Albrecht Jun 25, 2024 ▶ 34:38 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Albrecht: Vision is not essential for most coding and reasoning agent tasks
“And actually we found that for most of the kind of like code writing and reasoning problems that we care about, the visual part isn't really a huge important part of it.”
Josh Albrecht Jun 25, 2024 ▶ 45:17 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.