Disclosure certainty 4/5 debate potential 1/5

Albrecht: Imbue Manages Infrastructure with Three to Six Engineers

Josh Albrecht · State of the Art: Training 70B LLMs on 10,000 H100 clusters · Jun 25, 2024 · at 28:28

Josh Albrecht, CTO of Imbue, discusses the lean team size supporting Imbue's large-scale AI compute infrastructure.

0:00 / 0:13exact quote · 13.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Like our infrastructure team is like You know, it fluctuates from week to week, depending on like how many things are on fire and how much we need to build. But it's like between like three and six people, like it's small. It's not like some huge team of like tons and tons of engineers.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Josh Albrecht

Insight
Albrecht: LLM emergence is an artifact of non-linear evaluation metrics
“This emergent behavior that you're seeing, Is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
Josh Albrecht Jun 25, 2024 ▶ 55:56 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Opinion
Albrecht: Training on AWS prevents diagnosing low-level hardware errors
“And if we're just using, you know, AWS or some other cloud provider, These errors are still going to be there, and you're gonna have no way to know and no way to debug this and no way to diagnose what's going wrong.”
Josh Albrecht Jun 25, 2024 ▶ 19:49 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Josh Albrecht Jun 25, 2024 ▶ 1:01:43 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Albrecht: AgentBench paper's appendix examples are actually incorrect solutions
“Like we were looking at the agent bench paper, I think just last week for our paper club. And one of the things that we noticed is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not…”
Josh Albrecht Jun 25, 2024 ▶ 1:06:05 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Albrecht: Optimizing for competitive coding benchmarks does not create useful programmers
“Like, we do a lot of code generation, but we don't really do a lot on, like, code competition problems for the very, very hard ones, so that you can go very far down that route and make something like really good at those problems, but not actually that useful…”
Josh Albrecht Jun 25, 2024 ▶ 1:13:00 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Disclosure
Albrecht: Imbue avoids Kubernetes to keep cluster infrastructure simple to debug
“Less layers of infrastructure, less layers of abstraction, make it a lot easier to work with. Like we don't use Kubernetes, for example, I would just directly launch these things and it's just been much easier to debug this way.”
Josh Albrecht Jun 25, 2024 ▶ 34:38 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.