LLM Benchmarks

topic on 2 shows · 2 statements across 2 episodes

Latent Space the a16z Podcast

2 statements about LLM Benchmarks, every show

LATENT SPACE Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Josh Albrecht Jun 25, 2024 ▶ 1:01:43 State of the Art: Training 70B LLMs on 10,000 H100 clusters
a16z Opinion
Ghodsi: All current LLM evaluation benchmarks are completely bullshit
“I kind of think all the benchmarks are bullshit, and so all these, so all the LLM benchmarks, here's how it works.”
Ali Ghodsi Sep 25, 2023 ▶ 17:52 AI Food Fights in the Enterprise with Databricks' Ali Ghodsi

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.