BoolQ

2 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Jun 25, 2024 by Josh Albrecht · said 2 times in 1 episodes since 2024 · across every show →

Mentions by year

brought up most by Josh Albrecht (2)

tap a year for its mentions
0011212024episodesmentions
0112024episodes it came up in
0010.5212024episodesmentions per episode
2024 2 mentions in 1 episode

every mention, scene by scene, with the transcript →

Everything said about BoolQ, oldest first

Jun 25, 2024
Disclosure
Albrecht: Imbue is releasing a new reasoning benchmark and 11 cleaned evaluations
“We're releasing a whole bunch of different data there, a new benchmark about code, reasoning, understanding, as well as our own private versions of 11 different open source benchmarks. So things like PoolQ or ANLI, where we've gone through and kind of cleaned …”
Josh Albrecht Jun 25, 2024 ▶ 7:23 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Jun 25, 2024 neutral
Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Josh Albrecht Jun 25, 2024 ▶ 1:01:43 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.