UltraFeedback

2 statements across 2 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Jan 11, 2024 by Nathan Lambert · said 3 times in 2 episodes since 2024 · across every show →

Mentions by year

brought up most by Nathan Lambert (3)

tap a year for its mentions
00112120242025episodesmentions
01120242025episodes it came up in
0010.52120242025episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about UltraFeedback, oldest first

Jan 11, 2024 neutral
Assertion Supported
Lambert: DPO benchmark gains rely largely on the UltraFeedback dataset
“Everyone's using this ultra feedback data set and it boosts AlpacaVal, MTBench, TruthfulQA, and like the qualitative model a bit. We don't really know why.”
Nathan Lambert Jan 11, 2024 ▶ 29:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jul 31, 2025 neutral
Assertion Not checkable as stated
Lambert: Academia relied on UltraFeedback for open preference tuning for a year
“The academic community had been using this one data set since like all the way back in the hugging face models of like Zephyr beta is when this ultra feedback data set got popular. And still a year later is like this state of the art data set for open preferen…”
Nathan Lambert Jul 31, 2025 ▶ 3:15 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.