reinforcement learning with verifiable rewards

1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Nov 20, 2025 by Nathan Lambert · across every show →

Everything said about reinforcement learning with verifiable rewards, oldest first

Nov 20, 2025 neutral
Disclosure
Lambert: AI2 coined 'reinforcement learning with verifiable rewards' replicating Llama 3
“We spent a long time to try to replicate what we thought was close to Lama three post training with multiple stages and optimizers, which is the project that like came up with the name reinforcement learning with verifiable rewards with a bunch of people.”
Nathan Lambert Nov 20, 2025 ▶ 29:22 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.