Reinforcement Learning From Code Execution Feedback

topic on 1 show · 1 statements across 1 episodes

the MAD Podcast

1 statements about Reinforcement Learning From Code Execution Feedback, every show

MAD Insight
Kant: Programmatic RL can scale magnitudes larger than human feedback
“And it's that RL loop that is very interesting because since it's programmatic, since we have an Oracle of truth, we can scale this up far larger, right? Magnitudes larger than what you can do with human feedback today.”
Eiso Kant Dec 20, 2023 ▶ 21:27 The Race to Build the Ultimate AI Programmer | Poolside CTO Eiso Kant

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.