Reinforcement learning from code execution feedback

1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Dec 20, 2023 by Eiso Kant · across every show →

Everything said about Reinforcement learning from code execution feedback, oldest first

Dec 20, 2023 positive
Insight
Kant: Programmatic RL can scale magnitudes larger than human feedback
“And it's that RL loop that is very interesting because since it's programmatic, since we have an Oracle of truth, we can scale this up far larger, right? Magnitudes larger than what you can do with human feedback today.”
Eiso Kant Dec 20, 2023 ▶ 21:27 The Race to Build the Ultimate AI Programmer | Poolside CTO Eiso Kant
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.