reinforcement learning
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Jan 31, 2025 by Dave Morin · across every show →
Everything said about reinforcement learning, oldest first
Jan 31, 2025 positive
Morin: DeepSeek's paper unifies fine-tuning and optimization in one RL equation
“They do have this one part of one of the papers that talks about a unified paradigm around reinforcement learning, which is pretty cool. They've got a Very beautiful piece of math that kind of brings together all these different fine tuning and optimization te…”