Off Policy Learning
topic on 1 show · 1 statements across 1 episodes
1 statements about Off Policy Learning, every show
Nair: 2017–2022 academic RL breakthroughs failed because researchers overfit to benchmarks
“A lot of the methods that people were really excited about is, like you know, off policy learning, like, value functions, like, these kind of things, and somehow that, that stuff hasn't really panned out, I would say, and it's not exactly clear why, but in the…”