Pre Training Paradigm
topic on 1 show · 2 statements across 1 episodes
2 statements about Pre Training Paradigm, every show
Hendrycks: RL-based reasoning models are improving faster than pre-training did
“That is separate from the new reasoning paradigm that has emerged in the past year which is where you train models to on math and coding types of questions with reinforcement learning, and that has a very steep slope, and I don't see any signs of that slowing …”