Reward Model Loss Function
topic on 1 show · 1 statements across 1 episodes
1 statements about Reward Model Loss Function, every show
Lambert: Anthropic and OpenAI reward model loss functions are mathematically identical
“Fun fact is that these loss functions Look different and anthropic in opening eyes papers, but they're just literally just log transform. So if you start like expantiating both sides and taking the log of both sides, you'll like converge on one of the two, the…”