reward model loss function
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Jan 11, 2024 by Nathan Lambert · across every show →
Everything said about reward model loss function, oldest first
Jan 11, 2024 neutral
Lambert: Anthropic and OpenAI reward model loss functions are mathematically identical
“Fun fact is that these loss functions Look different and anthropic in opening eyes papers, but they're just literally just log transform. So if you start like expantiating both sides and taking the log of both sides, you'll like converge on one of the two, the…”