Reasoning Benchmarks
topic on 1 show · 1 statements across 1 episodes
1 statements about Reasoning Benchmarks, every show
Mascorro: DeepSeek-R1-Zero Improved Math Scores but Struggled with Readability and Language Switching
“R one zero, which in a way was a very interesting model because it showed that it improved in some reasoning benchmarks and math benchmarks. But eventually didn't do really well on other things, right? Like it was switching between languages. I think that was …”