reasoning benchmarks
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Mar 5, 2025 by Marco Mascorro · across every show →
Everything said about reasoning benchmarks, oldest first
Mar 5, 2025 neutral
Mascorro: DeepSeek-R1-Zero Improved Math Scores but Struggled with Readability and Language Switching
“R one zero, which in a way was a very interesting model because it showed that it improved in some reasoning benchmarks and math benchmarks. But eventually didn't do really well on other things, right? Like it was switching between languages. I think that was …”