GSM-HK

2 statements across 2 episodes · 1 bullish · 0 bearish · 2 people on the record · first statement Jan 19, 2025 by Yining Zhang · across every show →

Everything said about GSM-HK, oldest first

Jan 19, 2025 positive
Assertion Contradicted
Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Yining Zhang Jan 19, 2025 ▶ 12:38 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Mar 23, 2025 neutral
Assertion Supported
Agarwal: Synthetic Data Distillation Can Outperform Logits on Benchmarks
“It's not always the case that synthetic data distillation beats, oh, sorry, it's worse than logits. Like, sometimes logits can be worse off. So if you look at the T-five base, two-fifty million scenario on GSM-HK, On the last plot, you can see that synthetic d…”
Rishabh Agarwal Mar 23, 2025 ▶ 26:20 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.