DeepSeq
4 statements across 4 episodes · 1 bullish · 1 bearish · 4 people on the record · first statement Dec 23, 2024 by Luca Soldani · across every show →
Everything said about DeepSeq, oldest first
Dec 23, 2024
Soldani: Frontier LLM pre-training requires at least 50,000 GPUs
“To give you a sense of, like, how I personally think about research budget for each part of the language model pipeline is, like, on the pre-training side, you can maybe do something with a thousand GPUs. Really, you want 10,000. And, like, if you want real es…”
Jan 19, 2025 positive
Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Jun 21, 2026 negative
Malde: Western open-source AI lags Chinese models at trillion-parameter scale
“I think America or the Western world has some work to do still. Like obviously having one trillion parameter models like Kimi or like an amazing models like GLM and DeepSeq, I don't think we're quite there yet for that size of model.”