DeepSeq
company on 7 shows · 8 statements across 8 episodes
More or Less
We Live to Build
Latent Space
No Priors
the MAD Podcast
the a16z Podcast
TBPN
8 statements about DeepSeq, every show
DeepSeek models are the hardest to support due to architectural novelties
“I think that, like, obviously the DeepSeq models tend to be the most challenging as they have, like, the most novel architectural stuff going on model over model.”
Malde: Western open-source AI lags Chinese models at trillion-parameter scale
“I think America or the Western world has some work to do still. Like obviously having one trillion parameter models like Kimi or like an amazing models like GLM and DeepSeq, I don't think we're quite there yet for that size of model.”
Gil: Chinese open-source AI models rank among highest on benchmarks
“Some of the highest
Scoring models against benchmarks now are Chinese models on the open source side.
On the closer side, it's still a lot of the US models, but things like Quinn, DeepSeq, et cetera, are doing very well.”
Gonzalez: ChatGPT is the best generative AI tool, followed by DeepSeek
“I would say the best one of all of them is ChatGBT and then DeepSeq. That's for sure, because that's where a lot of the resources are coming from, and many GBTs come from ChatGBT.”
DeepSeek runs each single model replica across more than 300 GPUs
“DeepSeq actually that company itself was running and still running this model over more than 300 GPUs. So think about this deployment. One replica is 300 GPUs, and there are so many different, so many more replicas.”
Jessica Lessin: Alibaba and Chinese Government Proposed Investing in DeepSeek
“The information did break the news today that Alibaba and the Chinese government are proposing that they invest in DeepSeq.”
Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Soldani: Frontier LLM pre-training requires at least 50,000 GPUs
“To give you a sense of, like, how I personally think about research budget for each part of the language model pipeline is, like, on the pre-training side, you can maybe do something with a thousand GPUs. Really, you want 10,000. And, like, if you want real es…”