DeepSeq

company on 7 shows · 8 statements across 8 episodes

More or Less We Live to Build Latent Space No Priors the MAD Podcast the a16z Podcast TBPN

8 statements about DeepSeq, every show

DeepSeek models are the hardest to support due to architectural novelties
“I think that, like, obviously the DeepSeq models tend to be the most challenging as they have, like, the most novel architectural stuff going on model over model.”
Philip Kiely Aug 3, 2026 ▶ 16:01 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Malde: Western open-source AI lags Chinese models at trillion-parameter scale
“I think America or the Western world has some work to do still. Like obviously having one trillion parameter models like Kimi or like an amazing models like GLM and DeepSeq, I don't think we're quite there yet for that size of model.”
Ronak Malde Jun 21, 2026 ▶ 14:41 ⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
NO PRIORS Assertion Supported
Gil: Chinese open-source AI models rank among highest on benchmarks
“Some of the highest Scoring models against benchmarks now are Chinese models on the open source side. On the closer side, it's still a lot of the US models, but things like Quinn, DeepSeq, et cetera, are doing very well.”
Elad Gil Jan 8, 2026 ▶ 15:14 NVIDIA’s Jensen Huang on Reasoning Models, Robotics, and Refuting the “AI Bubble” Narrative
Gonzalez: ChatGPT is the best generative AI tool, followed by DeepSeek
“I would say the best one of all of them is ChatGBT and then DeepSeq. That's for sure, because that's where a lot of the resources are coming from, and many GBTs come from ChatGBT.”
Alexandra Gonzalez Jul 22, 2025 ▶ 28:52 She Managed $6B Across 220 Countries - Then Knocked on Old Bosses' Doors
MAD Assertion Not checkable as stated
DeepSeek runs each single model replica across more than 300 GPUs
“DeepSeq actually that company itself was running and still running this model over more than 300 GPUs. So think about this deployment. One replica is 300 GPUs, and there are so many different, so many more replicas.”
Lin Qiao Mar 27, 2025 ▶ 36:10 Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
MORE OR LESS Assertion Open · timeframe Feb 2025
Jessica Lessin: Alibaba and Chinese Government Proposed Investing in DeepSeek
“The information did break the news today that Alibaba and the Chinese government are proposing that they invest in DeepSeq.”
Jessica Lessin Feb 21, 2025 ▶ 18:52 #87: SKIMS To Revive Nike? AI Drama & The Satoshi Mystery · More or Less Podcast
LATENT SPACE Assertion Contradicted
Yining Zhang: DeepSeek V3 scores 94.6 on GSM8K, outperforming Llama 405B
“Yeah, I think even they use the FP-A to quantization, the benchmark result is very good, such as something like GSM-HK. The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source AIM, even the LAMA 400 zero five billion.”
Yining Zhang Jan 19, 2025 ▶ 12:38 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Soldani: Frontier LLM pre-training requires at least 50,000 GPUs
“To give you a sense of, like, how I personally think about research budget for each part of the language model pipeline is, like, on the pre-training side, you can maybe do something with a thousand GPUs. Really, you want 10,000. And, like, if you want real es…”
Luca Soldani Dec 23, 2024 ▶ 11:11 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.