DeepSeek

includes DeepSeek R1, DeepSeek V3, DeepSeek OCR, DeepSeek Code, DeepSeek R One, DeepSeek Reasoner Model, DeepSeek V3.2

23 statements across 19 episodes · 17 bullish · 3 bearish · 18 people on the record · first statement May 31, 2024 by Mark Huang · said 307 times in 69 episodes since 2024 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (35), Alessio Fanelli (19), Yining Zhang (18), Nathan Lambert (11), Ahmad Awais (10), Kyle Kranen (7), Elie Bakouch (7), William Beauchamp (6)

tap a year for its mentions
001252525050202420252026episodesmentions
02550202420252026episodes it came up in
002.525550202420252026episodesmentions per episode
2026 69 mentions in 15 episodes 5 per episode
2025 231 mentions in 48 episodes 5 per episode
2024 7 mentions in 6 episodes 1 per episode

every mention, scene by scene, with the transcript →

Everything said about DeepSeek, oldest first

May 31, 2024 positive
Opinion
Huang: DeepSeek's Multi-Head Latent Attention is novel and underrated
“Underrated specific instance would be, like, the deep seek paper. I'd never seen it before, but, like, the multi-head latent attention, like, that was really unexpected to me because, like, I thought I'd seen every Not every type, obviously, but, like, every w…”
Mark Huang May 31, 2024 ▶ 1:08:55 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Dec 23, 2024 bullish
Assertion Supported
Soldani: 2024 open models rival frontier performance of closed models
“You have models that are, you know, reveling frontier level performance of what you can get from closed models, from like Quen, from Deep Seek. We got Lama III.”
Luca Soldani Dec 23, 2024 ▶ 1:15 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Jan 2, 2025 negative
Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Nathan Lambert Jan 2, 2025 ▶ 6:26 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Jan 19, 2025 positive
Assertion Supported
Zhang: DeepSeek team officially recommends SGLang as its inference engine
“That's why SGLAN is the recommended LLM engine by the DeepSeq team.”
Yining Zhang Jan 19, 2025 ▶ 28:06 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Jan 19, 2025 positive
Assertion Not checkable as stated
Zhang: Baidu and ByteDance Internal Models Use DeepSeek-Like MoE Architectures
“As far as I know, some companies such as Baidu or Baidu Dance, they are internal, the dominant AOM, they use the MOE architecture, and their ways, I think, is similar to the DeepSeq MOE model.”
Yining Zhang Jan 19, 2025 ▶ 13:41 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Jan 26, 2025 positive
Assertion Supported
Beauchamp: DeepSeek slashes inference costs by shrinking KV cache
“What DeepSeek have achieved that's quite special is they've got this amazing inference engine. They've been able to reduce the size of the KV cache significantly. And then by being able to do that, they're able to significantly reduce their inference costs.”
William Beauchamp Jan 26, 2025 ▶ 23:30 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Feb 6, 2025 positive
Assertion Supported
Colvin: AI providers are centralizing around OpenAI's API standard
“I think the truth is that everyone is centralizing around OpenAI's SD API as the one to do. So DeepSeek support that. Grok with a K support that. Olama also does it. Well, I mean, if there is that library right now, it's more or less the OpenAI SDK.”
Samuel Colvin Feb 6, 2025 ▶ 30:32 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Mar 7, 2025 positive
Assertion Not checkable as stated
DeepSeek Revealed Attention Kernel Optimizations Kept Secret by Big Labs
“Deep seeks recent open sourcing of their various code components that they use to train that model, which I think outside of the big labs was not really well known to, right, it was not really well known how to write kind of a kernel that's optimized for this …”
Misha Laskin Mar 7, 2025 ▶ 13:37 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Mar 23, 2025 neutral
Assertion Supported
Agarwal: DeepSeek Distilled Reasoning Models Using Correctness-Filtered Synthetic Data
“What they did was they took the best model they had, they generated a bunch of samples, they filtered them based on correctness, because these were on tasks like coding, mathematical problem solving, and a bunch of those things. So they saw whatever, what are …”
Rishabh Agarwal Mar 23, 2025 ▶ 13:57 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Mar 23, 2025 positive
Insight
Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another mod…”
Rishabh Agarwal Mar 23, 2025 ▶ 14:26 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Jun 13, 2025 positive
Opinion
DeepSeek's open releases pulled global AI progress forward by six months
“But I think that what it did is it pulled forward progress in AI by like six months.”
Chris Lattner Jun 13, 2025 ▶ 57:54 The Shape of Compute (Chris Lattner of Modular)
Jul 2, 2025 positive
Insight
Morris: Weight deltas can reconstruct a competitor's proprietary fine-tuning dataset
“There's some tricks to it, but it's basically just like gradient based selection based on this weight difference. And it seems to be okay. Like it can get us pretty good training data. So I guess if you actually wanted to use this, it would be like your compet…”
Jack Morris Jul 2, 2025 ▶ 1:04:42 Information Theory for Language Models: Jack Morris
Jul 31, 2025 neutral
Assertion Supported
Lambert: DeepSeek-R1 starts solving math questions immediately without explicit planning
“If you look at DeepSeq R-One and you ask it a hard math question, it's not like, here's my plan of attack. It just starts.”
Nathan Lambert Jul 31, 2025 ▶ 40:10 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Oct 20, 2025 negative
Assertion Not checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Elie Bakouch Oct 20, 2025 ▶ 59:13 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Nov 3, 2025 positive
Assertion Supported
AMD funded community DeepSeek kernel development through the GPU Mode Discord
“AMD's actually done this. Like there's some deep seek, like through GPU modes, discord, there have been some like deep seek kernels that they say you have, you know, guaranteed access to compute for, write them.”
Quentin Anthony Nov 3, 2025 ▶ 57:48 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Nov 14, 2025 positive
Assertion Supported
Mechanistic Interpretability Scales Without Bottlenecks to Large Models Like DeepSeek
“There's no gap for scale. Like, they've shown that even for the biggest open source models, you like, even like DeepSeq's big models, you, they can do it. And then in general, like, scaling is not the bottleneck.”
Deedy Das Nov 14, 2025 ▶ 52:11 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Dec 30, 2025 bullish
Assertion Not checkable as stated
Nair: OpenAI already possessed a superior model during the DeepSeek release
“The feeling in OpenAI is that like, well, I think we had a better model already at the time, right?”
Ashvin Nair Dec 30, 2025 ▶ 31:21 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Dec 31, 2025 positive
Insight
McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know…”
Josh McGrath Dec 31, 2025 ▶ 12:42 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Mar 8, 2026 positive
Assertion Supported
DeepSeek 128k context fits in 8GB KV cache versus Llama's 80GB
“For context like the total I think the total context length of DeepSeq is a 128,000 tokens, or it might be 256,000 with rope extension. That entire context, I think it's a 128,000, fits into eight gigabytes. And previously context, like I think the Lama four o…”
Kyle Kranen Mar 8, 2026 ▶ 52:08 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Jun 3, 2026 neutral
Assertion Contradicted
Hong: DeepSeek dissolved its formal reasoning team over strategic shift
“And we have since, for example, Deep Seek All right. Like originally having a formal team and then later dissolve that team because of strategic direction change.”
Carina Hong Jun 3, 2026 ▶ 1:14:16 Scaling Past Informal AI - Carina Hong, Axiom Math
Jun 6, 2026 positive
Assertion Not checkable as stated
Awais: Claude 3.7 Max is already CommandCode's second most used model
“But they will only be for deep seek when to 3.7 max is the second most used model on command code right now. It's just two or three days old.”
Ahmad Awais Jun 6, 2026 ▶ 39:43 ⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
Jun 6, 2026 negative
Assertion Not checkable as stated
Awais: Claude Code hides 50+ tool-call failures per session on DeepSeek
“In CloudCode you know, they hide a lot of the errors behind control O, right? So you don't even know that, you know, you have like 50 plus tool call failures plus per session. You're just sitting there and you're like, oh, why is DeepSeq so slow?”
Ahmad Awais Jun 6, 2026 ▶ 12:21 ⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
Jun 6, 2026 bullish
Disclosure
CommandCode offers 600 million DeepSeek tokens for $1 per month
“We launched a Go plan with just dollar one per month to, where you can do like, six hundred million tokens of DeepSQL for Pro in it, just to prove like, open models are actually really, really good, and they are catching up, right?”
Ahmad Awais Jun 6, 2026 ▶ 16:44 ⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.