DeepSeek, every mention
167 scenes, the whole family · ← back to DeepSeek
tap a year for its mentions
every year anyone Shawn Wang 35Alessio Fanelli 19Yining Zhang 18Nathan Lambert 11Ahmad Awais 10Kyle Kranen 7Elie Bakouch 7William Beauchamp 6Rishabh Agarwal 6Philip Kiely 6
Verbatim, from the transcripts: the passages where DeepSeek comes up
The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 16:01 Philip Kiely And then when there's a new model with a newer architecture, I think that, like, obviously the DeepSeq models tend to be the most challenging as they have, like, the most novel architectural stuff going on, uh, model over model. 2 times in the scene
- ▶ 20:12 Philip Kiely And, and ultimately what you get out of the system is all of a sudden you have Kimi vision, GLM weights and deep seek attention all in one model.
- ▶ 1:08:32 Alessio Fanelli I think on your guys' end, you see a lot of, okay, one day it's GLM, Kimmy, DeepSeq, uh, Minimax, throw in the others.
- ▶ 1:13:35 Philip Kiely And so, for example, when DeepSeq R-One came out, it was, you know, it was six hundred seventy one billion parameters, which at the time was really huge, and I think did a lot to push us to really quickly adopt Blackwell and get good at…
- ▶ 1:34:48 Philip Kiely Just coding model that we had access to, and it would do an equally good job of optimizing, uh, DeepSeq or Kimi or something. 2 times in the scene
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
- ▶ 11:15 unnamed speaker Deep Seek of the West.
- ▶ 24:42 unnamed speaker We're very much like, ah, okay, look, it's like, you know, on par with Kimmy, DeepSeek, whatnot, the small ones, Gemma level.
- ▶ 1:16:45 unnamed speaker I will call out that one of the branches of research is DeepSeq OCR, which is, can you just throw away the text tokenizer and just only vision?
- ▶ 1:18:50 Eiso Kant And I love what deep seek and others are trying.
- ▶ 1:20:42 Eiso Kant Uh, and, uh, you know, you started with Deep Seek of the West and, and, and, uh, I think that's, uh, 3 times in the scene
⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
- ▶ 14:26 Ronak Malde And then all of a sudden, DeepSeq v three comes out.
- ▶ 14:29 Ronak Malde And everyone is like, wow, they're using frontier paradigms. 2 times in the scene
⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
- ▶ 4:48 Ahmad Awais So it is the 25th of May today, and about literally 25 days ago, everybody was talking about how DeepSeek is so good. 4 times in the scene
- ▶ 9:21 Ahmad Awais So this was me being like, you know, I just discovered why and how Deepsea can outperform Opus. 2 times in the scene
- ▶ 13:36 unnamed speaker Being thoughtful, like, I wonder if it's only DeepSeq that's like showing this kind of stuff, or is it just a general open models trick? 4 times in the scene
- ▶ 16:24 Ahmad Awais It's like, so one of the first things that happened was I think we might be the most used coding agent out there right now for DeepSeq. 4 times in the scene
- ▶ 39:16 unnamed speaker Uh, you know, I, I, I was reminded that actually DeepSeek announced that they're going to do DeepSeek code at some point. 4 times in the scene
- ▶ 39:16 unnamed speaker Uh, you know, I, I, I was reminded that actually DeepSeek announced that they're going to do DeepSeek code at some point.
Scaling Past Informal AI - Carina Hong, Axiom Math
- ▶ 14:31 Carina Hong Mass Arena, which is this organization that evaluates a lot of LLMs, found the best LLM, Deep Seek, got 103 points out of a 120 point exam.
- ▶ 1:14:16 Carina Hong And we have since, for example, Deep Seek
⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
- ▶ 23:01 Shawn Wang I think it was Deep Link.
⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
- ▶ 16:18 unnamed speaker You get access to Deep Seek, Meta, Moonshot, and up here somewhere it also mentions, 24 hour fine tuning.
Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
- ▶ 10:46 Marc Andreessen It was O-one and then R-one that basically answered that question and basically said, oh, no, we're going to be able to actually turn this into something that's going to work in the real world.
- ▶ 29:22 Marc Andreessen Um, you know, I think deep seek was like a gift to the world. 2 times in the scene
- ▶ 29:56 Marc Andreessen And then our one comes out and it's just like, there's the code and there's the paper.
Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
- ▶ 38:37 Guillaume Lample Uh, I think, like, the DeepSeq one also, like, very, a lot of impact.
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 45:28 Kyle Kranen It's like, like a DeepSeq style architecture. 5 times in the scene
- ▶ 1:14:18 Kyle Kranen Uh, you can run a lolly seek. 2 times in the scene
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- ▶ 36:48 Shawn Wang Honestly, like, uh, it's very interesting, this, uh, sort of moonshot AI, and this is a tangent, we're not, we're not really gonna focus on this very much, but, um, you know how, like, the, the sort of AI tigers out of China were, were… 2 times in the scene
The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
- ▶ 18:16 Shawn Wang So the, a simple example was vision, um, can on a pixel level encode text and DeepSeek had this, uh, DeepSeek LCR paper that did that.
Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
- ▶ 10:00 Mark Bissell Like you look at Quen or, um, R-one and they have sort of like this CCP bias in them and-
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 25:23 George Cameron And the Deep Seek jump. 4 times in the scene
- ▶ 25:36 Micah Hill-Smith Well, a couple of weeks, it was, it was Boxing Day in New Zealand, uh, when, when Deep Seek v three came out and I like, we'd been tracking Deep Seek and a bunch of the other global players that were less known. 3 times in the scene
- ▶ 26:23 Micah Hill-Smith Um, the world really noticed when they followed that up with the RL working on top of E three and R one succeeding like a few weeks later.
- ▶ 1:04:09 Shawn Wang Deep Seek was a major pusher of fine grained experts, let's call it.
- ▶ 1:12:12 Shawn Wang I took the latest deep seek paper and, uh, I, you know, they had some descriptions of their search agents and their coding agents.
[State of Research Funding] Beyond NSF, Slingshots, Open Frontiers — Andy Konwinski, Laude Institute
- ▶ 19:21 Andy Konwinski Moonshot, Kimmy, Deep Seek.
[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv
- ▶ 5:12 unnamed speaker Oh yeah, DeepSeek. 3 times in the scene
- ▶ 5:30 unnamed speaker And those might be a little bit more expensive if, if you're using their API offering, but I think like Deep Seek OCR is very, very good.
- ▶ 12:05 unnamed speaker Compared to, like, it's not state-of-the-art, but you compare that to, like, Deep Seek R-one or, you know, Gemini 2.5 is, like, pretty good.
- ▶ 20:47 unnamed speaker Was that DeepSeek? 6 times in the scene
- ▶ 20:49 unnamed speaker Literally, DeepSeek R one, or...?
[State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
- ▶ 15:06 Shawn Wang So, so I, uh, Deep Seek came out with, uh, V 3.2 recently. 2 times in the scene
- ▶ 15:06 Shawn Wang So, so I, uh, Deep Seek came out with, uh, V 3.2 recently.
[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
- ▶ 12:24 unnamed speaker Um, GRPO came out in the DeepSeq math paper, which, uh, when it came out, I read it, and I was like, okay, this is kind of cool. 2 times in the scene
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 30:36 unnamed speaker Um, did the Deep Seek moment this year, also this year, crazy, uh, change anything internally? 3 times in the scene
The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
- ▶ 30:57 Shawn Wang Like deep seek.
Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
⚡ Inside GitHub’s AI Revolution: Jared Palmer Reveals Agent HQ & The Future of Coding Agents
- ▶ 23:32 unnamed speaker Uh, especially with open vision models like DeepSeq OCR and, and Omo OCR, uh, it, it, like, just give it a few more turns of the scaling.
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 7:08 unnamed speaker I think when Deep Sea Card One came out, everybody was, like, going crazy about, oh, they went, like, so deep in, like, the actual GPU code to, like, improve things.
- ▶ 52:42 unnamed speaker I think there's obviously been the rise of DeepSeq since then. 3 times in the scene
- ▶ 54:06 unnamed speaker So between understanding, I don't know, R-one and understanding a Luther train model that is like fully open, what, what's the delta between the two and the amount of work that you can do? 3 times in the scene
⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
- ▶ 8:09 Elie Bakouch The rest of it, I think the model architecture, we, with, like, the Quen, Deepsea, Kimi architecture, we're already at a point where this is already not optimal, but this is, like, pretty, pretty advanced. 3 times in the scene
- ▶ 8:54 Elie Bakouch And for example, a good, uh, a good way to view that is that, uh, DeepSeq rig three is still using the same Adam parameter than, uh, Lama two.
- ▶ 29:24 Elie Bakouch And I think DeepSeq is using that as well.
- ▶ 38:07 Elie Bakouch But yeah, if I go back here, there is this, um, this NSA, which is like the, the new deep seek attention.
- ▶ 59:24 Elie Bakouch To, because basically the, the, the, the coin tree or even deep seek, they don't release, release all the ablation data.
Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
- ▶ 21:01 Shawn Wang It was very strange because I think when deep seek first, like started talking about it, it was viewed as an optimization.
A Technical History of Generative Media
- ▶ 14:22 Batuhan Taskaya You guys see this on like, you know, even for stuff like, you know, QAn, DeepSeq, whatever, people want, like, even if,
Better Data is All You Need — Ari Morcos, Datology
- ▶ 38:29 Alessio Fanelli Same with deep seek.
- ▶ 42:40 Ari Morcos But as soon as you break that assumption, and I think deep seek showed that already you can get a frontier model for a marginal cost of a couple million dollars, um, that's gone down.
Long Live Context Engineering - with Jeff Huber of Chroma
- ▶ 36:12 Jeff Huber I think it was called, unfortunately, Ragar one, where they like teach, uh, deep cigar one, you know, kind of give it the tool of how to retrieve.
⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
- ▶ 32:17 Mitesh Agrawal I mean, even the attention mechanism changing when DeepSeq, well, not changing, but like if DeepSeq opened it to the world, which is the MHLA stuff, I think if you had burned or etched in the, the architecture of doing Transformer the way… 3 times in the scene
Greg Brockman on OpenAI's Road to AGI
- ▶ 47:41 unnamed speaker Sliding window attention, the very fine-grained mixture of experts, which I think DeepSeek popularized, rope, yarn, attention sinks, any, anything that, you know, I think stood out to you and, uh, the choices made for GPT-OSS?
The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
- ▶ 13:38 Shawn Wang So, so far, Deep Seek, obviously, like, still one of the biggest news of the, the year, what was a big gift to, to, uh, to, to the, the inference providers and the sort of relative decline of Lama and the disappointment Lama four was, uh,… 3 times in the scene
- ▶ 14:12 Stephanie Palazzolo Or, I even, I know even whenever, like, the DeepSeek R-One model first came out, one of the major, like, startups in the space was, like, yeah, 75% of the usage that we see for R-One is people trying to distill it into, like, 2 times in the scene
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 5:40 Nathan Lambert No, that was like in it being taking off because it was after deep seek, but it's like when people like that have the acronym on the slides and that's, it's also very clear of like RLHF is four letters.
- ▶ 18:52 Nathan Lambert And then deep seek R one is still the canonical recipe on a like reasoning only model.
- ▶ 19:19 Shawn Wang It's kind of cool that like people are, you know, taking variations on it, but also I don't know if DeepSeek is going to come up with R two and just blow away everyone with whatever is next.
- ▶ 23:05 Nathan Lambert It's like DeepSeq R-one to the new R-one, it goes down.
- ▶ 38:37 Nathan Lambert What I would say that we have already done with O-one and R-one, which is you do a lot of RL, you show the inference time scaling works and you get really high benchmark numbers. 2 times in the scene
- ▶ 41:23 Nathan Lambert Like, I don't think Deep Seek has, doesn't have it built in, but it probably could do it.
- ▶ 46:57 Shawn Wang There's one case where, with O-one and sort of the, the sort of Q-star ideas, there was one case where it was sort of overhyped in some sense, but now it's coming back with O-one Pro and DeepThink. 2 times in the scene
- ▶ 1:00:15 Nathan Lambert Which DeepSeq mentioned, but that's one thing to go. 2 times in the scene
- ▶ 1:09:34 Nathan Lambert It's like DeepSeq.
- ▶ 1:15:48 Alessio Fanelli Um, any parting thoughts on how you're gonna build the American Deep Seek? 5 times in the scene
⚡️Using RFT to Build Clinical Superintelligence
- ▶ 5:59 Brendan Fortuna And it's the same technique that's used to train these state-of-the-art reasoning models, like O-three, you know, R-one, uh, Claude-four.
AI is Eating Search
- ▶ 53:14 Robert McCloy And if you look, like, go look back at DeepSeq, look at when Grok first launched a mobile app, things like that, and you can look at similar, you can look at similar data from, like, similar web, for example. 2 times in the scene
Cline: The Collaborative AI Coder
- ▶ 46:25 unnamed speaker Like maybe if your organization you're forced into using like a very small model, that's not very good at search and replace like a deep seek or something.
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
- ▶ 6:45 Alessio Fanelli For me, it was Mistral Small being the best open source model above Foro and DeepSeq VIII.
Information Theory for Language Models: Jack Morris
- ▶ 1:03:19 Jack Morris The first thing is we assume access to two checkpoints, which I think is probably not the case in Gemma, but in the, in the case of deep seek, if you download the four hundred billion parameter model weights, it's this giant file and you… 3 times in the scene
- ▶ 1:05:34 Shawn Wang Decently often for the open model labs, like even Deep Seek R R one like has released an update.
Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
- ▶ 15:58 Spooks (Swyx) And one of the famous findings from DeepSeek is that MCTS wasn't that useful to them.
The Shape of Compute (Chris Lattner of Modular)
- ▶ 19:56 Chris Lattner And so you have a lot of crazy forms of attention, like the deep seek things that just came out and like all this stuff is always changing.
- ▶ 43:01 Shawn Wang Um, I, I don't know how to make this happen, but like, I think you win when Mistral, Meta, Deep Seek, and Quinn adopts you and like ship you natively, right?
- ▶ 52:50 Alessio Fanelli There's the more recent, maybe open source thing, which is DeepSeq, obviously. 9 times in the scene
⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
- ▶ 13:09 Alex Duffy Well, if you Deep Seek, is this a Deep Seek game? 4 times in the scene
- ▶ 13:20 Alex Duffy And so what you're looking at, just for context, you're playing as Deep Seek reasoner, which you can see in the bottom right.
- ▶ 24:22 unnamed speaker Here's what Deep Seek's personality is.
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 27:04 Will Brown I think my last, I don't know, this isn't like, I wasn't the first person to do this, but like, it was pretty clear to me, like after one and before R one, that like RL was going to work and that that was going to intersect with agents…
- ▶ 30:03 Will Brown A model like R one, like a hundred percent of the time it is going to use its think tokens.
- ▶ 34:16 Will Brown If you're doing like an R-one and people are like, oh, math is easy to verify.
⚡️Open Questions in Agentic RL — Will Brown (Prime Intellect)
- ▶ 0:48 Will Brown We have things like R-one, which are great at kind of the single-term math and code reasoning problems.
- ▶ 2:55 Will Brown Single-turn RLVR world, the models like O-one and R-one, and we want to, like, have these things become more agentic, and it seems like the path is to incorporate reinforcement learning into this process.
What is an RL environment? w/ Nous Research's Roger Jin
⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
- ▶ 10:03 unnamed speaker So for instance, one of the interesting things we saw was that DeepSeq was quite decent at lab play. 2 times in the scene
- ▶ 10:58 Shawn Wang Um, you know, I think it's very significant that Claude is so much better than DeepSeq and OpenPlay.
page 1 of 2 · 100 scenes per page · newest episode first next →