DeepSeek, every mention
167 scenes, the whole family · ← back to DeepSeek
tap a year for its mentions
every year anyone Shawn Wang 35Alessio Fanelli 19Yining Zhang 18Nathan Lambert 11Ahmad Awais 10Kyle Kranen 7Elie Bakouch 7William Beauchamp 6Rishabh Agarwal 6Philip Kiely 6
Verbatim, from the transcripts: the passages where DeepSeek comes up
Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
- ▶ 26:27 Charles Packer It's like pretty consistent across like both 3.7 deep seek, three mini, which all like the way you actually scale the x-axis here is fundamentally quite different in each case with 3.7 extended thinking mode.
- ▶ 26:45 Charles Packer And then to, I guess, enable this sort of like to send, to tell like R one on the API, whether or not to like go to 10 K tokens versus like two K tokens, um, also is like a, you know, different mechanism.
The #1 SWE-Bench Verified Agent
- ▶ 27:24 Guy Gur-Ari So I think the DeepSeq paper R one was very good understanding DPO and the variants like GRPO is very popular now.
The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
- ▶ 13:45 Rishabh Agarwal And I guess a better version of this is you can do something like best defend, which is, let's say this was something which was used in DeepSeq. 3 times in the scene
- ▶ 13:50 Rishabh Agarwal So, uh, if you saw that paper, they trained the DeepSeq reasoner model, and then they distilled the, the, basically that model to other models. 2 times in the scene
- ▶ 43:26 Rishabh Agarwal It's annoying to use, like, let's say if you're trying to distill from DeepSeq, you can't serve that huge model and have, like, a live teacher running because it is going to be expensive.
Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
- ▶ 13:37 Misha Laskin You know, I, I think, uh, I think there are going to be a lot of move 37, uh, and, you know, one of them, so that, right, uh, for example, we're building, right, we're training these, uh, large language models, and maybe, maybe kind of a… 2 times in the scene
Browserbase: Browser Infrastructure For Your AI Agents
- ▶ 1:19 Paul Klein I took my first vacation, actually, two weeks ago, and Operator came out on the first day, and then a week later, Deep Seek came out.
Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
- ▶ 15:05 unnamed speaker Oh, any takeaways from like deep seek? 3 times in the scene
S1: the $6 DeepSeek R1 Competitor (ft. Entropix)
- ▶ 0:12 unnamed speaker Uh, DeepSeek is all the rage. 3 times in the scene
- ▶ 0:25 unnamed speaker Uh, and we wind up looking for a guest and Tim has been blogging about R one, S one, um, and, and all the other stuff. 3 times in the scene
- ▶ 2:00 unnamed speaker Like we, we've debated having a deep sea blog, unless you actually even wrote one, but I was like, I don't think we're saying anything new here.
- ▶ 2:34 unnamed speaker Um, R one, uh, I was posting on it about, about it on blue sky
- ▶ 2:56 unnamed speaker Um, should we double click more on R-One, or are we, should we go into S-One? 2 times in the scene
- ▶ 3:06 unnamed speaker Um, so S-One came out more recently, and, um, we cloned Deep Sea for six dollars, is, is, uh, what the, the headline is? 3 times in the scene
- ▶ 13:07 unnamed speaker I mean, R one, R one did both are there. 2 times in the scene
Bee AI: The Wearable Ambient Agent
- ▶ 55:07 Ethan Sutin You do need a sonnet level model to, like, execute things like the agent, and, um, we will be having a subscription for, like, features like that, because it's, you know, although now with the R-one, like, we'll see, uh, we haven't… 2 times in the scene
- ▶ 55:22 Shawn Wang A deep seek?
smol agents are all you need
- ▶ 14:25 Aymeric (Emmerich) I tried R one, but R one is a bit under, uh, O one, uh, with small agents. 2 times in the scene
- ▶ 21:38 Swyx (Marcos Swix) January was already super exciting with DeepSeek and blowing it wide open.
- ▶ 21:46 Swyx (Marcos Swix) All the R-one derivatives, I think, are happening as well. 2 times in the scene
The AI Architect: Bret Taylor
- ▶ 1:33:29 Bret Taylor Now, uh, I just saw like someone post one of, they distilled one of the deep seek models and just made it really small and, you know, it's doing these chains of thoughts so fast, you know, it's achieving latency numbers, I think.
Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
- ▶ 29:35 Shawn Wang So for example, if I wanted to use Deep Seek, I'm out of luck because Pydance API doesn't have Deep Seek yet. 3 times in the scene
Why every AI Engineer needs an AI Gateway (ft Portkey.ai CEO)
- ▶ 2:59 Rohit Agarwal I think DeepSeq was the first model where within the 1:06 hours of it launching, we had so many pull requests on our 2 times in the scene
- ▶ 4:07 unnamed speaker And then they add Grok, they add Gemini, all these other things, they add DeepSeq, and then it starts to be a whole thing inside of the code base.
- ▶ 19:16 Rohit Agarwal So you could always say, okay, if it comes to node A, always hit GPT-FORO, and node B is DeepSeek R-ONE,
Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- ▶ 15:52 unnamed speaker Well, uh, the, uh, the, the sort of apparent metagame from Paul Gauffier on Ader is that you use R-One as an architect and Sonnet as a code model, and that's apparently the best combination that beats O-One for him.
- ▶ 34:15 Shawn Lewis The next question is, does our one do as well?
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 22:42 William Beauchamp And it's similar, I guess, to how DeepSeek have been able to produce such a compelling model when compared to someone like an open AI, right? 6 times in the scene
- ▶ 22:54 unnamed speaker V-three, yeah.
The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
- ▶ 1:24 Shawn Wang Yeah, so, okay, we're meeting to talk about what happened after DeepSeq's launch. 2 times in the scene
- ▶ 2:07 unnamed speaker So on Monday, when Monday morning, DeepSeek one was released, we were like thinking, okay, maybe we can use Curator. 4 times in the scene
- ▶ 7:10 unnamed speaker So with R one, it's the first time we actually get the reasoning traces in, in open source. 2 times in the scene
- ▶ 10:28 Shawn Wang Mostly I agree on, on this discussion, especially from the R-one paper. 2 times in the scene
- ▶ 11:58 Shawn Wang It seems like at least R-one is generalizing very well. 4 times in the scene
- ▶ 12:25 unnamed speaker So if you look at how some of DeepSeq's distilled models are trained, they have. 4 times in the scene
- ▶ 15:21 unnamed speaker The, the, the, the simple thing that we changed is changing out QWQ for DeepSeq R-One. 7 times in the scene
- ▶ 20:10 unnamed speaker Like, DeepSeek, they, they in fact said that at least for one of their R-one-zero models, like, they didn't even do SFT on it. 3 times in the scene
- ▶ 23:18 unnamed speaker I mean, it's amazing to see how since this R-one model came out every three hours or something.
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 0:43 unnamed speaker And, you know, you are lead software engineer on the model performance team, and you guys recently shipped DeepSeq v. three as one of the many models that you do host. 2 times in the scene
- ▶ 1:03 unnamed speaker So we can take this a number of directions, but I think one thing we wanted to just get off the bat on was to start with, uh, DeepSeq and, and maybe, and then we'll work our way backwards back to SGLang. 3 times in the scene
- ▶ 1:22 Yining Zhang Yeah, because, uh, DeepSeq VIII, I think, is currently considered the leading open source LLMs based on the benchmark results and the chat area results. 6 times in the scene
- ▶ 1:49 unnamed speaker One of the interesting things is, like, they are bootstrapped, like, uh, you know, very private, small lab. 2 times in the scene
- ▶ 4:56 Yining Zhang So before the DeepSig-Fee-Seventy-B, I think we haven't encountered that issue for that so large weights. 4 times in the scene
- ▶ 8:05 unnamed speaker It seems like, you know, I think Noam Shazir started talking about, um, sort of training natively quantize, and I think that's what DeepSea seems to have done, at least they said in their paper. 2 times in the scene
- ▶ 13:05 unnamed speaker Uh, speak too much about, do you see, because obviously we don't know, uh, that much on this, unless we work on the team, but 2 times in the scene
- ▶ 15:41 unnamed speaker Well, then one more thing, I guess, maybe more commercially relevant, um, D-Six API pricing is very competitive.
- ▶ 18:28 unnamed speaker Can you maybe quickly run people through how do you go from taking the deep seek V three weights to like actually run it? 2 times in the scene
- ▶ 26:57 Yining Zhang I think for the common use case, maybe not, not the DeepSeq VIII, for the common use case, I think SGLAN's performance is better than FLM, and its usability is better than TensorFlow TLM. 2 times in the scene
- ▶ 27:40 Yining Zhang And also in DeepSeq, oh, sorry, in the SGLAN version 0.4, we also support the DP attention for DeepSeq, and in the latest 3 times in the scene
- ▶ 34:25 Yining Zhang Yeah, we, we support some, uh, DeepSeq optimization, such as MLA optimization, DPR tension optimization, and we also support the serial overhead CPU schedule.
- ▶ 35:50 unnamed speaker And when you think about a model that is, you know, as large as DeepSeq v three, especially, uh, like having better KVcache reutilization is great.
- ▶ 50:03 Yining Zhang When we released the DeepSeq feed story support, we have some community user, something like Cursor. 2 times in the scene
- ▶ 55:36 unnamed speaker I think this is a really good dive into both BaseNet and SG Lang, and a little bit of DeepSig V three, which people are very interested in.
- ▶ 55:43 unnamed speaker I'm trying to talk to them as well, because obviously they're, uh, they're, they're a fascinating lab.
The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
- ▶ 6:21 Nathan Lambert The two I've highlighted are from Deep Seek and Quinn, and a lot of people in this room have probably seen them. 2 times in the scene
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 37:29 Shawn Wang How much of a moat is there in this like proprietary sort of training data that they've, uh, presumably accomplished because like even deep seek, it was able to do it and they had, you know, two months notice to do this, to do R one.
- ▶ 37:29 Shawn Wang How much of a moat is there in this like proprietary sort of training data that they've, uh, presumably accomplished because like even deep seek, it was able to do it and they had, you know, two months notice to do this, to do R one.
- ▶ 1:20:43 Shawn Wang Flash thinking is not on here, as well as all the other QWQs, R-ones, and, and all the other sort of thinking models.
- ▶ 1:36:51 Alessio Fanelli R-one, 2 times in the scene
Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
- ▶ 1:15 Luca Soldani Um, you have, uh, models that are, you know, reveling frontier level performance of what you can get from closed models, from like Quen, from Deep Seek.
- ▶ 11:30 Luca Soldani And, like, if you want real estate of the art, you know, your DeepSeq minimum is like 50,000.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 1:49:04 unnamed speaker Um, uh, so like, you know, it's, it, I don't know if you have any commentary on, like, uh, Mixtral, DeepSeq, Snowflake, Quen, uh, all these, um, proliferation of, uh, MOEs, MOE models that seem to all be sparse upcycle, because,
State of the Art: Training 70B LLMs on 10,000 H100 clusters
- ▶ 54:39 Jonathan Frankle The deep seek folks have done an awesome job looking at like batch size warmup.
How to train a Million Context LLM — with Mark Huang of Gradient.ai
- ▶ 1:08:55 Mark Huang Underrated, uh, uh, uh, specific instance would be, like, the deep seek paper.
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 52:57 unnamed speaker Uh, so anyway, I think moving on to next week's paper, uh, I was thinking of doing a deep seat MOE paper.
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 1:01:43 unnamed speaker Um, going from, like, mixed trial being eight experts, um, to, like, the DeepSeq MOE models, I don't know if you saw them, being, like, 30, 60 experts, and you can see it keep going up, I guess.
← previous page 2 of 2 · 100 scenes per page · newest episode first