DeepSeek, every mention

167 scenes, the whole family · ← back to DeepSeek

tap a year for its mentions
001252525050202420252026episodesmentions
02550202420252026episodes it came up in
002.525550202420252026episodesmentions per episode

every year anyone Shawn Wang 35Alessio Fanelli 19Yining Zhang 18Nathan Lambert 11Ahmad Awais 10Kyle Kranen 7Elie Bakouch 7William Beauchamp 6Rishabh Agarwal 6Philip Kiely 6

Verbatim, from the transcripts: the passages where DeepSeek comes up

loading…

Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin) Apr 21, 2025 · 2 mentions

  • ▶ 26:27 Charles Packer It's like pretty consistent across like both 3.7 deep seek, three mini, which all like the way you actually scale the x-axis here is fundamentally quite different in each case with 3.7 extended thinking mode.
  • ▶ 26:45 Charles Packer And then to, I guess, enable this sort of like to send, to tell like R one on the API, whether or not to like go to 10 K tokens versus like two K tokens, um, also is like a, you know, different mechanism.

The #1 SWE-Bench Verified Agent Apr 2, 2025 · 1 mention

  • ▶ 27:24 Guy Gur-Ari So I think the DeepSeq paper R one was very good understanding DPO and the variants like GRPO is very popular now.

The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind Mar 23, 2025 · 6 mentions

  • ▶ 13:45 Rishabh Agarwal And I guess a better version of this is you can do something like best defend, which is, let's say this was something which was used in DeepSeq. 3 times in the scene
  • ▶ 13:50 Rishabh Agarwal So, uh, if you saw that paper, they trained the DeepSeq reasoner model, and then they distilled the, the, basically that model to other models. 2 times in the scene
  • ▶ 43:26 Rishabh Agarwal It's annoying to use, like, let's say if you're trying to distill from DeepSeq, you can't serve that huge model and have, like, a live teacher running because it is going to be expensive.

Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin] Mar 7, 2025 · 2 mentions

  • ▶ 13:37 Misha Laskin You know, I, I think, uh, I think there are going to be a lot of move 37, uh, and, you know, one of them, so that, right, uh, for example, we're building, right, we're training these, uh, large language models, and maybe, maybe kind of a… 2 times in the scene

Browserbase: Browser Infrastructure For Your AI Agents Feb 28, 2025 · 1 mention

  • ▶ 1:19 Paul Klein I took my first vacation, actually, two weeks ago, and Operator came out on the first day, and then a week later, Deep Seek came out.

Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era Feb 28, 2025 · 3 mentions

  • ▶ 15:05 unnamed speaker Oh, any takeaways from like deep seek? 3 times in the scene

S1: the $6 DeepSeek R1 Competitor (ft. Entropix) Feb 26, 2025 · 15 mentions

  • ▶ 0:12 unnamed speaker Uh, DeepSeek is all the rage. 3 times in the scene
  • ▶ 0:25 unnamed speaker Uh, and we wind up looking for a guest and Tim has been blogging about R one, S one, um, and, and all the other stuff. 3 times in the scene
  • ▶ 2:00 unnamed speaker Like we, we've debated having a deep sea blog, unless you actually even wrote one, but I was like, I don't think we're saying anything new here.
  • ▶ 2:34 unnamed speaker Um, R one, uh, I was posting on it about, about it on blue sky
  • ▶ 2:56 unnamed speaker Um, should we double click more on R-One, or are we, should we go into S-One? 2 times in the scene
  • ▶ 3:06 unnamed speaker Um, so S-One came out more recently, and, um, we cloned Deep Sea for six dollars, is, is, uh, what the, the headline is? 3 times in the scene
  • ▶ 13:07 unnamed speaker I mean, R one, R one did both are there. 2 times in the scene

Bee AI: The Wearable Ambient Agent Feb 17, 2025 · 3 mentions

  • ▶ 55:07 Ethan Sutin You do need a sonnet level model to, like, execute things like the agent, and, um, we will be having a subscription for, like, features like that, because it's, you know, although now with the R-one, like, we'll see, uh, we haven't… 2 times in the scene
  • ▶ 55:22 Shawn Wang A deep seek?

smol agents are all you need Feb 13, 2025 · 5 mentions

The AI Architect: Bret Taylor Feb 11, 2025 · 1 mention

  • ▶ 1:33:29 Bret Taylor Now, uh, I just saw like someone post one of, they distilled one of the deep seek models and just made it really small and, you know, it's doing these chains of thoughts so fast, you know, it's achieving latency numbers, I think.

Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire Feb 6, 2025 · 3 mentions

  • ▶ 29:35 Shawn Wang So for example, if I wanted to use Deep Seek, I'm out of luck because Pydance API doesn't have Deep Seek yet. 3 times in the scene

Why every AI Engineer needs an AI Gateway (ft Portkey.ai CEO) Feb 5, 2025 · 4 mentions

  • ▶ 2:59 Rohit Agarwal I think DeepSeq was the first model where within the 1:06 hours of it launching, we had so many pull requests on our 2 times in the scene
  • ▶ 4:07 unnamed speaker And then they add Grok, they add Gemini, all these other things, they add DeepSeq, and then it starts to be a whole thing inside of the code base.
  • ▶ 19:16 Rohit Agarwal So you could always say, okay, if it comes to node A, always hit GPT-FORO, and node B is DeepSeek R-ONE,

Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis Jan 28, 2025 · 2 mentions

  • ▶ 15:52 unnamed speaker Well, uh, the, uh, the, the sort of apparent metagame from Paul Gauffier on Ader is that you use R-One as an architect and Sonnet as a code model, and that's apparently the best combination that beats O-One for him.
  • ▶ 34:15 Shawn Lewis The next question is, does our one do as well?

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research Jan 26, 2025 · 7 mentions

  • ▶ 22:42 William Beauchamp And it's similar, I guess, to how DeepSeek have been able to produce such a compelling model when compared to someone like an open AI, right? 6 times in the scene
  • ▶ 22:54 unnamed speaker V-three, yeah.

The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1 Jan 24, 2025 · 29 mentions

  • ▶ 1:24 Shawn Wang Yeah, so, okay, we're meeting to talk about what happened after DeepSeq's launch. 2 times in the scene
  • ▶ 2:07 unnamed speaker So on Monday, when Monday morning, DeepSeek one was released, we were like thinking, okay, maybe we can use Curator. 4 times in the scene
  • ▶ 7:10 unnamed speaker So with R one, it's the first time we actually get the reasoning traces in, in open source. 2 times in the scene
  • ▶ 10:28 Shawn Wang Mostly I agree on, on this discussion, especially from the R-one paper. 2 times in the scene
  • ▶ 11:58 Shawn Wang It seems like at least R-one is generalizing very well. 4 times in the scene
  • ▶ 12:25 unnamed speaker So if you look at how some of DeepSeq's distilled models are trained, they have. 4 times in the scene
  • ▶ 15:21 unnamed speaker The, the, the, the simple thing that we changed is changing out QWQ for DeepSeq R-One. 7 times in the scene
  • ▶ 20:10 unnamed speaker Like, DeepSeek, they, they in fact said that at least for one of their R-one-zero models, like, they didn't even do SFT on it. 3 times in the scene
  • ▶ 23:18 unnamed speaker I mean, it's amazing to see how since this R-one model came out every three hours or something.

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 35 mentions

  • ▶ 0:43 unnamed speaker And, you know, you are lead software engineer on the model performance team, and you guys recently shipped DeepSeq v. three as one of the many models that you do host. 2 times in the scene
  • ▶ 1:03 unnamed speaker So we can take this a number of directions, but I think one thing we wanted to just get off the bat on was to start with, uh, DeepSeq and, and maybe, and then we'll work our way backwards back to SGLang. 3 times in the scene
  • ▶ 1:22 Yining Zhang Yeah, because, uh, DeepSeq VIII, I think, is currently considered the leading open source LLMs based on the benchmark results and the chat area results. 6 times in the scene
  • ▶ 1:49 unnamed speaker One of the interesting things is, like, they are bootstrapped, like, uh, you know, very private, small lab. 2 times in the scene
  • ▶ 4:56 Yining Zhang So before the DeepSig-Fee-Seventy-B, I think we haven't encountered that issue for that so large weights. 4 times in the scene
  • ▶ 8:05 unnamed speaker It seems like, you know, I think Noam Shazir started talking about, um, sort of training natively quantize, and I think that's what DeepSea seems to have done, at least they said in their paper. 2 times in the scene
  • ▶ 13:05 unnamed speaker Uh, speak too much about, do you see, because obviously we don't know, uh, that much on this, unless we work on the team, but 2 times in the scene
  • ▶ 15:41 unnamed speaker Well, then one more thing, I guess, maybe more commercially relevant, um, D-Six API pricing is very competitive.
  • ▶ 18:28 unnamed speaker Can you maybe quickly run people through how do you go from taking the deep seek V three weights to like actually run it? 2 times in the scene
  • ▶ 26:57 Yining Zhang I think for the common use case, maybe not, not the DeepSeq VIII, for the common use case, I think SGLAN's performance is better than FLM, and its usability is better than TensorFlow TLM. 2 times in the scene
  • ▶ 27:40 Yining Zhang And also in DeepSeq, oh, sorry, in the SGLAN version 0.4, we also support the DP attention for DeepSeq, and in the latest 3 times in the scene
  • ▶ 34:25 Yining Zhang Yeah, we, we support some, uh, DeepSeq optimization, such as MLA optimization, DPR tension optimization, and we also support the serial overhead CPU schedule.
  • ▶ 35:50 unnamed speaker And when you think about a model that is, you know, as large as DeepSeq v three, especially, uh, like having better KVcache reutilization is great.
  • ▶ 50:03 Yining Zhang When we released the DeepSeq feed story support, we have some community user, something like Cursor. 2 times in the scene
  • ▶ 55:36 unnamed speaker I think this is a really good dive into both BaseNet and SG Lang, and a little bit of DeepSig V three, which people are very interested in.
  • ▶ 55:43 unnamed speaker I'm trying to talk to them as well, because obviously they're, uh, they're, they're a fascinating lab.

The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024] Jan 2, 2025 · 2 mentions

  • ▶ 6:21 Nathan Lambert The two I've highlighted are from Deep Seek and Quinn, and a lot of people in this room have probably seen them. 2 times in the scene

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 5 mentions

  • ▶ 37:29 Shawn Wang How much of a moat is there in this like proprietary sort of training data that they've, uh, presumably accomplished because like even deep seek, it was able to do it and they had, you know, two months notice to do this, to do R one.
  • ▶ 37:29 Shawn Wang How much of a moat is there in this like proprietary sort of training data that they've, uh, presumably accomplished because like even deep seek, it was able to do it and they had, you know, two months notice to do this, to do R one.
  • ▶ 1:20:43 Shawn Wang Flash thinking is not on here, as well as all the other QWQs, R-ones, and, and all the other sort of thinking models.
  • ▶ 1:36:51 Alessio Fanelli R-one, 2 times in the scene

Best of 2024: Open Models [LS LIVE! at NeurIPS 2024] Dec 23, 2024 · 2 mentions

  • ▶ 1:15 Luca Soldani Um, you have, uh, models that are, you know, reveling frontier level performance of what you can get from closed models, from like Quen, from Deep Seek.
  • ▶ 11:30 Luca Soldani And, like, if you want real estate of the art, you know, your DeepSeq minimum is like 50,000.

The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka Jul 5, 2024 · 1 mention

  • ▶ 1:49:04 unnamed speaker Um, uh, so like, you know, it's, it, I don't know if you have any commentary on, like, uh, Mixtral, DeepSeq, Snowflake, Quen, uh, all these, um, proliferation of, uh, MOEs, MOE models that seem to all be sparse upcycle, because,

State of the Art: Training 70B LLMs on 10,000 H100 clusters Jun 25, 2024 · 1 mention

How to train a Million Context LLM — with Mark Huang of Gradient.ai May 31, 2024 · 1 mention

A Comprehensive Overview of Large Language Models - Latent Space Paper Club Mar 15, 2024 · 1 mention

  • ▶ 52:57 unnamed speaker Uh, so anyway, I think moving on to next week's paper, uh, I was thinking of doing a deep seat MOE paper.

A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate Feb 28, 2024 · 1 mention

  • ▶ 1:01:43 unnamed speaker Um, going from, like, mixed trial being eight experts, um, to, like, the DeepSeq MOE models, I don't know if you saw them, being, like, 30, 60 experts, and you can see it keep going up, I guess.
← previous page 2 of 2 · 100 scenes per page · newest episode first
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.