DeepSeek, every mention
46 scenes, the whole family · ← back to DeepSeek
tap a year for its mentions
every year anyone Sebastian Raschka 32Matt Turck 25Lin Qiao 13Jeremy Howard 3Douwe Kiela 3Benedict Evans 3Dan Fu 2Yann Dubois 1Nathan Benaich 1Mitch Trojanowski 1
Verbatim, from the transcripts: the passages where DeepSeek comes up
How to Build Autonomous, Long-Horizon AI Agents | Basis
- ▶ 20:03 Mitch Trojanowski Versus if you look, you know, if you fast forward a bit and you look at the, like the DeepSeq, um, R-one paper, uh, where they effectively laid out, you know, what I think all the labs were doing at that time, or at least OpenAI was doing…
Cerebras CEO: Why GPUs Can't Do Fast Inference
- ▶ 12:13 Matt Turck So this, oh, Huawei sent chips for deep seek and this emergence of a full stack Chinese, um, AI factory for like a better term.
OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
- ▶ 27:53 Dan Roberts If you look at, like, the DeepSeq algorithm,
OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
- ▶ 38:08 Yann Dubois From models like Kimi or, or from DeepSeq models, it seems that they are closer to one million data points.
State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
- ▶ 2:55 Sebastian Raschka Take a GPT one or two model and with a few, I mean, few lines of code almost, you can transform it into the latest, uh, let's say deep seek version, 3.2 architecture.
- ▶ 15:31 Sebastian Raschka I would say deep seek kind of like restarted that trend in 2024 in December with deep seek version three, they had an OE model before, but I think this is like the one that everyone looked at because that made such a big splash that people… 4 times in the scene
- ▶ 15:31 Sebastian Raschka I would say deep seek kind of like restarted that trend in 2024 in December with deep seek version three, they had an OE model before, but I think this is like the one that everyone looked at because that made such a big splash that people… 2 times in the scene
- ▶ 16:21 Sebastian Raschka Uh, DeepSig version 3.2, where they changed attention mechanism.
- ▶ 19:55 Sebastian Raschka So RLVR was kind of like popularized by deep seek R one, which was based on deep seek version three. 3 times in the scene
- ▶ 19:55 Sebastian Raschka So RLVR was kind of like popularized by deep seek R one, which was based on deep seek version three. 2 times in the scene
- ▶ 20:06 Sebastian Raschka And with that, they also introduced the GRPO algorithm you, you mentioned, but it, well, they go well together because they make it more efficient, the whole thing, but it doesn't have to be.
- ▶ 26:03 Sebastian Raschka And so my statement that it is not so promising or useful was mainly based on the R one paper where they had a final paragraph at the bottom.
- ▶ 31:06 Sebastian Raschka Because if you also look at the numbers of how much it costs, uh, just GPU hours deep seek version three, they had like a five million dollar price tag on that given the, I think two dollars per GPU, they assume whether that's a correct… 3 times in the scene
- ▶ 31:06 Sebastian Raschka Because if you also look at the numbers of how much it costs, uh, just GPU hours deep seek version three, they had like a five million dollar price tag on that given the, I think two dollars per GPU, they assume whether that's a correct… 2 times in the scene
- ▶ 31:06 Sebastian Raschka Because if you also look at the numbers of how much it costs, uh, just GPU hours deep seek version three, they had like a five million dollar price tag on that given the, I think two dollars per GPU, they assume whether that's a correct…
- ▶ 37:10 Sebastian Raschka I mean, if I look at DeepSeq, for example, because, I mean, I'm, I'm always picking here DeepSeq in this podcast because, uh, I think they have a really nice trajectory of models. 4 times in the scene
- ▶ 37:29 Sebastian Raschka I mean, if you look at a version three and then R one, and then they had version 3.2 model with the sparse attention mechanism, and then also this math version two with a self refinement and everything.
- ▶ 37:29 Sebastian Raschka I mean, if you look at a version three and then R one, and then they had version 3.2 model with the sparse attention mechanism, and then also this math version two with a self refinement and everything.
- ▶ 37:29 Sebastian Raschka I mean, if you look at a version three and then R one, and then they had version 3.2 model with the sparse attention mechanism, and then also this math version two with a self refinement and everything.
- ▶ 46:42 Sebastian Raschka And I think why I think that is if you use DeepSeq locally or use the platform, let's say you use a local LLM and use ChatGPT.
- ▶ 52:55 Sebastian Raschka Maybe, you know, with DeepSeq version three, not so much, but that's almost like a different community, like the Tinkerer community, like me, like small system, uh, 2 times in the scene
- ▶ 56:57 Sebastian Raschka Because I mean, back then it was GPT one, GPT two, GPT three, and now it's GPT four, 4.14 point, I think two or five, five, five, one, five, two, and all the models, they are iterating now, or even the same with deep seek.
The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
AI Eats the World: Benedict Evans on What Really Matters Now
- ▶ 10:58 Benedict Evans But for the moment, you know, and this is, was kind of my point about Deep Seek. 3 times in the scene
- ▶ 11:08 Matt Turck Deep Seek or Deep Research?
Jeremy Howard on Building 5,000 AI Products with 14 People (Answer AI Deep-Dive)
- ▶ 0:50 Matt Turck In this episode, we dig into the dialogue engineering workflow Jeremy's team uses to code and build alongside models like Quen and DeepSeq, and discuss Jeremy's no mercy verdict on Devon and the coding agent hype.
- ▶ 3:33 Jeremy Howard When and deep seek, I guess we've decided we like them very much.
- ▶ 6:29 Matt Turck A few months into that big, you know, deep seek moment that some call the spotting moment for AI and all the things. 6 times in the scene
- ▶ 10:38 Jeremy Howard And companies like DeepSeq showing we really do have lots and lots of opportunity to bring costs down.
- ▶ 33:42 Jeremy Howard Um, like it helps, like, understand the details of the technology well, so I kind of, these things like DeepSea Car One or whatever don't seem, they don't come out of the blue, they don't seem like wild jumps or whatever, you know, I can,…
Snowflake CEO on Winning the AI Arms Race
- ▶ 1:19:54 Matt Turck Deep Seek. 6 times in the scene
Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
- ▶ 42:33 Matt Turck This year, we're in a deep seek kind of a, you know, kind of, kind of moment. 3 times in the scene
Why This Ex-Meta Leader is Rethinking AI Infrastructure | Lin Qiao, CEO, Fireworks AI
- ▶ 0:04 Matt Turck Today, my guest is Lin Shao, the CEO of Fireworks AI, an LLM inference platform that enables companies like Cursor, Uber, and DoorDash to build AI product experiences on top of hundreds of open source models like DeepSeek, Quen, Lama, and… 2 times in the scene
- ▶ 35:57 Lin Qiao Uh, so human in the loop, uh, that we have, uh, take DeepSeq for example. 4 times in the scene
- ▶ 44:06 Lin Qiao And deep seek model come in as a raw model. 2 times in the scene
- ▶ 44:09 Lin Qiao Um, and we add function calling to v three model and which
- ▶ 55:43 Lin Qiao And it's clear that Deep Seek has, has created a big dent there, but the meaning is not about Deep Seek itself by itself. 6 times in the scene
Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
- ▶ 0:41 Matt Turck We started the conversation with Dao's thoughts on the latest AI model innovations, including GPT-Four .5, Sonnet-Four .7, and of course, DeepSeq.
- ▶ 2:57 Douwe Kiela For me, the most exciting thing by far is DeepSeek, where that really, I think, changed the narrative in the AI ecosystem around what's possible, and who actually is an incumbent, and what is the moat that some of these companies have.
- ▶ 7:23 Matt Turck So going back to, uh, DeepSeek that you mentioned, uh, a minute ago, what's sort of fascinating is that, uh, a couple of weeks ago, whenever DeepSeek came out, uh, it felt like a Sputnik moment. 2 times in the scene
- ▶ 8:51 Douwe Kiela It's, it's like when you have a bunch of GPUs and you know how to train a language model, uh, and, and if you can train up a pretty good base model, so that's their deep seek V three, then you can give that reasoning capabilities…
- ▶ 44:50 Douwe Kiela Yeah, I, I think synthetic data, um, I, so O-one, uh, and, and I guess sort of DeepSeq showed that synthetic data is actually pretty valuable, right?
From Selfie to Studio: Captions CEO on AI Video for 10M Creators
- ▶ 31:34 Gaurav Misra And Deep Seek and all that, right?
Farewell, Chatbots: AI Agents Are Taking Over Customer Service | Mike Murchison, CEO, Ada
- ▶ 9:32 Matt Turck Maybe in February or March you'll be running on top of a deep seek. 2 times in the scene
State of AI 2024: Frontier Models, AI Geopolitics, Robotics | Nathan Benaich, Air Street Capital
- ▶ 10:34 Nathan Benaich Um, like, their models like Alibaba's Quen, um, and then this spin out from a quantitative hedge fund, uh, DeepSeek, which publishes code models and others, and they've been actually a very, uh, lively contributor to open source, and on…