DeepSeek V3

product on 10 shows · 11 statements across 7 episodes · said 93 times in 29 episodes since 2025

Latent Space 27 the a16z Podcast 20 TBPN 17 the MAD Podcast 11 All-In 8 Big Technology 6 No Priors 1 the Startup Ideas Podcast 1 Catalyst 1 20VC 1

Mentions by year, every show

tap a year for its mentions
004013802520252026episodesmentions
0132520252026episodes it came up in
0021342520252026episodesmentions per episode

Latent Space 27the a16z Podcast 20TBPN 17the MAD Podcast 11All-In 8Big Technology 620VC 1Catalyst 12 more shows

2026 15 mentions in 5 episodes 3 per episode
2025 78 mentions in 24 episodes 3 per episode

every mention on every show, scene by scene, with the transcript →

11 statements about DeepSeek V3, every show

MAD Assertion Supported
DeepSeek-R1 cost roughly $300,000 to train, 10x cheaper than DeepSeek-V3
“Deep seek version three, they had like a five million dollar price tag on that given the, I think two dollars per GPU, they assume whether that's a correct assumption or not it's a different question, but if you compare it relative to the cost of R one, I thin…”
Sebastian Raschka Jan 29, 2026 ▶ 31:06 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
LATENT SPACE Assertion Supported
Bakouch: DeepSeek-V3 uses the same Adam optimizer parameters as Llama 2
“And for example, a good a good way to view that is that DeepSeq rig three is still using the same Adam parameter than Lama two.”
Elie Bakouch Oct 20, 2025 ▶ 8:54 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
BIG TECHNOLOGY Assertion Supported
Kantrowitz: Kimi K2 Scores 65.8 on SWE-Bench, Trailing Claude 4 Opus
“Claude IV Opus gets a 72.5 on that. Kimi K-II gets 65.8, so not far behind. And just to, you know, give some context, Deep Seek V-III, which everybody was going crazy over, gets a 38.”
Alex Kantrowitz Jul 21, 2025 ▶ 45:32 Grok's AI Lovebot, Aqui-Hire-Sition Backlash, OpenAI's ChatGPT Agent Debuts
BIG TECHNOLOGY Assertion Supported
Patel: Model inference costs dropped 60x from GPT-4 to DeepSeek-V3
“And likewise, when we look at from GPT-IV to DeepSeq VIII it's fallen roughly 600 X in cost. Right. So we're not quite at that 1200 X, but it has fallen 600 X in cost from 60 dollars to less than you know, to about a dollar. Right. Or to less than a dollar. So…”
Dylan Patel Apr 23, 2025 ▶ 29:14 Generative AI 101: Tokens, Pre-training, Fine-tuning, Reasoning — With SemiAnalysis CEO Dylan Patel
BIG TECHNOLOGY Prediction Held up
Patel: Meta's next Llama model will match DeepSeek-V3's cost efficiency
“And Meta's Meta is going to release their new llama soon enough. Right. And that one is going to be, you know, a similar level of cost decrease probably similar areas, deep seek V three.”
Dylan Patel Apr 23, 2025 ▶ 30:27 Generative AI 101: Tokens, Pre-training, Fine-tuning, Reasoning — With SemiAnalysis CEO Dylan Patel
a16z Assertion Supported
Mascorro: DeepSeek-V3 features 256 experts, far exceeding typical open-source models
“We talk about it as 256 experts, which is a large, a relative large number of experts in terms of at least open source models that we've seen out there.”
Marco Mascorro Mar 5, 2025 ▶ 9:00 DeepSeek, Reasoning Models, and the Future of LLMs
a16z Assertion Supported
Appenzeller: DeepSeek spent $5.5 million at market rates to train V3
“They quoted 5.5 million dollars, I think, you know, at market rates to train it.”
Guido Appenzeller Mar 5, 2025 ▶ 18:14 DeepSeek, Reasoning Models, and the Future of LLMs
ALL-IN Assertion Supported
David Sacks notes DeepSeek V3 self-identified as ChatGPT-4
“A month ago, we had a press cycle in Silicon Valley when DeepSeq's V-three model came out, that DeepSeq V-three was self-identifying as ChatGPT. When you would ask it, who are you? Like, what model are you? Five out of eight times, V-three would tell you that …”
David Sacks Jan 31, 2025 ▶ 34:56 DeepSeek Panic, US vs China, OpenAI $40B?, and Doge Delivers with Travis Kalanick and David Sacks
Zhang: DeepSeek-V3 is currently the leading open-source LLM
“Yeah, because DeepSeq VIII, I think, is currently considered the leading open source LLMs based on the benchmark results and the chat area results.”
Yining Zhang Jan 19, 2025 ▶ 1:22 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
LATENT SPACE Assertion Supported
Zhang: DeepSeek-V3 cannot run on a single 8xH100 GPU node
“You need, I think 671 gigabytes for the weights, and you also need an extra memory for the KV cache, so it's not possible to run that on H-one hundred.”
Yining Zhang Jan 19, 2025 ▶ 2:46 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
LATENT SPACE Assertion Not checkable as stated
Zhang: Cursor Team Contacted SGLang Over DeepSeek-V3 Support
“And when we released the DeepSeq feed story support some employee from the Cursor team Also very interested in our implementation and ask, reach out and ask some questions from us.”
Yining Zhang Jan 19, 2025 ▶ 50:25 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.