Feb 2, 2025 · 12m · allin

AI Czar David Sacks Explains the DeepSeek Freak Out

David Sacks · 7m spoken Chamath Palihapitiya · 2m spoken David Friedberg · 48s spoken Jason Calacanis · 11s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the All-In Podcast, David Sacks and his co-hosts analyze the technical breakthroughs, financial realities, and geopolitical implications of Chinese AI firm DeepSeek's new reasoning model. They debunk viral cost myths while exploring how open-source innovation and extreme compute efficiency are shifting value across the global tech landscape.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 96.9% of the talking time here. How this is scored →

The hosts as informed peer 5.3 Guest teaching 2.8 Guest disagreement 1.8 The hosts pushing back 2.8
05100:0010:000:00–4:39 · The hosts as informed peer 0/10 The DeepSeek Reaction & Drivers of the Story David Sacks delivers an uninterrupted monologue synthesizing the global reaction to DeepSeek, breaking down open-source dynamics and comparing base LLMs to reasoning models. Because no hosts speak or interact during this opening section, all host-side scores remain zero.4:39–8:43 · The hosts as informed peer 5/10 Debunking the $6 Million Cost Myth & Real Compute Hardware Costs Hosts Jason Calacanis, TK, and Chamath Palihapitiya interject with questions regarding training costs, white paper validity, and GPU ownership. Sacks educates the group by citing semiconductor analyst estimates that DeepSeek utilizes a 50,000-hopper GPU cluster worth over a billion dollars.8:43–11:27 · The hosts as informed peer 9/10 DeepSeek's Algorithmic Innovations & Bypassing CUDA Chamath Palihapitiya takes the lead, reframing the narrative around algorithmic breakthroughs forced by compute constraints. He demonstrates high technical expertise by detailing DeepSeek's use of GRPO over standard PPO reinforcement learning and their bypass of Nvidia CUDA using low-level PTX code.11:27–12:27 · The hosts as informed peer 7/10 Market Impact, Value Chain Shift & The Electricity Analogy David Friedberg introduces an economic macro perspective, highlighting how model commoditization shifts value upstream or to downstream users, drawing an analogy to the early U.S. electricity market. The tone is entirely collaborative and exploratory.0:00–4:39 · Guest teaching 0/10 The DeepSeek Reaction & Drivers of the Story David Sacks delivers an uninterrupted monologue synthesizing the global reaction to DeepSeek, breaking down open-source dynamics and comparing base LLMs to reasoning models. Because no hosts speak or interact during this opening section, all host-side scores remain zero.4:39–8:43 · Guest teaching 6/10 Debunking the $6 Million Cost Myth & Real Compute Hardware Costs Hosts Jason Calacanis, TK, and Chamath Palihapitiya interject with questions regarding training costs, white paper validity, and GPU ownership. Sacks educates the group by citing semiconductor analyst estimates that DeepSeek utilizes a 50,000-hopper GPU cluster worth over a billion dollars.8:43–11:27 · Guest teaching 3/10 DeepSeek's Algorithmic Innovations & Bypassing CUDA Chamath Palihapitiya takes the lead, reframing the narrative around algorithmic breakthroughs forced by compute constraints. He demonstrates high technical expertise by detailing DeepSeek's use of GRPO over standard PPO reinforcement learning and their bypass of Nvidia CUDA using low-level PTX code.11:27–12:27 · Guest teaching 2/10 Market Impact, Value Chain Shift & The Electricity Analogy David Friedberg introduces an economic macro perspective, highlighting how model commoditization shifts value upstream or to downstream users, drawing an analogy to the early U.S. electricity market. The tone is entirely collaborative and exploratory.0:00–4:39 · Guest disagreement 1/10 The DeepSeek Reaction & Drivers of the Story David Sacks delivers an uninterrupted monologue synthesizing the global reaction to DeepSeek, breaking down open-source dynamics and comparing base LLMs to reasoning models. Because no hosts speak or interact during this opening section, all host-side scores remain zero.4:39–8:43 · Guest disagreement 2/10 Debunking the $6 Million Cost Myth & Real Compute Hardware Costs Hosts Jason Calacanis, TK, and Chamath Palihapitiya interject with questions regarding training costs, white paper validity, and GPU ownership. Sacks educates the group by citing semiconductor analyst estimates that DeepSeek utilizes a 50,000-hopper GPU cluster worth over a billion dollars.8:43–11:27 · Guest disagreement 3/10 DeepSeek's Algorithmic Innovations & Bypassing CUDA Chamath Palihapitiya takes the lead, reframing the narrative around algorithmic breakthroughs forced by compute constraints. He demonstrates high technical expertise by detailing DeepSeek's use of GRPO over standard PPO reinforcement learning and their bypass of Nvidia CUDA using low-level PTX code.11:27–12:27 · Guest disagreement 1/10 Market Impact, Value Chain Shift & The Electricity Analogy David Friedberg introduces an economic macro perspective, highlighting how model commoditization shifts value upstream or to downstream users, drawing an analogy to the early U.S. electricity market. The tone is entirely collaborative and exploratory.0:00–4:39 · The hosts pushing back 0/10 The DeepSeek Reaction & Drivers of the Story David Sacks delivers an uninterrupted monologue synthesizing the global reaction to DeepSeek, breaking down open-source dynamics and comparing base LLMs to reasoning models. Because no hosts speak or interact during this opening section, all host-side scores remain zero.4:39–8:43 · The hosts pushing back 4/10 Debunking the $6 Million Cost Myth & Real Compute Hardware Costs Hosts Jason Calacanis, TK, and Chamath Palihapitiya interject with questions regarding training costs, white paper validity, and GPU ownership. Sacks educates the group by citing semiconductor analyst estimates that DeepSeek utilizes a 50,000-hopper GPU cluster worth over a billion dollars.8:43–11:27 · The hosts pushing back 6/10 DeepSeek's Algorithmic Innovations & Bypassing CUDA Chamath Palihapitiya takes the lead, reframing the narrative around algorithmic breakthroughs forced by compute constraints. He demonstrates high technical expertise by detailing DeepSeek's use of GRPO over standard PPO reinforcement learning and their bypass of Nvidia CUDA using low-level PTX code.11:27–12:27 · The hosts pushing back 1/10 Market Impact, Value Chain Shift & The Electricity Analogy David Friedberg introduces an economic macro perspective, highlighting how model commoditization shifts value upstream or to downstream users, drawing an analogy to the early U.S. electricity market. The tone is entirely collaborative and exploratory.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 100% · guest 0%0:00 · the hosts 100% · guest 0%3:00 · the hosts 100% · guest 0%3:00 · the hosts 100% · guest 0%6:00 · the hosts 87.2% · guest 12.8%6:00 · the hosts 87.2% · guest 12.8%9:00 · the hosts 100% · guest 0%9:00 · the hosts 100% · guest 0%12:00 · the hosts 100% · guest 0%12:00 · the hosts 100% · guest 0%
Sharpest disagreement ▶ 8:43 Calling out analyst incentives and bias

Chamath forcefully dismisses external claims about training costs by arguing that semiconductor analysts and market participants are driven by inherent biases toward Nvidia.

Hardest push from the hosts ▶ 6:15 TK challenging Sacks on white paper stress testing

TK interrupts Sacks' cost debunking to point out that DeepSeek's published white paper allows developers to stress-test and verify their cost efficiency claims directly.

Biggest teaching moment ▶ 6:51 Sacks breaking down DeepSeek's true hardware cluster

Sacks educates the panel on the distinction between single training run costs and overall R&D infrastructure by detailing DeepSeek's 50,000 GPU cluster.

The host holds their own ▶ 9:40 Chamath detailing GRPO and PTX architectural innovations

Chamath demonstrates substantial technical expertise by breaking down how DeepSeek bypassed CUDA via PTX assembly programming and implemented GRPO reinforcement learning.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The DeepSeek Reaction & Drivers of the Story 0010 David Sacks delivers an uninterrupted monologue synthesizing the global reaction to DeepSeek, breaking down open-source dynamics and comparing base LLMs to reasoning models. Because no hosts speak or interact during this opening section, all host-side scores remain zero.
Debunking the $6 Million Cost Myth & Real Compute Hardware Costs 5624 Hosts Jason Calacanis, TK, and Chamath Palihapitiya interject with questions regarding training costs, white paper validity, and GPU ownership. Sacks educates the group by citing semiconductor analyst estimates that DeepSeek utilizes a 50,000-hopper GPU cluster worth over a billion dollars.
DeepSeek's Algorithmic Innovations & Bypassing CUDA 9336 Chamath Palihapitiya takes the lead, reframing the narrative around algorithmic breakthroughs forced by compute constraints. He demonstrates high technical expertise by detailing DeepSeek's use of GRPO over standard PPO reinforcement learning and their bypass of Nvidia CUDA using low-level PTX code.
Market Impact, Value Chain Shift & The Electricity Analogy 7211 David Friedberg introduces an economic macro perspective, highlighting how model commoditization shifts value upstream or to downstream users, drawing an analogy to the early U.S. electricity market. The tone is entirely collaborative and exploratory.

Statements from this episode (8)

Assertion Supported
Sacks: DeepSeek release erased $1 trillion in market cap in one day
“It's rare that a model release is going to be a global news story or Cause a trillion dollars of market cap decline in, in one day.”
David Sacks Feb 2, 2025 ▶ 0:26
Assertion Not checkable as stated
Sacks: China's AI lag behind US narrowed to 3 to 6 months
“If you had asked most people in the industry a few weeks ago, how far behind is China on AI models, they would say six to 12 months. And now I think they might say something more like three to six months, right?”
David Sacks Feb 2, 2025 ▶ 4:10
Assertion Not checkable as stated
Sacks: US AI models' final training runs cost tens of millions
“The final training run cost was more in the tens of millions of dollars about nine or 10 months ago. And so, you know, it's not six million versus a billion, okay?”
David Sacks Feb 2, 2025 ▶ 5:35
Assertion Supported
Sacks: DeepSeek's compute cluster costs over $1 billion
“You add up the cost of a compute cluster with 50,000 plus hoppers, and it's going to be over a billion dollars. So this idea that you've got this scrappy company that did it for only six million, just not true. They have a substantial compute cluster that they…”
David Sacks Feb 2, 2025 ▶ 7:58
Assertion Supported
Palihapitiya: DeepSeek created GRPO algorithm to slash AI memory requirements
“They invented a totally different algorithm. There was the orthodoxy. Right? This thing called PPO that everybody used, and they were like, no, we're going to use something else called, I think it's called GRPO or something. It uses a lot less computer memory,…”
Chamath Palihapitiya Feb 2, 2025 ▶ 9:55
Assertion Partly supported
Palihapitiya: DeepSeek bypassed Nvidia CUDA software lock-in using PTX
“Everybody is used to building models and compiling through CUDA, which is Nvidia's proprietary language, which I've said for a couple of times is their biggest moat, but it's also the biggest threat factor for lock-in. And these guys worked totally around CUDA…”
Chamath Palihapitiya Feb 2, 2025 ▶ 10:24
Insight
Palihapitiya: Smaller $2M seed rounds force AI startups to innovate
“When the AI company wakes up and rolls out of bed and some VC gives them two hundred million dollars, maybe that's not the right answer for a series A or a seed. And maybe the right answer is two million so that they do these deep seek like innovations.”
Chamath Palihapitiya Feb 2, 2025 ▶ 11:10
Prediction Not checkable as stated
Friedberg: Rapid AI model commoditization will shift value creation elsewhere
“At the end of the day, if model performance continues to improve, get cheaper, and it's so competitive that it commoditizes much faster than anyone even Thought. Then the value is going to be created somewhere else in the value chain.”
David Friedberg Feb 2, 2025 ▶ 11:58
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 460 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.