DeepSeek-R1

also referred to as: deepseek r1

part of DeepSeek

7 statements across 5 episodes · 2 bullish · 1 bearish · 5 people on the record · first statement Jan 24, 2025 by Shawn Wang · said 74 times in 25 episodes since 2025 · across every show →

Mentions by year

brought up most by Shawn Wang (11), Will Brown (5), Nathan Lambert (4), Swyx (Marcos Swix) (2), Stephanie Palazzolo (2), Rishabh Agarwal (2), Marc Andreessen (2), Ethan Sutin (2)

tap a year for its mentions
004013802520252026episodesmentions
0132520252026episodes it came up in
0021342520252026episodesmentions per episode
2026 5 mentions in 4 episodes 1 per episode
2025 69 mentions in 21 episodes 3 per episode

every mention, scene by scene, with the transcript →

Everything said about DeepSeek-R1, oldest first

Jan 24, 2025 neutral
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Shawn Wang Jan 24, 2025 ▶ 10:32 The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
Feb 13, 2025 negative
Insight
Roucher: DeepSeek-R1 ranks slightly below OpenAI o1 on smolagents tasks
“I tried R one, but R one is a bit under O one with small agents. And I think this is also a matter of formatting. Like sometimes the model struggles to just output them, the code snippets in the correct way that we expect.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 14:25 smol agents are all you need
Apr 21, 2025 bullish
Assertion Supported
Packer: Sleep-time compute offers Pareto improvements across Claude 3.7 and DeepSeek
“It's like pretty consistent across like both 3.7 deep seek, three mini, which all like the way you actually scale the x-axis here is fundamentally quite different in each case with 3.7 extended thinking mode. The parameter you provide to scale it is different …”
Charles Packer Apr 21, 2025 ▶ 26:27 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
Jul 31, 2025 positive
Insight
Lambert: Reasoning models solved basic skills; planning is the next frontier
“So I came up with four and the foundational one was skills, which is What I would say that we have already done with O-one and R-one, which is you do a lot of RL, you show the inference time scaling works and you get really high benchmark numbers. And then the…”
Nathan Lambert Jul 31, 2025 ▶ 38:34 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 neutral
Assertion Supported
Lambert: DeepSeek-R1 starts solving math questions immediately without explicit planning
“If you look at DeepSeq R-One and you ask it a hard math question, it's not like, here's my plan of attack. It just starts.”
Nathan Lambert Jul 31, 2025 ▶ 40:10 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 neutral
Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Nathan Lambert Jul 31, 2025 ▶ 22:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Feb 5, 2026 neutral
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Mark Bissell Feb 5, 2026 ▶ 10:08 Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.