The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 23 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Awais: Claude Code hides 50+ tool-call failures per session on DeepSeek
“In CloudCode you know, they hide a lot of the errors behind control O, right? So you don't even know that, you know, you have like 50 plus tool call failures plus per session. You're just sitting there and you're like, oh, why is DeepSeq so slow?”
Ahmad Awais Jun 6, 2026 ▶ 12:21 ⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
Insight
McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know…”
Josh McGrath Dec 31, 2025 ▶ 12:42 [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
Assertion Not checkable as stated
Nair: OpenAI already possessed a superior model during the DeepSeek release
“The feeling in OpenAI is that like, well, I think we had a better model already at the time, right?”
Ashvin Nair Dec 30, 2025 ▶ 31:21 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Assertion Not checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Elie Bakouch Oct 20, 2025 ▶ 59:13 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Opinion
DeepSeek's open releases pulled global AI progress forward by six months
“But I think that what it did is it pulled forward progress in AI by like six months.”
Chris Lattner Jun 13, 2025 ▶ 57:54 The Shape of Compute (Chris Lattner of Modular)
Assertion Contradicted
Hong: DeepSeek dissolved its formal reasoning team over strategic shift
“And we have since, for example, Deep Seek All right. Like originally having a formal team and then later dissolve that team because of strategic direction change.”
Carina Hong Jun 3, 2026 ▶ 1:14:16 Scaling Past Informal AI - Carina Hong, Axiom Math
Insight
Morris: Weight deltas can reconstruct a competitor's proprietary fine-tuning dataset
“There's some tricks to it, but it's basically just like gradient based selection based on this weight difference. And it seems to be okay. Like it can get us pretty good training data. So I guess if you actually wanted to use this, it would be like your compet…”
Jack Morris Jul 2, 2025 ▶ 1:04:42 Information Theory for Language Models: Jack Morris
Assertion Supported
Beauchamp: DeepSeek slashes inference costs by shrinking KV cache
“What DeepSeek have achieved that's quite special is they've got this amazing inference engine. They've been able to reduce the size of the KV cache significantly. And then by being able to do that, they're able to significantly reduce their inference costs.”
William Beauchamp Jan 26, 2025 ▶ 23:30 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Assertion Not checkable as stated
Zhang: Baidu and ByteDance Internal Models Use DeepSeek-Like MoE Architectures
“As far as I know, some companies such as Baidu or Baidu Dance, they are internal, the dominant AOM, they use the MOE architecture, and their ways, I think, is similar to the DeepSeq MOE model.”
Yining Zhang Jan 19, 2025 ▶ 13:41 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Opinion
Huang: DeepSeek's Multi-Head Latent Attention is novel and underrated
“Underrated specific instance would be, like, the deep seek paper. I'd never seen it before, but, like, the multi-head latent attention, like, that was really unexpected to me because, like, I thought I'd seen every Not every type, obviously, but, like, every w…”
Mark Huang May 31, 2024 ▶ 1:08:55 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Assertion Supported
Mechanistic Interpretability Scales Without Bottlenecks to Large Models Like DeepSeek
“There's no gap for scale. Like, they've shown that even for the biggest open source models, you like, even like DeepSeq's big models, you, they can do it. And then in general, like, scaling is not the bottleneck.”
Deedy Das Nov 14, 2025 ▶ 52:11 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Assertion Supported
Agarwal: DeepSeek Distilled Reasoning Models Using Correctness-Filtered Synthetic Data
“What they did was they took the best model they had, they generated a bunch of samples, they filtered them based on correctness, because these were on tasks like coding, mathematical problem solving, and a bunch of those things. So they saw whatever, what are …”
Rishabh Agarwal Mar 23, 2025 ▶ 13:57 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Insight
Agarwal: Synthetic data distillation bypasses model vocabulary and tokenizer mismatches
“The one nice thing about this kind of distillation is it doesn't matter if you have a vocabulary mismatch, because we're not using the next token distribution or probability labels. You can distill from one model which uses some random tokenizer to another mod…”
Rishabh Agarwal Mar 23, 2025 ▶ 14:26 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Assertion Not checkable as stated
DeepSeek Revealed Attention Kernel Optimizations Kept Secret by Big Labs
“Deep seeks recent open sourcing of their various code components that they use to train that model, which I think outside of the big labs was not really well known to, right, it was not really well known how to write kind of a kernel that's optimized for this …”
Misha Laskin Mar 7, 2025 ▶ 13:37 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Assertion Supported
Colvin: AI providers are centralizing around OpenAI's API standard
“I think the truth is that everyone is centralizing around OpenAI's SD API as the one to do. So DeepSeek support that. Grok with a K support that. Olama also does it. Well, I mean, if there is that library right now, it's more or less the OpenAI SDK.”
Samuel Colvin Feb 6, 2025 ▶ 30:32 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Assertion Supported
Zhang: DeepSeek team officially recommends SGLang as its inference engine
“That's why SGLAN is the recommended LLM engine by the DeepSeq team.”
Yining Zhang Jan 19, 2025 ▶ 28:06 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Nathan Lambert Jan 2, 2025 ▶ 6:26 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Assertion Supported
Soldani: 2024 open models rival frontier performance of closed models
“You have models that are, you know, reveling frontier level performance of what you can get from closed models, from like Quen, from Deep Seek. We got Lama III.”
Luca Soldani Dec 23, 2024 ▶ 1:15 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Assertion Not checkable as stated
Awais: Claude 3.7 Max is already CommandCode's second most used model
“But they will only be for deep seek when to 3.7 max is the second most used model on command code right now. It's just two or three days old.”
Ahmad Awais Jun 6, 2026 ▶ 39:43 ⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
Disclosure
CommandCode offers 600 million DeepSeek tokens for $1 per month
“We launched a Go plan with just dollar one per month to, where you can do like, six hundred million tokens of DeepSQL for Pro in it, just to prove like, open models are actually really, really good, and they are catching up, right?”
Ahmad Awais Jun 6, 2026 ▶ 16:44 ⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
Assertion Supported
DeepSeek 128k context fits in 8GB KV cache versus Llama's 80GB
“For context like the total I think the total context length of DeepSeq is a 128,000 tokens, or it might be 256,000 with rope extension. That entire context, I think it's a 128,000, fits into eight gigabytes. And previously context, like I think the Lama four o…”
Kyle Kranen Mar 8, 2026 ▶ 52:08 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Assertion Supported
AMD funded community DeepSeek kernel development through the GPU Mode Discord
“AMD's actually done this. Like there's some deep seek, like through GPU modes, discord, there have been some like deep seek kernels that they say you have, you know, guaranteed access to compute for, write them.”
Quentin Anthony Nov 3, 2025 ▶ 57:48 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Assertion Supported
Lambert: DeepSeek-R1 starts solving math questions immediately without explicit planning
“If you look at DeepSeq R-One and you ask it a hard math question, it's not like, here's my plan of attack. It just starts.”
Nathan Lambert Jul 31, 2025 ▶ 40:10 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.