The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

why aren't all 5,957 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 100 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Webster: AI guardrails are a commodity and easy to build
“Building a business on guardrails really scares me because there are so many incumbents that can come in and eat your lunch. And it's like, It's not actually that hard to build a guardrail. I don't know if I'm gonna make people angry by saying that maybe some …”
Ian Webster Oct 24, 2025 ▶ 32:20 Breaking AI to Fix It: Ian Webster's Journey from Discord's Clyde to Promptfoo's $18M Series A
Opinion
Corbitt: GRPO is likely a dead end due to parallel rollout constraints
“The big downside, the huge downside of GRPO, and I think actually the reason why GRPO actually is likely to be a dead end, and we probably will not be continue using it indefinitely. The fact that you need to have these parallel rollouts in order to train on i…”
Kyle Corbitt Oct 16, 2025 ▶ 22:46 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Assertion Not checkable as stated
Corbitt: Prompt optimization methods like JEPA failed OpenPipe's agent benchmarks
“It didn't work on the problems we tried it on. It just didn't. It got like a minor boost over the sort of like more naive prompt we had and was just like, it was like, okay, Just kind of like our naive prompt with our model gets maybe like 50% on this benchmar…”
Kyle Corbitt Oct 16, 2025 ▶ 37:08 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year. And my money is on, they have to do a coin. Like it's, I'm not a crypto guy at all, but like, y…”
Shawn Wang Oct 16, 2025 ▶ 49:13 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Opinion
Lenz: Model providers should not dictate enterprise AI policies
“Right now, if you're using a model, you're taking in their own policy. Even if I want to use GPT-OSS, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave, etc. And I don't thi…”
Barak Lenz Oct 11, 2025 ▶ 40:33 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Opinion
Agarwal: No AI incident troubleshooting competitor genuinely works in production
“And I don't think we've seen any other company in our space having something actually work at production.”
Anish Agarwal Oct 5, 2025 ▶ 25:49 ⚡️Traversal: Causal ML and Reinforcement Learning
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Andrew Feldman Oct 1, 2025 ▶ 3:05 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Disclosure
Ball: Amp core team has 8 people, ships 15 times daily without code reviews
“I think we're around eight people now on the AMP core team, and we still don't do formal code reviews. We still push to main. We still ship 15 times every day.”
Thorsten Ball Sep 25, 2025 ▶ 10:55 Amp: The Emperor Has No Clothes
Prediction Not checkable as stated
Slack: AI tools like Cursor and Claude Code will peak and decline
“A lot of these other tools that are great, like cloud code and codex and cursor and so on that they've forgotten what made them great and what made them grow so fast, which is building the very best product. And they built it in a way that's too overfit on the…”
Quinn Slack Sep 25, 2025 ▶ 22:29 Amp: The Emperor Has No Clothes
Prediction Not checkable as stated
Slack: Top AI labs face a major customer stampede within two months
“I think we are one or two months away from a possible news cycle. That is the foundation model companies have spent billions of dollars in capex and hired like crazy. And now, you know, they're no longer the best in this realm and there's a huge stampede away …”
Quinn Slack Sep 25, 2025 ▶ 30:49 Amp: The Emperor Has No Clothes
Prediction Not checkable as stated
Ball: Complex sub-agent workflows will result in user hangovers
“A lot of the features what we see, you know, where people build like elaborate workflows, like I have my custom slash commands and they trigger custom Custom sub-agents and they in turn trigger custom MCP tool calls behind which again another model is doing in…”
Thorsten Ball Sep 25, 2025 ▶ 35:22 Amp: The Emperor Has No Clothes
Opinion
Slack: AI prompt enhancers are a 'bullshit feature' that does not work
“Yeah, so prompt enhancer, that's a bullshit feature that doesn't actually work. The theory behind it is nuts, because what helps LLMs is not tricks and phrasing your prompt in a certain way. It's fundamentally information that you have in your head that you ca…”
Quinn Slack Sep 25, 2025 ▶ 40:46 Amp: The Emperor Has No Clothes
Opinion
Slack: MCP as user-facing tech creates high token costs and bad UX
“As a user facing technology, it is such a common failure mode where a user will go and add in some MCP servers. Auth is a huge pain, but let's say they get over that hurdle. Then they have, I don't know, 50 tools exposed that often are too low level granularit…”
Quinn Slack Sep 25, 2025 ▶ 41:44 Amp: The Emperor Has No Clothes
Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer With a context length of, you know, 256,000 or more, they're not using a true transformer. What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Diego Bachman Sep 23, 2025 ▶ 2:58 ⚡️ Beyond Transformers with Power Retention
Opinion
Gorkem: Hosting LLMs is a bad business due to Google search competition
“Language models, hosting language models is not a good business. At the time we thought, okay, we are going to be competing against OpenAI and Anthropic and all these labs. Turned, turned out that it was even worse because the killer application of language mo…”
Gorkem Yurtseven Sep 8, 2025 ▶ 8:50 A Technical History of Generative Media
Opinion
Morcos: The Transformer is just one of many equivalently good architectures
“And one of my like more controversial viewpoints, I think, is that I think the transformer is a great advance to be sure, but I think it's one of a very large Set of equivalently good architectures that we could have found. And there are many, many ways we cou…”
Ari Morcos Aug 29, 2025 ▶ 12:35 Better Data is All You Need — Ari Morcos, Datology
Opinion
Morcos: GPT-4.5 and Llama 4 show limits of naive mega-model scaling
“And I think that's what we've seen to some extent with the failure of the mega models, right? With 4.5 and Lama four and others. I think that there is a challenge of just continuing to do that naively and you have to figure out how to break it.”
Ari Morcos Aug 29, 2025 ▶ 27:07 Better Data is All You Need — Ari Morcos, Datology
Insight
Morcos: Post-training alignment is ineffective long-term compared to pre-training alignment
“Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to tak…”
Ari Morcos Aug 29, 2025 ▶ 53:37 Better Data is All You Need — Ari Morcos, Datology
Opinion
Huber: Silicon Valley treats AGI as a secular religion
“I think AGI is also a religion. It has a problem of evil. We don't have enough intelligence. It has a solution, a deus ex machina. It has the second coming of Christ that AGI, the singularity is going to come. It's going to save humanity because we will now ha…”
Jeff Huber Aug 19, 2025 ▶ 48:45 Long Live Context Engineering - with Jeff Huber of Chroma
Opinion
Sohmers: Hardening silicon for specific AI models is obsolete in months
“Doing any of that, like, hardening for specific Model things. I don't think lasts more than, you know, two or three months at the rate that the industry moves at.”
Thomas Sohmers Aug 18, 2025 ▶ 30:11 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Opinion
Agrawal: Cerebras and Groq offer cloud APIs because their software struggles
“You know, if you look at Cerebrus and Grok and others, they've really tried to do this, their cloud kind of portal. And for us, you know, it shows two things. One is the difficulty of software for third party to implement that, that that's why they're kind of …”
Mitesh Agrawal Aug 18, 2025 ▶ 39:13 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Prediction Not checkable as stated
Ermon: Diffusion models could become the dominant architecture over autoregressive models
“I'm pretty optimistic about a future where diffusion models Can become the dominant solution. I've seen it happen before with GANs a few years ago, so I wouldn't be surprised if that's the case also here.”
Stefano Ermon Aug 4, 2025 ▶ 18:35 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Opinion
Mohan: Bearish on outsourced eval startups because AI companies must own evaluations
“And I guess maybe one of the things I'm a little bearish on is If another company comes out and solves eval properly for a bunch of different verticals, what was the company that they were selling to really doing? What are they really doing at that point? If t…”
Varun Mohan Jul 28, 2025 ▶ 41:05 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Opinion
Ramachandran: SWE-bench and HumanEval do not reflect real professional software engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have SweeBench, that's cool, no actual Professional work looks like Sweebench, like human eval, same thing.”
Anshul Ramachandran Jul 28, 2025 ▶ 1:50:27 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Prediction Not checkable as stated
Hou: 99% of AI editor rules file contents will be automatically inferred
“We strongly believe that having a rules file, you know, we do allow users to add a rules file, we strongly believe that a rules file is a crutch. You know, by the end of twenty-twenty-five, 99% of the things that you're gonna put in a rules file will be interp…”
Kevin Hou Jul 28, 2025 ▶ 2:57:45 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Prediction Open · timeframe Dec 2029
Kamradt Predicts ARC-AGI-3 Benchmark Will Remain Unbeaten For 3 Years
“And then V three, our durability estimate for that is three years. And that's what we're aiming for is 36 months for V three.”
Greg Kamradt Jul 18, 2025 ▶ 28:35 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Opinion
Kamradt: Any AGI Definition Involving Profit Has Ulterior Motives
“Any AGI definition that involves money has ulterior motives. I mean, simple as that. Money has nothing to do with intelligence, right?”
Greg Kamradt Jul 18, 2025 ▶ 31:25 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
What-if
Morris: ChatGPT could likely have been built using RNNs instead of Transformers
“And I think like, we honestly probably could have gotten this with RNNs. I know like the scaling laws paper shows that RNNs have worse curves for scaling, but probably people would have been like, I bet you could have built chat GPT with a very sophisticated R…”
Jack Morris Jul 2, 2025 ▶ 1:10:28 Information Theory for Language Models: Jack Morris
Opinion
Zach Lloyd: Warp's coding agent outperforms both Cursor and Windsurf
“Warp can code, and Warp code is at a level that's, like, comparable to cloud code. I think it's actually better than, like, Cursor's agent. It's probably better than Windsurf since they aren't, they don't have Sonnet for.”
Zach Lloyd Jun 25, 2025 ▶ 21:37 ⚡️Warp 2.0: the Agentic Development Environment - Zach Lloyd and Ben Holmes
Opinion
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Noam Brown Jun 19, 2025 ▶ 52:30 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Noam Brown Jun 19, 2025 ▶ 53:31 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Chris Lattner Jun 13, 2025 ▶ 12:25 The Shape of Compute (Chris Lattner of Modular)
Opinion
Lattner says C++ 'sucks' for modern accelerated compute
“Let me be the first to tell you, and I can say this now, I feel comfortable saying this, that C++ sucks.”
Chris Lattner Jun 13, 2025 ▶ 13:38 The Shape of Compute (Chris Lattner of Modular)
Opinion
Lattner calls vLLM a 'hot mess' due to too many stakeholders
“VLM seems much more like a massive community with a lot of stakeholders, a lot of stuff going on, and it's kind of a hot mess.”
Chris Lattner Jun 13, 2025 ▶ 23:52 The Shape of Compute (Chris Lattner of Modular)
Opinion
Ameisen: Stochastic Parrots Label Ignores Complex Multi-Step LLM Reasoning
“It's, like, activating many different distributed representations, like, combining them, and sort of, like, doing something pretty complicated. And so, yeah, I think it's funny, because in my opinion, that's like, yeah, like, oh god, stochastic parrots is not …”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:02:24 The Utility of Interpretability — Emmanuel Amiesen
Insight
Optimal future AI coding interfaces will not evolve from traditional IDEs
“Our take is that it is very unlikely that the optimal UI or the optimal interaction pattern for this new software development where humans spend much less time writing code. I think it's very unlikely that that optimal interaction pattern will be found by iter…”
Matan Grinberg May 29, 2025 ▶ 31:55 The AI Coding Factory
Insight
Embiricos: Scaffolding-heavy AI agents are limited by developers' mental capacity
“A lot of, like, agents that I see are really impressive, but it's basically, like, part of what's impressive is it's like a bunch of developers building this, like, really bespoke state machine around a bunch of, like, short model calls, and so then the upper …”
Alexander Embiricos May 16, 2025 ▶ 30:37 ChatGPT Codex: The Missing Manual
Opinion
Bergum: The standalone vector database infrastructure category is dying
“I'm not saying that the companies are dying, right? I'm just saying that the separate infrastructure category is dying, right? Because you have vector search capabilities in almost any DB technology nowadays, right?”
Jo Kristian Bergum Apr 19, 2025 ▶ 4:10 The Rise and Fall of the Vector DB category: Jo Kristian Bergum (ex-Chief Scientist, Vespa)
Prediction Not checkable as stated
Bergum: Pinecone will not endure like MongoDB because it is too narrow
“So there's always like this convergence, but MongoDB kind of, it sticks, but I don't think that for Pinecoin that was originally leading that movement, it won't like stick in the same way. It's too narrow. It's too, too narrow.”
Jo Kristian Bergum Apr 19, 2025 ▶ 7:56 The Rise and Fall of the Vector DB category: Jo Kristian Bergum (ex-Chief Scientist, Vespa)
Prediction Not checkable as stated
Conrad: Hyperscalers will probably lose significant money reselling Nvidia GPUs
“My intuition is that the hyperscalers are probably going to lose a lot of money, and they know they're going to lose a lot of money on reselling NVIDIA GPUs at least.”
Evan Conrad Apr 11, 2025 ▶ 6:28 SF Compute: Commoditizing Compute
Prediction Not checkable as stated
Conrad: Decentralized compute networks will never beat co-located InfiniBand clusters
“I just don't really think this is gonna ever be more efficient than a fully interconnected cluster with Infiniband, or, you know, whatever sort of next spec might be. Like, I could be completely wrong, but Speedolite is really hard to beat. And regardless of w…”
Evan Conrad Apr 11, 2025 ▶ 34:33 SF Compute: Commoditizing Compute
Prediction Not checkable as stated
Conrad: The AI VC bubble will pop and fail to return capital
“So what you've done by not having a future is you've inflated the venture capital market. And that is a bubble that's totally going to pop at some point. Like a lot of the companies are not going to work. And the valuations are not going to work. And what's go…”
Evan Conrad Apr 11, 2025 ▶ 59:31 SF Compute: Commoditizing Compute
Opinion
Colvin: Agent framework engineering quality lags far behind Python ecosystem
“Looking in general at the ecosystem of agent frameworks, the engineering quality is far below that of the rest of the Python ecosystem.”
Samuel Colvin Feb 6, 2025 ▶ 8:11 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Insight
Beauchamp: ML model performance plateaus near human-level on an S-curve
“With machine learning, first of all, you see that the performance of the models follows an S-curve. So it's not like it just goes off to infinity, right? And the S curve, it kind of plateaus around human level performance.”
William Beauchamp Jan 26, 2025 ▶ 9:39 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Prediction Not checkable as stated
Beauchamp: Timeline for AGI has certainly been pushed out
“I think the odds, the timeline for AGI has certainly been pushed out, right?”
William Beauchamp Jan 26, 2025 ▶ 52:02 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Assertion Not checkable as stated
Zhang: Meta Failed at Training MoE Models for Llama Series
“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
Yining Zhang Jan 19, 2025 ▶ 14:53 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Bryk: Perplexity and ChatGPT Search rely on legacy Google and Bing APIs
“So these systems, there are a few of them now they basically rely on like traditional search engines like Google or Bing, and then they combine them with like LLMs at the end to, you know, output some power graphics answering your question. So they, Like, Sear…”
Will Bryk Jan 10, 2025 ▶ 21:16 Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
Opinion
Comfyanonymous criticizes Gradio for coupling interface logic with backends
“Yeah, Gradio, I don't like Gradio. It's bad. Like, the, that's one of the reasons why, like, Automatic was very bad. It's great because the problem with Gradio, it forces you to, well, not forces you, but it kind of makes your interface logic and your back-end…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 28:56 AI Engineering for Art - with comfyanonymous
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.