The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 790 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Agarwal: No AI incident troubleshooting competitor genuinely works in production
“And I don't think we've seen any other company in our space having something actually work at production.”
Anish Agarwal Oct 5, 2025 ▶ 25:49 ⚡️Traversal: Causal ML and Reinforcement Learning
Opinion
Slack: AI prompt enhancers are a 'bullshit feature' that does not work
“Yeah, so prompt enhancer, that's a bullshit feature that doesn't actually work. The theory behind it is nuts, because what helps LLMs is not tricks and phrasing your prompt in a certain way. It's fundamentally information that you have in your head that you ca…”
Quinn Slack Sep 25, 2025 ▶ 40:46 Amp: The Emperor Has No Clothes
Opinion
Slack: MCP as user-facing tech creates high token costs and bad UX
“As a user facing technology, it is such a common failure mode where a user will go and add in some MCP servers. Auth is a huge pain, but let's say they get over that hurdle. Then they have, I don't know, 50 tools exposed that often are too low level granularit…”
Quinn Slack Sep 25, 2025 ▶ 41:44 Amp: The Emperor Has No Clothes
Opinion
Gorkem: Hosting LLMs is a bad business due to Google search competition
“Language models, hosting language models is not a good business. At the time we thought, okay, we are going to be competing against OpenAI and Anthropic and all these labs. Turned, turned out that it was even worse because the killer application of language mo…”
Gorkem Yurtseven Sep 8, 2025 ▶ 8:50 A Technical History of Generative Media
Opinion
Morcos: The Transformer is just one of many equivalently good architectures
“And one of my like more controversial viewpoints, I think, is that I think the transformer is a great advance to be sure, but I think it's one of a very large Set of equivalently good architectures that we could have found. And there are many, many ways we cou…”
Ari Morcos Aug 29, 2025 ▶ 12:35 Better Data is All You Need — Ari Morcos, Datology
Opinion
Morcos: GPT-4.5 and Llama 4 show limits of naive mega-model scaling
“And I think that's what we've seen to some extent with the failure of the mega models, right? With 4.5 and Lama four and others. I think that there is a challenge of just continuing to do that naively and you have to figure out how to break it.”
Ari Morcos Aug 29, 2025 ▶ 27:07 Better Data is All You Need — Ari Morcos, Datology
Opinion
Huber: Silicon Valley treats AGI as a secular religion
“I think AGI is also a religion. It has a problem of evil. We don't have enough intelligence. It has a solution, a deus ex machina. It has the second coming of Christ that AGI, the singularity is going to come. It's going to save humanity because we will now ha…”
Jeff Huber Aug 19, 2025 ▶ 48:45 Long Live Context Engineering - with Jeff Huber of Chroma
Opinion
Sohmers: Hardening silicon for specific AI models is obsolete in months
“Doing any of that, like, hardening for specific Model things. I don't think lasts more than, you know, two or three months at the rate that the industry moves at.”
Thomas Sohmers Aug 18, 2025 ▶ 30:11 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Opinion
Agrawal: Cerebras and Groq offer cloud APIs because their software struggles
“You know, if you look at Cerebrus and Grok and others, they've really tried to do this, their cloud kind of portal. And for us, you know, it shows two things. One is the difficulty of software for third party to implement that, that that's why they're kind of …”
Mitesh Agrawal Aug 18, 2025 ▶ 39:13 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Opinion
Mohan: Bearish on outsourced eval startups because AI companies must own evaluations
“And I guess maybe one of the things I'm a little bearish on is If another company comes out and solves eval properly for a bunch of different verticals, what was the company that they were selling to really doing? What are they really doing at that point? If t…”
Varun Mohan Jul 28, 2025 ▶ 41:05 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Opinion
Ramachandran: SWE-bench and HumanEval do not reflect real professional software engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have SweeBench, that's cool, no actual Professional work looks like Sweebench, like human eval, same thing.”
Anshul Ramachandran Jul 28, 2025 ▶ 1:50:27 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Opinion
Kamradt: Any AGI Definition Involving Profit Has Ulterior Motives
“Any AGI definition that involves money has ulterior motives. I mean, simple as that. Money has nothing to do with intelligence, right?”
Greg Kamradt Jul 18, 2025 ▶ 31:25 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Opinion
Zach Lloyd: Warp's coding agent outperforms both Cursor and Windsurf
“Warp can code, and Warp code is at a level that's, like, comparable to cloud code. I think it's actually better than, like, Cursor's agent. It's probably better than Windsurf since they aren't, they don't have Sonnet for.”
Zach Lloyd Jun 25, 2025 ▶ 21:37 ⚡️Warp 2.0: the Agentic Development Environment - Zach Lloyd and Ben Holmes
Opinion
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Noam Brown Jun 19, 2025 ▶ 52:30 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Noam Brown Jun 19, 2025 ▶ 53:31 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Opinion
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Chris Lattner Jun 13, 2025 ▶ 12:25 The Shape of Compute (Chris Lattner of Modular)
Opinion
Lattner says C++ 'sucks' for modern accelerated compute
“Let me be the first to tell you, and I can say this now, I feel comfortable saying this, that C++ sucks.”
Chris Lattner Jun 13, 2025 ▶ 13:38 The Shape of Compute (Chris Lattner of Modular)
Opinion
Lattner calls vLLM a 'hot mess' due to too many stakeholders
“VLM seems much more like a massive community with a lot of stakeholders, a lot of stuff going on, and it's kind of a hot mess.”
Chris Lattner Jun 13, 2025 ▶ 23:52 The Shape of Compute (Chris Lattner of Modular)
Opinion
Ameisen: Stochastic Parrots Label Ignores Complex Multi-Step LLM Reasoning
“It's, like, activating many different distributed representations, like, combining them, and sort of, like, doing something pretty complicated. And so, yeah, I think it's funny, because in my opinion, that's like, yeah, like, oh god, stochastic parrots is not …”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:02:24 The Utility of Interpretability — Emmanuel Amiesen
Opinion
Bergum: The standalone vector database infrastructure category is dying
“I'm not saying that the companies are dying, right? I'm just saying that the separate infrastructure category is dying, right? Because you have vector search capabilities in almost any DB technology nowadays, right?”
Jo Kristian Bergum Apr 19, 2025 ▶ 4:10 The Rise and Fall of the Vector DB category: Jo Kristian Bergum (ex-Chief Scientist, Vespa)
Opinion
Colvin: Agent framework engineering quality lags far behind Python ecosystem
“Looking in general at the ecosystem of agent frameworks, the engineering quality is far below that of the rest of the Python ecosystem.”
Samuel Colvin Feb 6, 2025 ▶ 8:11 Agent Engineering with Pydantic + Graphs — with Samuel Colvin, CEO of Pydantic Logfire
Opinion
Comfyanonymous criticizes Gradio for coupling interface logic with backends
“Yeah, Gradio, I don't like Gradio. It's bad. Like, the, that's one of the reasons why, like, Automatic was very bad. It's great because the problem with Gradio, it forces you to, well, not forces you, but it kind of makes your interface logic and your back-end…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 28:56 AI Engineering for Art - with comfyanonymous
Opinion
Neubig: Anthropic's MCP Duplicates Existing APIs With Little Added Value
“We already have an API for GitHub. So why do we need an MCP for GitHub, right? You know, like GitHub has an API. The GitHub API is evolving. We can look up the GitHub API documentation. So it seems like kind of duplicated a little bit. And also they have a set…”
Graham Neubig Dec 25, 2024 ▶ 41:29 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Opinion
Ben Allal: Synthetic data may enrich the web rather than pollute it
“So personally, I wouldn't say the web is posted with synthetic data. Maybe it's even making it more rich.”
Loubna Ben Allal Dec 24, 2024 ▶ 4:35 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Opinion
Soldani: Open Model Bio-Risk Warnings Were a Lobbying Ploy
“You know, if you remember the beginning of this year, it was all about bio-risk of these open models. The whole thing fizzled out because there's been, finally there's been, like, rigorous research, not just this paper from coherent folks, but there's been rig…”
Luca Soldani Dec 23, 2024 ▶ 23:09 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Opinion
Ramachandran: SWE-bench and HumanEval do not reflect real professional engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have sweet bench. That's cool. No actual Professional work looks like Sweebench. Like, human eval, same thi…”
Anshul Ramachandran Dec 13, 2024 ▶ 13:17 Windsurf: The Enterprise AI IDE
Opinion
Crivello: OpenAI's GPT-4o Is Overhyped and Poor for Agents
“I think four O is overhyped. Frankly, we don't use four O. I don't think it's good for agentic behavior.”
Florent Crivello Nov 15, 2024 ▶ 32:07 Agents @ Work: Lindy.ai (with live demo!)
Opinion
Polu: The bulk of useful enterprise agent work can use APIs
“The bulk of the useful stuff that you can do within the company can be done through API. The data can be retrieved by API, the actions can be taken through API.”
Stanislas Polu Nov 11, 2024 ▶ 32:30 Agents @ Work: Dust.tt — with Stanislas Polu
Opinion
Drew Houston: Large language models are a rapidly self-commoditizing, bad business
“Large language models are a pretty bad Business from a, you know, you sort of take off your tech lens and just sort of business lens. Like there's sort of this weirdly self-commoditizing thing where, you know, models only have value if they're kind of on this…”
Drew Houston Oct 18, 2024 ▶ 29:50 Building the Silicon Brain - Drew Houston of Dropbox
Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Ankur Goyal Oct 11, 2024 ▶ 1:27:25 Production AI Engineering starts with Evals
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:24 Production AI Engineering starts with Evals
Opinion
Schulhoff: Role Prompting Does Not Improve Accuracy on Modern LLMs
“For accuracy-based tasks, like MMLU, you're trying to solve a math problem, and maybe you tell the AI that it's a math professor, and you expect it to have improved performance. I really don't think that works. I'm quite certain that doesn't work on more moder…”
Sander Schulhoff Sep 20, 2024 ▶ 17:08 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Opinion
Tay: Long context architecture is the future of AI over RAG
“And, yeah, I mean, I think long context is definitely the future, rather than rec. But I mean, they could be used in conjunction, like,”
Yi Tay Jul 5, 2024 ▶ 1:40:05 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Opinion
Big AI labs build AGI internally while giving the public child-proof apps
“In some sense, the big AI companies are incentivized and interested in building AGI internally, and giving everybody else a child-proof application.”
Joscha Bach Apr 27, 2024 ▶ 1:27:38 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Opinion
Bach: Prompting an LLM correctly can make it sentient to some degree
“It's a Weltgeist that gets possessed by a prompt. And if you possess it with the right prompt, then it can become sentient to some degree.”
Joscha Bach Apr 27, 2024 ▶ 1:44:52 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Opinion
Liu: App-layer AI startups should avoid hiring traditional machine learning engineers
“I think a lot of these app layer startups should not be hiring MLEs because they end up churning.”
Jason Liu Apr 24, 2024 ▶ 58:42 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Opinion
Chintala: Inference Moats from Fast CUDA Kernels Last Only Months
“I think, like, Together and Fireworks and all these people are trying to build some faster CUDA kernels and faster, like, you know, hardware kernels in general. But those modes only last for a month or two. Like, these ideas quickly propagate.”
Soumith Chintala Mar 6, 2024 ▶ 28:36 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Opinion
Yegge: Software engineers ignoring AI coding assistants risk career obsolescence
“If you're one of those engineers, man, you better start like, you know, planning another career. Okay. Because this stuff is in the future and it's honestly, it takes some effort to actually make coding assistants work today, right? You have to, you know, just…”
Steve Yegge Dec 17, 2023 ▶ 1:32:16 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Opinion
Patel: Frontier AI model training costs are effectively irrelevant
“In my opinion, I think that's a little bit spicy, but yeah, it's like training costs are irrelevant, right? Like GPT-IV, right? Like 20,000 A-one hundreds. That's like, I know it sounds like a lot of money.”
Dylan Patel Dec 5, 2023 ▶ 4:34 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Opinion
Patel: Fine-tuning existing small models for cloud use is useless
“Unless, unless you're fine tuning for on device use, I think fine tuning current existing models, especially the smaller ones is a useless waste of time, right?”
Dylan Patel Dec 5, 2023 ▶ 30:56 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Opinion
Royzen: AI dev tools do not need to own the IDE
“Somewhere where I disagree with him is that you need to own the IDE. I think like he made kind of some good points about, you know, not having platform risk in the long term, but some of the, you know, features that were mentioned, like suggesting diffs, for e…”
Michael Royzen Nov 3, 2023 ▶ 24:29 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Opinion
Howard: Meta 'blew it' on Code Llama due to catastrophic forgetting
“So Code Llama was a, I think it was like a five hundred billion token fine-tuning of Llama II using code. And also prose about code that Meta did. And honestly, they kind of blew it. Because Code Llama is good at coding, but it's bad at everything else.”
Jeremy Howard Oct 20, 2023 ▶ 43:26 The End of Finetuning — with Jeremy Howard of Fast.ai
Opinion
Howard: TensorFlow 2 was a failure that Google avoided internally
“Then in the end, you know, Google didn't follow through, which is fair enough, like, asking everybody to, you know, learn a new programming language is going to be tough, but, like, it was very obvious, very, very obvious at that time that TensorFlow II was go…”
Jeremy Howard Oct 20, 2023 ▶ 59:34 The End of Finetuning — with Jeremy Howard of Fast.ai
Opinion
Howard: RAG is an inefficient hack compared to fine-tuning
“RAG is like such a inefficient hack, really, isn't it? It's like, You know, segment up my data in some somewhat arbitrary way, embed it, ask questions about that, you know, hope that my embedding, you know, model embeds questions in the same embedding space as…”
Jeremy Howard Oct 20, 2023 ▶ 1:03:36 The End of Finetuning — with Jeremy Howard of Fast.ai
Opinion
Hotz: Analog computing for AI won't work and clockless chips aren't practical
“Analog computing just won't work. And clockless computing sure, it might work in theory, but your ETA tools are, maybe AIs will be able to design clockless chips, but not humans.”
George Hotz Jun 20, 2023 ▶ 7:50 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Opinion
Hotz: Systolic arrays are the wrong architectural choice for AI chips
“I think systolic arrays are the wrong choice. Systolic array, I think they have systolic arrays because that was the guy's PhD. And of course Amazon makes... They are very power efficient, but it becomes hard to schedule a lot of stuff. On them, if you're not …”
George Hotz Jun 20, 2023 ▶ 9:12 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Opinion
Swyx found Codex preferable to human marketing consultants for ads
“Codex has been my CMO for the last like two months. Basically running our ads and giving me advice on AEO and SEO and what have you. And I find it much more preferable to talking to a real human consultant, which I've also done. And it's roughly the same resul…”
Shawn Wang Sep 7, 2026 ▶ 33:35 Orbs: Shifting Coding to Cloud — Quinn Slack, Amp Code
Opinion
Untapped AI hardware opportunity is co-designing models for non-Nvidia architectures
“What I think is the most untapped opportunity right now frankly for Cerebrus, but frankly for the entire non-NVIDIA environment right? And the non-environ sorry community is that, you know, we're all running models that were designed for NVIDIA GPUs, right? An…”
Sean Lie Sep 2, 2026 ▶ 30:35 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.