The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

why aren't all 5,957 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 100 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Prediction Not checkable as stated
Swyx: AI models hit a 2T parameter wall, won't reach 10T
“As far as everyone is concerned, Claude, you know, Opus 3.5 is not coming out. GPT 4.5 is not coming out. And Gemini two, like we don't have pro whatever we've hit that wall, whatever that wall is. Maybe I'll call it like the two trillion parameter wall. Like …”
Shawn Wang Jan 1, 2025 ▶ 40:13 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Opinion
Neubig: Anthropic's MCP Duplicates Existing APIs With Little Added Value
“We already have an API for GitHub. So why do we need an MCP for GitHub, right? You know, like GitHub has an API. The GitHub API is evolving. We can look up the GitHub API documentation. So it seems like kind of duplicated a little bit. And also they have a set…”
Graham Neubig Dec 25, 2024 ▶ 41:29 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:12 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Opinion
Ben Allal: Synthetic data may enrich the web rather than pollute it
“So personally, I wouldn't say the web is posted with synthetic data. Maybe it's even making it more rich.”
Loubna Ben Allal Dec 24, 2024 ▶ 4:35 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Assertion Not checkable as stated
Fu: Embedding model quality barely matters for final RAG performance
“We had this experience over and over again where you could have any, an embedding model of any quality, so you could have a really, really bad embedding model, or you could have a really, really good one by, and by any measure of good, and for the final RAG ap…”
Dan Fu Dec 24, 2024 ▶ 33:00 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Opinion
Soldani: Open Model Bio-Risk Warnings Were a Lobbying Ploy
“You know, if you remember the beginning of this year, it was all about bio-risk of these open models. The whole thing fizzled out because there's been, finally there's been, like, rigorous research, not just this paper from coherent folks, but there's been rig…”
Luca Soldani Dec 23, 2024 ▶ 23:09 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Opinion
Ramachandran: SWE-bench and HumanEval do not reflect real professional engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have sweet bench. That's cool. No actual Professional work looks like Sweebench. Like, human eval, same thi…”
Anshul Ramachandran Dec 13, 2024 ▶ 13:17 Windsurf: The Enterprise AI IDE
What-if
Varun Mohan: Codeium would have failed if it used vLLM
“If we use VLLM, we would not be talking with you right now.”
Varun Mohan Dec 13, 2024 ▶ 57:27 Windsurf: The Enterprise AI IDE
Insight
Anthropic's Schluntz: Avoid agent frameworks and start from scratch with raw prompts
“I think with agent frameworks in general, they can certainly save you some like boilerplate, but I think there's actually this like downside of making agents too easy, where you end up very quickly, like building a much more complex system than you need. And s…”
Erik Schluntz Nov 28, 2024 ▶ 43:01 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Prediction Not checkable as stated
Specialized open-source expert models will outperform one-size-fits-all closed-source models
“And that's our prediction is With specialization, there will be a lot of expert models, really, really good, and even better than, like, one size fits all open source closed source model.”
Lin Qiao Nov 25, 2024 ▶ 33:45 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Assertion Not checkable as stated
Crivello: Lindy AI agents perform better than humans for most use cases
“I think the bar is it needs to be better than a human. And for most use cases we serve today, it is better than the human, especially if you put it on rails.”
Florent Crivello Nov 15, 2024 ▶ 30:22 Agents @ Work: Lindy.ai (with live demo!)
Opinion
Crivello: OpenAI's GPT-4o Is Overhyped and Poor for Agents
“I think four O is overhyped. Frankly, we don't use four O. I don't think it's good for agentic behavior.”
Florent Crivello Nov 15, 2024 ▶ 32:07 Agents @ Work: Lindy.ai (with live demo!)
Assertion Not checkable as stated
Crivello: Every major remote success story lost to an in-person competitor
“In every one of these examples, you have a co-located counterfactual that is sometimes orders of magnitude bigger.”
Florent Crivello Nov 15, 2024 ▶ 49:48 Agents @ Work: Lindy.ai (with live demo!)
Insight
Crivello: Remote work optimizes for cost; building creative software requires in-person collaboration
“If you're optimizing for cost, absolutely be remote. If you're optimizing for creativity, which I think that software and product building is a creative endeavor, if you're optimizing for creativity, it's kind of like composing an album. You can't do it on the…”
Florent Crivello Nov 15, 2024 ▶ 50:36 Agents @ Work: Lindy.ai (with live demo!)
Opinion
Polu: The bulk of useful enterprise agent work can use APIs
“The bulk of the useful stuff that you can do within the company can be done through API. The data can be retrieved by API, the actions can be taken through API.”
Stanislas Polu Nov 11, 2024 ▶ 32:30 Agents @ Work: Dust.tt — with Stanislas Polu
Opinion
Drew Houston: Large language models are a rapidly self-commoditizing, bad business
“Large language models are a pretty bad Business from a, you know, you sort of take off your tech lens and just sort of business lens. Like there's sort of this weirdly self-commoditizing thing where, you know, models only have value if they're kind of on this…”
Drew Houston Oct 18, 2024 ▶ 29:50 Building the Silicon Brain - Drew Houston of Dropbox
Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Ankur Goyal Oct 11, 2024 ▶ 1:27:25 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Goyal: Agent control flow and graph routing will move into models
“It feels very clear to me that this type of logic is going to be built into the model. Anytime there is control flow complexity or uncertainty complexity, I think the history of AI has been to push more and more into the model.”
Ankur Goyal Oct 11, 2024 ▶ 1:33:10 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Goyal: OpenAI o1 will make agentic frameworks obsolete
“And I think O-one is going to do that to agentic frameworks as well. Hey, I think To me, it seems very unlikely that the, you know, you and me sort of like sipping an espresso and thinking about how, like, different personified roles of people should interact …”
Ankur Goyal Oct 11, 2024 ▶ 1:34:00 Production AI Engineering starts with Evals
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:24 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Shunyu Yao predicts training models on human computer trajectories achieves AGI
“The simplest way to achieve AGI is literally just record the re-actuatory of every human being and just put them together, you know, like what do you have thought about? What do you have done? Let's say on the computer, right? Imagine like solid experiment. Li…”
Shunyu Yao Sep 27, 2024 ▶ 1:02:07 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Opinion
Schulhoff: Role Prompting Does Not Improve Accuracy on Modern LLMs
“For accuracy-based tasks, like MMLU, you're trying to solve a math problem, and maybe you tell the AI that it's a math professor, and you expect it to have improved performance. I really don't think that works. I'm quite certain that doesn't work on more moder…”
Sander Schulhoff Sep 20, 2024 ▶ 17:08 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Insight
Tay: Frontier AI researchers cannot maintain standard nine-to-five work-life balance
“You cannot be, like, checking out on, like, Friday, Saturday, Sunday, and, like, work at, like, nine to five if you want to, like, Make progress, or like, some people are just so good at detaching, like, ok, like, you know, like, eight pm, I'm not going to, my…”
Yi Tay Jul 5, 2024 ▶ 38:49 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Opinion
Tay: Long context architecture is the future of AI over RAG
“And, yeah, I mean, I think long context is definitely the future, rather than rec. But I mean, they could be used in conjunction, like,”
Yi Tay Jul 5, 2024 ▶ 1:40:05 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Insight
Albrecht: LLM emergence is an artifact of non-linear evaluation metrics
“This emergent behavior that you're seeing, Is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
Josh Albrecht Jun 25, 2024 ▶ 55:56 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Insight
Translating code proves LLMs possess causal, functional understanding
“If you ask the LLM to translate a bit of Python into a little bit of C, and it's performing this task, obviously it is understanding in the sense that it has a, The causal, functional model it implements.”
Joscha Bach Apr 27, 2024 ▶ 1:03:36 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Insight
Bach: Claude is implemented as an invariant pattern similarly to consciousness
“Claude exists only as a pattern. It's something that is a pattern in the activation of the transistors. And even transistors don't actually exist. They are A pattern in the atoms that we are able to see as an invariance because we tune the atoms in a particula…”
Joscha Bach Apr 27, 2024 ▶ 1:20:37 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Opinion
Big AI labs build AGI internally while giving the public child-proof apps
“In some sense, the big AI companies are incentivized and interested in building AGI internally, and giving everybody else a child-proof application.”
Joscha Bach Apr 27, 2024 ▶ 1:27:38 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Opinion
Bach: Prompting an LLM correctly can make it sentient to some degree
“It's a Weltgeist that gets possessed by a prompt. And if you possess it with the right prompt, then it can become sentient to some degree.”
Joscha Bach Apr 27, 2024 ▶ 1:44:52 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Supported
Joscha Bach: Only a Tiny Fraction of Wikimedia's Budget Goes to Servers
“The Wikimedia Foundation is publishing what they are paying the money for, and a very tiny fraction on this goes into running the servers, and the editors are working for free.”
Joscha Bach Apr 27, 2024 ▶ 1:53:10 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Insight
Liu: LLM observability startups ignore full systems; just use Postgres
“The issue really is the fact that these observability companies isn't actually doing observability for the system, it's just doing the LLM thing. Like I still end up using like Datadog, right? Or like, you know, Sentry to do, like, latency. And so I just have …”
Jason Liu Apr 24, 2024 ▶ 32:35 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Insight
Liu: Companies abandon LLM frameworks to regain control over prompts
“So much of it is changing that if you give control of these systems away too early, you end up ultimately wanting them back. Like many companies I know that I reach out or ones were like, oh, we're going off of the frameworks because now that we know what the …”
Jason Liu Apr 24, 2024 ▶ 52:08 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Opinion
Liu: App-layer AI startups should avoid hiring traditional machine learning engineers
“I think a lot of these app layer startups should not be hiring MLEs because they end up churning.”
Jason Liu Apr 24, 2024 ▶ 58:42 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Prediction Not checkable as stated
Chintala: George Hotz's TinyGrad requires major breakthroughs to match PyTorch
“There's no, like, I don't think, like, unless we have, like, great breakthroughs, like, George's vision is achievable, like, or, like, he should be thinking about a narrower problem, such as, I'm only gonna make this for, like, work for self-driving car con ne…”
Soumith Chintala Mar 6, 2024 ▶ 9:40 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Prediction Not checkable as stated
Chintala: Apple's MLX will fail server-side due to lack of differentiation
“If they end up expanding onto the server side, and they'll probably build something like PyTorch as well, right? Like, eventually, that'll where it will land. And I think there, they will kind of fail on the, like, lack of differentiation. Like, it wouldn't be…”
Soumith Chintala Mar 6, 2024 ▶ 18:12 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Opinion
Chintala: Inference Moats from Fast CUDA Kernels Last Only Months
“I think, like, Together and Fireworks and all these people are trying to build some faster CUDA kernels and faster, like, you know, hardware kernels in general. But those modes only last for a month or two. Like, these ideas quickly propagate.”
Soumith Chintala Mar 6, 2024 ▶ 28:36 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Insight
AI Productivity Gains Will Ultimately Increase Demand for Software Engineers
“Like, I think we can, you know, 10 x the amount of developers, and still, you know, have a lot of people making a lot of money, you know, building amazing software, and also being, while at the same time being more productive. Like, I never understood this, li…”
Erik Bernhardsson Feb 19, 2024 ▶ 1:01:23 Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
Insight
Hsu: Higher valuation for the same quality company is always worse
“Higher valuation, given the same company quality, is always worse.”
David Hsu Feb 7, 2024 ▶ 12:44 The State of AI in production — with David Hsu of Retool
Opinion
Yegge: Software engineers ignoring AI coding assistants risk career obsolescence
“If you're one of those engineers, man, you better start like, you know, planning another career. Okay. Because this stuff is in the future and it's honestly, it takes some effort to actually make coding assistants work today, right? You have to, you know, just…”
Steve Yegge Dec 17, 2023 ▶ 1:32:16 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Opinion
Patel: Frontier AI model training costs are effectively irrelevant
“In my opinion, I think that's a little bit spicy, but yeah, it's like training costs are irrelevant, right? Like GPT-IV, right? Like 20,000 A-one hundreds. That's like, I know it sounds like a lot of money.”
Dylan Patel Dec 5, 2023 ▶ 4:34 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Opinion
Patel: Fine-tuning existing small models for cloud use is useless
“Unless, unless you're fine tuning for on device use, I think fine tuning current existing models, especially the smaller ones is a useless waste of time, right?”
Dylan Patel Dec 5, 2023 ▶ 30:56 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Prediction Open · timeframe Dec 2026
Patel: OpenAI and Microsoft partnership will likely collapse within years
“Yeah, I expect in the next few years that the OpenAI and Microsoft probably falls apart too.”
Dylan Patel Dec 5, 2023 ▶ 46:55 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Opinion
Royzen: AI dev tools do not need to own the IDE
“Somewhere where I disagree with him is that you need to own the IDE. I think like he made kind of some good points about, you know, not having platform risk in the long term, but some of the, you know, features that were mentioned, like suggesting diffs, for e…”
Michael Royzen Nov 3, 2023 ▶ 24:29 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Prediction Not checkable as stated
Royzen: Fine-tuned open source models will beat proprietary in 2024
“So I think that even if a delta exists, in twenty-twenty-four, the delta between proprietary and open source won't be large enough that a startup like us, with a lot of data that we've collected, can take the data that we have, fine-tune an open source model, …”
Michael Royzen Nov 3, 2023 ▶ 38:25 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Jeremy Howard Oct 20, 2023 ▶ 15:41 The End of Finetuning — with Jeremy Howard of Fast.ai
Opinion
Howard: Meta 'blew it' on Code Llama due to catastrophic forgetting
“So Code Llama was a, I think it was like a five hundred billion token fine-tuning of Llama II using code. And also prose about code that Meta did. And honestly, they kind of blew it. Because Code Llama is good at coding, but it's bad at everything else.”
Jeremy Howard Oct 20, 2023 ▶ 43:26 The End of Finetuning — with Jeremy Howard of Fast.ai
Opinion
Howard: TensorFlow 2 was a failure that Google avoided internally
“Then in the end, you know, Google didn't follow through, which is fair enough, like, asking everybody to, you know, learn a new programming language is going to be tough, but, like, it was very obvious, very, very obvious at that time that TensorFlow II was go…”
Jeremy Howard Oct 20, 2023 ▶ 59:34 The End of Finetuning — with Jeremy Howard of Fast.ai
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.