The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 2,445 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 100 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Prediction Not checkable as stated
Hou: 99% of AI editor rules file contents will be automatically inferred
“We strongly believe that having a rules file, you know, we do allow users to add a rules file, we strongly believe that a rules file is a crutch. You know, by the end of twenty-twenty-five, 99% of the things that you're gonna put in a rules file will be interp…”
Kevin Hou Jul 28, 2025 ▶ 2:57:45 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Prediction Open · timeframe Dec 2029
Kamradt Predicts ARC-AGI-3 Benchmark Will Remain Unbeaten For 3 Years
“And then V three, our durability estimate for that is three years. And that's what we're aiming for is 36 months for V three.”
Greg Kamradt Jul 18, 2025 ▶ 28:35 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Prediction Not checkable as stated
Bergum: Pinecone will not endure like MongoDB because it is too narrow
“So there's always like this convergence, but MongoDB kind of, it sticks, but I don't think that for Pinecoin that was originally leading that movement, it won't like stick in the same way. It's too narrow. It's too, too narrow.”
Jo Kristian Bergum Apr 19, 2025 ▶ 7:56 The Rise and Fall of the Vector DB category: Jo Kristian Bergum (ex-Chief Scientist, Vespa)
Prediction Not checkable as stated
Conrad: Hyperscalers will probably lose significant money reselling Nvidia GPUs
“My intuition is that the hyperscalers are probably going to lose a lot of money, and they know they're going to lose a lot of money on reselling NVIDIA GPUs at least.”
Evan Conrad Apr 11, 2025 ▶ 6:28 SF Compute: Commoditizing Compute
Prediction Not checkable as stated
Conrad: Decentralized compute networks will never beat co-located InfiniBand clusters
“I just don't really think this is gonna ever be more efficient than a fully interconnected cluster with Infiniband, or, you know, whatever sort of next spec might be. Like, I could be completely wrong, but Speedolite is really hard to beat. And regardless of w…”
Evan Conrad Apr 11, 2025 ▶ 34:33 SF Compute: Commoditizing Compute
Prediction Not checkable as stated
Conrad: The AI VC bubble will pop and fail to return capital
“So what you've done by not having a future is you've inflated the venture capital market. And that is a bubble that's totally going to pop at some point. Like a lot of the companies are not going to work. And the valuations are not going to work. And what's go…”
Evan Conrad Apr 11, 2025 ▶ 59:31 SF Compute: Commoditizing Compute
Prediction Not checkable as stated
Beauchamp: Timeline for AGI has certainly been pushed out
“I think the odds, the timeline for AGI has certainly been pushed out, right?”
William Beauchamp Jan 26, 2025 ▶ 52:02 Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
Assertion Not checkable as stated
Zhang: Meta Failed at Training MoE Models for Llama Series
“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
Yining Zhang Jan 19, 2025 ▶ 14:53 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Assertion Supported
Bryk: Perplexity and ChatGPT Search rely on legacy Google and Bing APIs
“So these systems, there are a few of them now they basically rely on like traditional search engines like Google or Bing, and then they combine them with like LLMs at the end to, you know, output some power graphics answering your question. So they, Like, Sear…”
Will Bryk Jan 10, 2025 ▶ 21:16 Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
Prediction Not checkable as stated
Swyx: AI models hit a 2T parameter wall, won't reach 10T
“As far as everyone is concerned, Claude, you know, Opus 3.5 is not coming out. GPT 4.5 is not coming out. And Gemini two, like we don't have pro whatever we've hit that wall, whatever that wall is. Maybe I'll call it like the two trillion parameter wall. Like …”
Shawn Wang Jan 1, 2025 ▶ 40:13 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Loubna Ben Allal Dec 24, 2024 ▶ 4:12 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Assertion Not checkable as stated
Fu: Embedding model quality barely matters for final RAG performance
“We had this experience over and over again where you could have any, an embedding model of any quality, so you could have a really, really bad embedding model, or you could have a really, really good one by, and by any measure of good, and for the final RAG ap…”
Dan Fu Dec 24, 2024 ▶ 33:00 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Prediction Not checkable as stated
Specialized open-source expert models will outperform one-size-fits-all closed-source models
“And that's our prediction is With specialization, there will be a lot of expert models, really, really good, and even better than, like, one size fits all open source closed source model.”
Lin Qiao Nov 25, 2024 ▶ 33:45 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Assertion Not checkable as stated
Crivello: Lindy AI agents perform better than humans for most use cases
“I think the bar is it needs to be better than a human. And for most use cases we serve today, it is better than the human, especially if you put it on rails.”
Florent Crivello Nov 15, 2024 ▶ 30:22 Agents @ Work: Lindy.ai (with live demo!)
Assertion Not checkable as stated
Crivello: Every major remote success story lost to an in-person competitor
“In every one of these examples, you have a co-located counterfactual that is sometimes orders of magnitude bigger.”
Florent Crivello Nov 15, 2024 ▶ 49:48 Agents @ Work: Lindy.ai (with live demo!)
Prediction Not checkable as stated
Goyal: Agent control flow and graph routing will move into models
“It feels very clear to me that this type of logic is going to be built into the model. Anytime there is control flow complexity or uncertainty complexity, I think the history of AI has been to push more and more into the model.”
Ankur Goyal Oct 11, 2024 ▶ 1:33:10 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Goyal: OpenAI o1 will make agentic frameworks obsolete
“And I think O-one is going to do that to agentic frameworks as well. Hey, I think To me, it seems very unlikely that the, you know, you and me sort of like sipping an espresso and thinking about how, like, different personified roles of people should interact …”
Ankur Goyal Oct 11, 2024 ▶ 1:34:00 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Shunyu Yao predicts training models on human computer trajectories achieves AGI
“The simplest way to achieve AGI is literally just record the re-actuatory of every human being and just put them together, you know, like what do you have thought about? What do you have done? Let's say on the computer, right? Imagine like solid experiment. Li…”
Shunyu Yao Sep 27, 2024 ▶ 1:02:07 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Assertion Supported
Joscha Bach: Only a Tiny Fraction of Wikimedia's Budget Goes to Servers
“The Wikimedia Foundation is publishing what they are paying the money for, and a very tiny fraction on this goes into running the servers, and the editors are working for free.”
Joscha Bach Apr 27, 2024 ▶ 1:53:10 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Prediction Not checkable as stated
Chintala: George Hotz's TinyGrad requires major breakthroughs to match PyTorch
“There's no, like, I don't think, like, unless we have, like, great breakthroughs, like, George's vision is achievable, like, or, like, he should be thinking about a narrower problem, such as, I'm only gonna make this for, like, work for self-driving car con ne…”
Soumith Chintala Mar 6, 2024 ▶ 9:40 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Prediction Not checkable as stated
Chintala: Apple's MLX will fail server-side due to lack of differentiation
“If they end up expanding onto the server side, and they'll probably build something like PyTorch as well, right? Like, eventually, that'll where it will land. And I think there, they will kind of fail on the, like, lack of differentiation. Like, it wouldn't be…”
Soumith Chintala Mar 6, 2024 ▶ 18:12 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Prediction Open · timeframe Dec 2026
Patel: OpenAI and Microsoft partnership will likely collapse within years
“Yeah, I expect in the next few years that the OpenAI and Microsoft probably falls apart too.”
Dylan Patel Dec 5, 2023 ▶ 46:55 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Prediction Not checkable as stated
Royzen: Fine-tuned open source models will beat proprietary in 2024
“So I think that even if a delta exists, in twenty-twenty-four, the delta between proprietary and open source won't be large enough that a startup like us, with a lot of data that we've collected, can take the data that we have, fine-tune an open source model, …”
Michael Royzen Nov 3, 2023 ▶ 38:25 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Jeremy Howard Oct 20, 2023 ▶ 15:41 The End of Finetuning — with Jeremy Howard of Fast.ai
Assertion Not checkable as stated
Howard: JAX was a grassroots Google reaction against TensorFlow 2
“But I mean, in the meantime, I will say, you know, Google now does have a backup plan. You know, they have JAX, which was never a strategy. It was just a bunch of people who also recognized TensorFlow two as shit, and they just decided to build something else.”
Jeremy Howard Oct 20, 2023 ▶ 1:01:58 The End of Finetuning — with Jeremy Howard of Fast.ai
Prediction Not checkable as stated
Hotz: Tinybox will be 5x faster than H100 systems per dollar
“For 90% of most companies model training use cases, the tiny box will be five X faster for the same price.”
George Hotz Jun 20, 2023 ▶ 48:45 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Prediction Not checkable as stated
Hotz: Machines will replace all human labor in about 20 years
“I'm a believer that machines are gonna replace everything in about 20 years.”
George Hotz Jun 20, 2023 ▶ 55:44 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Assertion Not checkable as stated
Jenik: Accelerated Understanding achieved 5-trillion token inference context length
“Like we're able to train up to a trillion context input. We're able to train With like, even inputs, outputs, both trillion context length, we are able to do inference at five trillion contexts.”
Benedikt Jenik Sep 4, 2026 ▶ 6:25 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Assertion Open · timeframe Sep 2029
Anandkumar: Multi-physics models outperform single-physics models of equivalent parameter size
“And in fact, I was going to add that it turns out that having the model of the same size with multiple areas of physics does better than giving all of those parameters to each single physics. So if you had separate models and made them big enough as the origin…”
Anima Anandkumar Sep 4, 2026 ▶ 8:08 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Prediction Open · timeframe Dec 2027
Sean Lie: Next Cerebras Chip Will Run Frontier AI at 5,000 TPS
“And in particular, we designed it together with our next generation wafer chip that will be coming next year. And with that chip, We'll be pushing the performance even further. So we just saw a two X improvement this year with CS four. We're going to push it e…”
Sean Lie Sep 2, 2026 ▶ 11:01 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Sean Lie Sep 2, 2026 ▶ 23:58 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Not checkable as stated
Sean Lie: Cerebras has solved yield at scale and 3D packaging
“And so, you know, we already have solved yield at scale. For example, we've already solved how you can actually package in a three-dimensional way. And that's, You know, problems that Samsung, that D-Matrix, and everybody else are also gonna have to solve over…”
Sean Lie Sep 2, 2026 ▶ 40:25 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Prediction Not checkable as stated
Transformers will never scale to high-resolution 4D physics simulations
“So forget ever having a transformer for anything of this scale. All of the world's compute will not be enough. And first of all, they all have to be co-located to be able to ever do this. So that's why we need other architectures.”
Anima Anandkumar Aug 26, 2026 ▶ 29:41 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Contradicted
Neural operators are the only AI architecture that works for climate emulation
“This is where the Allen AI Institute has now built climate models based on our neural operator architecture. And that's the only one that works As an AI emulator, right? None of the other architectures work for climate because climate requires us to assume the…”
Anima Anandkumar Aug 26, 2026 ▶ 37:44 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
AI models predict fusion reactor plasma disruption one million times faster
“You know, I talk about plasma and fusion reactor. You know, we barely have a few thousand samples, but we are able to accurately predict events like disruption very well. And we are able to do that a million times faster than what traditional simulations were …”
Anima Anandkumar Aug 26, 2026 ▶ 43:50 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Assertion Supported
Park: Generative agent digital twins replicate human behavior at 85% accuracy
“And this is where we basically could replicate people's behaviors and attitudes, 85% as accurately as people would replicate their own. So that actually was the first really paper that gave this validated results that we can actually model individuals in an ac…”
Joon Sung Park Aug 21, 2026 ▶ 29:20 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Assertion Not checkable as stated
Park: Frontier models hit only 20-30% accuracy predicting niche human behavior
“Where in some cases, the model performance of frontier models go all the way down to 20, 30%. Especially if you go into that more niche population on topics that our customers will actually care about. On more gen pop, it might be around 50 to 60%.”
Joon Sung Park Aug 21, 2026 ▶ 31:17 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Prediction Not checkable as stated
Park: Future simulations will cost as much to run as training models
“My hunch here is I do think in the next Some number of years, we will start creating simulations that will actually cost as much as training a foundation model.”
Joon Sung Park Aug 21, 2026 ▶ 49:58 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Prediction Not checkable as stated
Krentsel: AI takeoff flywheel will emerge from agents iterating on harness code
“There's a moment like that happening now, and this is the moment that's going to be a flywheel, I feel like, because this takeoff moment will come from being iterating in the same layer that you are producing, I think.”
Alex Krentsel Aug 15, 2026 ▶ 45:39 Exo: Harnesses should see their own code and logs — Alex Krentsel, UC Berekeley / Google Research
Assertion Not checkable as stated
McPartlon: AI models are nearing direct output of viable drug molecules
“We're kind of at the inflection point now. We're really seeing this internally at CHI, where the models are getting pretty close to, like, producing Molecules that could eventually, or like are very close to drugs.”
Matt McPartlon Aug 11, 2026 ▶ 48:31 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Prediction Not checkable as stated
Patil: Biomolecular AI models will be as massive and impactful as LLMs
“And, you know, I think this class of models is going to be like just as big, just as impactful as LLMs, but it's almost like the compute market, like kind of doesn't realize that yet, both in the capacity sense, but also in like the software stack sense.”
Neil Patil Aug 11, 2026 ▶ 1:03:48 🔬They Thought the Model Was Broken — Matt McPartlon & Neil Patil, Chai Discovery
Assertion Not checkable as stated
Companies developing fused mega kernels rarely run them in production
“Even the companies that have worked at or people that have spoken to who work at companies that do fused mega kernels, they, Very, very often don't end up running those in production because the TRTL and modular kernels that will launch are faster because you …”
Ali Taha Aug 3, 2026 ▶ 58:38 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Prediction Not checkable as stated
NVIDIA's Vera Rubin GPU architecture effectively renders mega kernel research obsolete
“The GPU is, is It's designed in such a way that it basically kills megakernels. You don't need to use megakernels that much anymore. So it seems like that entire research field goes into, like, won't be continued, but yeah.”
Ali Taha Aug 3, 2026 ▶ 59:20 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Prediction Not checkable as stated
Long-form AI video generation must switch to autoregressive architectures over diffusion
“Autoregressive video seems to me like that is the bet that the future is going to be making. But there are no good open source autoregressive video models out there today. And that seems to be, if you want to get like an hour movie, if you want to see video mo…”
Ali Taha Aug 3, 2026 ▶ 1:18:07 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Prediction Not checkable as stated
Kant: Catching up to frontier AI labs will soon become unfeasible
“We've got a small window before models are Really impacting recursive self-improvement to a level where catching up otherwise might become unfeasible.”
Eiso Kant Jul 22, 2026 ▶ 33:51 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.