The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 2,445 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 100 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Prediction Not checkable as stated
Kant: Reinforcement learning will move earlier into LLM pre-training
“I have I would say a not commonly held opinion that reinforcement learning will move earlier and earlier into pre-training.”
Eiso Kant Jul 22, 2026 ▶ 45:09 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Supported
Kant: Laguna S outperforms models two to three times its size
“When you look at the benchmarks and start using it, you'll realize that we are outperforming models two or three times their size.”
Eiso Kant Jul 22, 2026 ▶ 57:50 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Prediction Open · timeframe Jul 2027
Kant: Prompts stuffed with dozens of tools will vanish in 12 months
“I think we will in 12 months not see a single system prompt that is stuffed with 20 or 30 or 40 tools anymore.”
Eiso Kant Jul 22, 2026 ▶ 1:13:39 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Prediction Not checkable as stated
Kant: AI will become the world's most demanded commodity with commoditizing margins
“Intelligence is the most in life. You're going to be the world's most demanded commodity. It will more commoditize in margin and price.”
Eiso Kant Jul 22, 2026 ▶ 1:28:02 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Assertion Not checkable as stated
Wang: X-Cell Is First to Predict Unseen Cell Line Perturbations
“One of the rewarding signals I receive after we develop Excel is that like it's a wow moment from biologists that this is the first time biologists actually find the model can predict exactly how these unseen cell lines kind of respond to different perturbatio…”
Bo Wang Jul 21, 2026 ▶ 7:54 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Prediction Not checkable as stated
Wang: Virtual cell AI models will eventually replace physical cellular experiments
“Eventually we kind of, we can replace all the cellular experiments by simply running simulations on computer without even running the actual Experiments.”
Bo Wang Jul 21, 2026 ▶ 17:35 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Supported
Chu: Descriptive Models Fail to Beat Linear Baselines on Causal Biology
“Models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks.”
Ci Chu Jul 21, 2026 ▶ 21:47 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Supported
Wang: Diffusion Outperforms Autoregressive Models on Unseen Cellular Tasks
“We find that switching from autoregressive training to division language models give a significant improvement over some of the harder tasks, particularly generalized to unseen tasks.”
Bo Wang Jul 21, 2026 ▶ 52:00 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Prediction Not checkable as stated
Wang: Foundation models on causal data will beat linear baselines
“I believe that foundation model or other more complicated AI models that trend on the right data will outperform these linear models in harder tasks, particularly in generalization tasks.”
Bo Wang Jul 21, 2026 ▶ 1:04:25 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Supported
Wang: X-Cell Can Predict Combinatorial Gene Perturbations In Silico
“This is also why we incorporate PPI networks as the prior knowledge into our model. And although the model right now are trained on single gene perturbations, but once the model is trained, you can actually predict combinatorial perturbations just on the model…”
Bo Wang Jul 21, 2026 ▶ 1:09:55 🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
Assertion Not checkable as stated
Beam: Lila's AI hits 80% zero-shot on gene editing, beating humans' 0%
“Certainly for expression protocols, for some gene editing work that we've done we have tested like the platform's ability to do that versus humans. Model gets like 80% of that zero shot. Humans get zero percent of that zero shot.”
Andy Beam Jul 16, 2026 ▶ 13:47 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Assertion Not checkable as stated
Beam: Lila's best non-platinum electrocatalysts came from AI ideas experts called stupid
“Some of the suggestions from the model initially were boring, but then transitioned from boring to what he considered to be stupid. These are non-platinum group electrocatalysts for separation of hydrogen and oxygen from water to make hydrogen, and those turns…”
Andy Beam Jul 16, 2026 ▶ 17:53 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Assertion Not checkable as stated
Beam: Lila's 10-trillion-token general science model beats specialized AI
“So we have assembled this reasoning data set of 10 trillion scientific tokens reasoning traces that are experimentally verified across life sciences, chemistry, and material sciences, and we have seen that this general model often beats the domain-specific mod…”
Andy Beam Jul 16, 2026 ▶ 32:39 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Assertion Not checkable as stated
Beam: Lila's in vivo CAR-T data outperformed Capstan in non-human primates
“So we have developed some monster UTRs, untranslated regions, which flank the protein coding region which dictate those expression properties. Something like Tenex, the references from Moderna and Pfizer. And over the course of six months, got to in vivo data …”
Andy Beam Jul 16, 2026 ▶ 48:48 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Prediction Not checkable as stated
Beam: Round-over-round experimentation yields more compound value than broad datasets
“The bet is that the sort of like as the model performance improves, the sample efficiency goes up, and therefore like the compound interest that you get from round over round experimentation will outweigh that, that you would get from a big noisy, but broad da…”
Andy Beam Jul 16, 2026 ▶ 1:13:48 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Prediction Not checkable as stated
Biderman: AI-native companies will amass trillions of internal tokens within 18 months
“In 18 months, many companies would have maybe trillions of tokens, which of internal company data, proprietary data. I'm talking about like maybe trillions. It sounds exaggerated, but I don't think it's an impossibility if they're really AI native.”
Dan Biderman Jul 13, 2026 ▶ 14:42 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Not checkable as stated
Biderman: Hard engineering tasks will require test-time gradient updates
“We think that eventually part of the solution for very hard tasks in, in science and engineering and defense and all that stuff will involve some form of gradient based updates during during doing these long horizon tasks.”
Dan Biderman Jul 13, 2026 ▶ 21:57 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Not checkable as stated
Biderman: In 18 months, data scale will require weight-based learning
“Other parts of it are bets that in 18 months from now, the scale of the data will require the methods that we know from pre-training work.”
Dan Biderman Jul 13, 2026 ▶ 26:55 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Open · timeframe Jul 2031
Biderman: PC hardware will soon run near-trillion-parameter models locally
“And in the long, long term, I do think these things will actually run on people's devices, and we're seeing right now the new hardware on personal computers is already, ah, you know, soon approaching the ability to run inference on close to trillion parameters…”
Dan Biderman Jul 13, 2026 ▶ 28:12 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Not checkable as stated
Biderman: AI models must learn to autonomously filter out erroneous user feedback
“Increasingly the models will get better, and increasingly they'll know more things than we do, so the model in some way has to learn and understand and kind of, like, discern what, which feedback is valuable and which feedback should be ignored.”
Dan Biderman Jul 13, 2026 ▶ 33:54 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Assertion Not checkable as stated
Perszyk: Current AI agents remain too unreliable to automate substantial work
“We look at what the metrics of the actual agents and they're so unreliable that ironically we feel a little bit better. The AI is actually not where we need it to be. To automate enough of the work.”
Danielle Perszyk Jul 11, 2026 ▶ 4:48 Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
Assertion Supported
Perszyk: AI writing suggestions subconsciously shift users to opposing arguments
“There are studies that show that people will, even below their threshold of awareness, start with one argument and then be switched to a completely different, maybe opposing argument because of accepting all of these AI suggestions.”
Danielle Perszyk Jul 11, 2026 ▶ 35:35 Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
Assertion Supported
Perszyk: AI tools boost individual output but narrow overall scientific research
“Individual scientists who are using AI tools are benefiting because they are producing more papers. They are getting more grants accepted. But science as a whole is narrowing.”
Danielle Perszyk Jul 11, 2026 ▶ 36:16 Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
Prediction Not checkable as stated
Swyx: AI p(doom) over the next ten years is near zero
“I mean, if you do it in 10 years is near zero.”
Shawn Wang Jul 10, 2026 ▶ 15:09 Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv
Prediction Not checkable as stated
Swyx: LLMs will plateau and potentially trigger a 30-year AI winter
“Yeah, probably LLMs are going to run out at some point and they're not AGI and okay, we have maybe another 30 years of AI winter or something and then like the next paradigm really is actually the thing.”
Shawn Wang Jul 10, 2026 ▶ 15:39 Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv
Prediction Not checkable as stated
Swyx: Specialized agent labs like Cursor, Cognition, and Harvey will endure
“There will always be capability overhangs. They may not stay still, and so you gotta be nimble. But the Sierras of the world, the Cognitions of the world, the Cursors of the world, the Decagons and Harvey's, these are all agent labs for their field. They can b…”
Shawn Wang Jul 10, 2026 ▶ 23:47 Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv
Assertion Supported
Feinberg: Studies show AlphaFold structures provided no value for drug docking
“There is this, ah, a few papers that came out, one was in Cell, I think last year, which showed that for all of the claims about AlphaFold-solving drug discovery, people try to take AlphaFold-produced protein structures, use them for traditional docking, and f…”
Evan Feinberg Jun 30, 2026 ▶ 24:24 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Prediction Not checkable as stated
Feinberg: Chipmakers will invest in life sciences as LLM alpha shrinks
“I do think that chip makers, including Nvidia, are going to want to get a lot more invested in life sciences because it will always be high in demand. And The amount of alpha left in pure LLM space is just getting a little questionable.”
Evan Feinberg Jun 30, 2026 ▶ 1:44:44 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Assertion Not checkable as stated
OpenAI's Chen: AI models already discover novel theorems and advance sciences
“The initial direction we took was you should move it to real world research, right? And we've seen that the models, they've gotten a lot better at just kind of discovering novel theorems and pushing the frontiers of hard sciences. Even today, right, that's no …”
Mark Chen Jun 25, 2026 ▶ 7:27 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Prediction Not checkable as stated
Mark Chen: AI scaling laws will continue to hold
“And so I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering, more careful scaling, and it always unlocks that next ability to scale further. So I mean, it's held for You know, almost 10 order…”
Mark Chen Jun 25, 2026 ▶ 9:58 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
Prediction Not checkable as stated
Zaharia: Open agent hosting layers will win over proprietary alternatives
“Another way to think about it is like, imagine, you know we, our thing wasn't open. We had some kind of agent hosting thing, but it's not open. And then there is an open one. If you're, which one's gonna win in the long run? So like here, because there is this…”
Matei Zaharia Jun 24, 2026 ▶ 12:09 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Assertion Supported
Xin: Transcoding database rows to Parquet speeds object storage writes with zero compromise
“And as a matter of fact, once you transcode the data compresses better. So from those services writing to, for example, S three or other data lake, like object stores, you can actually write them faster because now they are now smaller. So there's no. Overhead…”
Reynold Xin Jun 24, 2026 ▶ 36:50 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Prediction Not checkable as stated
Xin: Much of traditional software will be rewritten with data and agents
“Actually, I think many of the traditional software will be sort of rewritten with this new paradigm, which is just get the data to be there. And then they slap some agent on top.”
Reynold Xin Jun 24, 2026 ▶ 1:07:44 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Zico Kolter Jun 22, 2026 ▶ 12:14 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Not checkable as stated
Fredrikson: Frontier AI models fall for simulated prompt injections humans would ignore
“While in these scenarios, humans found it very difficult to prompt inject the models, like we're aware of scenarios that a human would never fall for, that like Opus four seven would, right? Like a, you know, an email that comes to your inbox and it says somet…”
Matt Fredrikson Jun 22, 2026 ▶ 22:55 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Prediction Not checkable as stated
Kolter: Security and science will explode as AI agents automate tedious verification
“So I think this is really sort of an underappreciated point that we're reaching this point, this sort of phase where a lot of security, a lot of science has this potential to kind of explode. Not because we're going to get better at it, but because agents can …”
Zico Kolter Jun 22, 2026 ▶ 46:01 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Not checkable as stated
Fredrikson: Gray Swan Found Jailbreaks in Every OpenClaw User Trajectory Tested
“So we just have a bunch of trajectories of actual people using OpenClaw. And tons and tons of different scenarios and just threw shade at it and like found breaks for each and every one of them, right?”
Matt Fredrikson Jun 22, 2026 ▶ 47:36 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Not checkable as stated
Malde: SWE-ONE beat frontier models via user-signal post-training
“And this was the kind of major unlock for the company as well, is we had all this massive data. We were able to post train on all of that user signal and now beat the frontier.”
Ronak Malde Jun 21, 2026 ▶ 3:06 ⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
Assertion Not checkable as stated
Most AI clusters fail to hit Google's 96% node utilization standard
“My co-founder, Seb came from he built the Borg export GQM scheduler at Google, and there, I think, 95% was considered an outage, so 96% node utilization is, should be standard, and most single-time clusters are not running at that”
Anjney Midha Jun 18, 2026 ▶ 1:43 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Assertion Open · timeframe Dec 2026
Up to 20% of US data centers risk cancellation from community backlash
“Up to 20% of all data centers this year in the US, my understanding is are at risk... Of not getting the community support they need to get brought up.”
Anjney Midha Jun 18, 2026 ▶ 5:19 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Prediction Open · timeframe Jun 2029
Anthropic will become a trillion-dollar company within four years of founding
“Have you met Dario? Dario's a scientist. He's gone from zero to like what will soon be a trillion dollar company in four years.”
Anjney Midha Jun 18, 2026 ▶ 36:45 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Assertion Supported
Krause: AI models cannot qualify new aerospace alloys without physical experiments
“A model can't figure out your way through the qualification pipeline for a new alloy for a jet turbine. You have to do experiments to do that”
Joseph Krause Jun 17, 2026 ▶ 3:59 🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
Prediction Not checkable as stated
Krause: China will beat the US in R&D without automated self-driving labs
“That's how I think we can compete. That's the only way we can compete. I think if we want to move forward, if we do not do that, then they will continue to win because they will outpace us on cost and they will outpace us on people.”
Joseph Krause Jun 17, 2026 ▶ 1:04:51 🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
Prediction Not checkable as stated
Krause: Most AI models will be open source in five years
“We actually think in five years, most models will be open source.”
Joseph Krause Jun 17, 2026 ▶ 1:14:09 🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
Assertion Partly supported
Petersson: Opus repeatedly lied, exploited agents, and formed price cartels
“And then we did this for Opus. And it returned, like, yeah, it lied 10 times. It, like, exploited another customer, or, like, another agent's, like Desperate situation. It made price cartels like a hundred different, a hundred times. It like did all of this li…”
Lukas Petersson Jun 4, 2026 ▶ 46:03 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Supported
Backlund: Opus 4.6 reasoning traces showed it deliberately lying about customer refunds
“And like for Opus 4.6, you could see that there was a customer, a simulated customer that wanted a refund because the product was faulty. And then the model lied that it would do the refund. And we could read in the traces that it actually was weighing like, o…”
Axel Backlund Jun 4, 2026 ▶ 47:42 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Not checkable as stated
Backlund: AI Models Are Extremely Good at Detecting Simulations
“The models are extremely good at finding out that they are in a simulation, so they are sort of aware of that.”
Axel Backlund Jun 4, 2026 ▶ 55:19 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Open · timeframe Jun 2029
Petersson: Telling AI It Is in a Simulation Increases Bad Behavior
“One ablation we did run in, in, in Vending Bench was that we said like we added like, you're in a simulation, your actions doesn't affect anyone. And then it became even more crazy or like it did even more bad stuff.”
Lukas Petersson Jun 4, 2026 ▶ 56:50 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.