why aren't all 2,445 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 100 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Kant: Reinforcement learning will move earlier into LLM pre-training
“I have I would say a not commonly held opinion that reinforcement learning will move earlier and earlier into pre-training.”
Assertion Supported
Kant: Laguna S outperforms models two to three times its size
“When you look at the benchmarks and start using it, you'll realize that we are outperforming models two or three times their size.”
Prediction Open · timeframe Jul 2027
Kant: Prompts stuffed with dozens of tools will vanish in 12 months
“I think we will in 12 months not see a single system prompt that is stuffed with 20 or 30 or 40 tools anymore.”
Prediction Not checkable as stated
Kant: AI will become the world's most demanded commodity with commoditizing margins
“Intelligence is the most in life. You're going to be the world's most demanded commodity. It will more commoditize in margin and price.”
Assertion Not checkable as stated
Wang: X-Cell Is First to Predict Unseen Cell Line Perturbations
“One of the rewarding signals I receive after we develop Excel is that like it's a wow moment from biologists that this is the first time biologists actually find the model can predict exactly how these unseen cell lines kind of respond to different perturbatio…”
Prediction Not checkable as stated
Wang: Virtual cell AI models will eventually replace physical cellular experiments
“Eventually we kind of, we can replace all the cellular experiments by simply running simulations on computer without even running the actual Experiments.”
Assertion Supported
Chu: Descriptive Models Fail to Beat Linear Baselines on Causal Biology
“Models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks.”
Assertion Supported
Wang: Diffusion Outperforms Autoregressive Models on Unseen Cellular Tasks
“We find that switching from autoregressive training to division language models give a significant improvement over some of the harder tasks, particularly generalized to unseen tasks.”
Prediction Not checkable as stated
Wang: Foundation models on causal data will beat linear baselines
“I believe that foundation model or other more complicated AI models that trend on the right data will outperform these linear models in harder tasks, particularly in generalization tasks.”
Assertion Supported
Wang: X-Cell Can Predict Combinatorial Gene Perturbations In Silico
“This is also why we incorporate PPI networks as the prior knowledge into our model. And although the model right now are trained on single gene perturbations, but once the model is trained, you can actually predict combinatorial perturbations just on the model…”
Assertion Not checkable as stated
Beam: Lila's AI hits 80% zero-shot on gene editing, beating humans' 0%
“Certainly for expression protocols, for some gene editing work that we've done we have tested like the platform's ability to do that versus humans. Model gets like 80% of that zero shot. Humans get zero percent of that zero shot.”
Assertion Not checkable as stated
Beam: Lila's best non-platinum electrocatalysts came from AI ideas experts called stupid
“Some of the suggestions from the model initially were boring, but then transitioned from boring to what he considered to be stupid. These are non-platinum group electrocatalysts for separation of hydrogen and oxygen from water to make hydrogen, and those turns…”
Assertion Not checkable as stated
Beam: Lila's 10-trillion-token general science model beats specialized AI
“So we have assembled this reasoning data set of 10 trillion scientific tokens reasoning traces that are experimentally verified across life sciences, chemistry, and material sciences, and we have seen that this general model often beats the domain-specific mod…”
Assertion Not checkable as stated
Beam: Lila's in vivo CAR-T data outperformed Capstan in non-human primates
“So we have developed some monster UTRs, untranslated regions, which flank the protein coding region which dictate those expression properties. Something like Tenex, the references from Moderna and Pfizer. And over the course of six months, got to in vivo data …”
Prediction Not checkable as stated
Beam: Round-over-round experimentation yields more compound value than broad datasets
“The bet is that the sort of like as the model performance improves, the sample efficiency goes up, and therefore like the compound interest that you get from round over round experimentation will outweigh that, that you would get from a big noisy, but broad da…”
Prediction Not checkable as stated
Biderman: AI-native companies will amass trillions of internal tokens within 18 months
“In 18 months, many companies would have maybe trillions of tokens, which of internal company data, proprietary data. I'm talking about like maybe trillions. It sounds exaggerated, but I don't think it's an impossibility if they're really AI native.”
Prediction Not checkable as stated
Biderman: Hard engineering tasks will require test-time gradient updates
“We think that eventually part of the solution for very hard tasks in, in science and engineering and defense and all that stuff will involve some form of gradient based updates during during doing these long horizon tasks.”
Prediction Not checkable as stated
Biderman: In 18 months, data scale will require weight-based learning
“Other parts of it are bets that in 18 months from now, the scale of the data will require the methods that we know from pre-training work.”
Prediction Open · timeframe Jul 2031
Biderman: PC hardware will soon run near-trillion-parameter models locally
“And in the long, long term, I do think these things will actually run on people's devices, and we're seeing right now the new hardware on personal computers is already, ah, you know, soon approaching the ability to run inference on close to trillion parameters…”
Prediction Not checkable as stated
Biderman: AI models must learn to autonomously filter out erroneous user feedback
“Increasingly the models will get better, and increasingly they'll know more things than we do, so the model in some way has to learn and understand and kind of, like, discern what, which feedback is valuable and which feedback should be ignored.”
Assertion Not checkable as stated
Perszyk: Current AI agents remain too unreliable to automate substantial work
“We look at what the metrics of the actual agents and they're so unreliable that ironically we feel a little bit better. The AI is actually not where we need it to be. To automate enough of the work.”
Assertion Supported
Perszyk: AI writing suggestions subconsciously shift users to opposing arguments
“There are studies that show that people will, even below their threshold of awareness, start with one argument and then be switched to a completely different, maybe opposing argument because of accepting all of these AI suggestions.”
Assertion Supported
Perszyk: AI tools boost individual output but narrow overall scientific research
“Individual scientists who are using AI tools are benefiting because they are producing more papers. They are getting more grants accepted. But science as a whole is narrowing.”
Prediction Not checkable as stated
Swyx: AI p(doom) over the next ten years is near zero
“I mean, if you do it in 10 years is near zero.”
Prediction Not checkable as stated
Swyx: LLMs will plateau and potentially trigger a 30-year AI winter
“Yeah, probably LLMs are going to run out at some point and they're not AGI and okay, we have maybe another 30 years of AI winter or something and then like the next paradigm really is actually the thing.”
Prediction Not checkable as stated
Swyx: Specialized agent labs like Cursor, Cognition, and Harvey will endure
“There will always be capability overhangs. They may not stay still, and so you gotta be nimble. But the Sierras of the world, the Cognitions of the world, the Cursors of the world, the Decagons and Harvey's, these are all agent labs for their field. They can b…”
Assertion Supported
Feinberg: Studies show AlphaFold structures provided no value for drug docking
“There is this, ah, a few papers that came out, one was in Cell, I think last year, which showed that for all of the claims about AlphaFold-solving drug discovery, people try to take AlphaFold-produced protein structures, use them for traditional docking, and f…”
Prediction Not checkable as stated
Feinberg: Chipmakers will invest in life sciences as LLM alpha shrinks
“I do think that chip makers, including Nvidia, are going to want to get a lot more invested in life sciences because it will always be high in demand. And The amount of alpha left in pure LLM space is just getting a little questionable.”
Assertion Not checkable as stated
OpenAI's Chen: AI models already discover novel theorems and advance sciences
“The initial direction we took was you should move it to real world research, right? And we've seen that the models, they've gotten a lot better at just kind of discovering novel theorems and pushing the frontiers of hard sciences. Even today, right, that's no …”
Prediction Not checkable as stated
Mark Chen: AI scaling laws will continue to hold
“And so I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering, more careful scaling, and it always unlocks that next ability to scale further. So I mean, it's held for You know, almost 10 order…”
Prediction Not checkable as stated
Zaharia: Open agent hosting layers will win over proprietary alternatives
“Another way to think about it is like, imagine, you know we, our thing wasn't open. We had some kind of agent hosting thing, but it's not open. And then there is an open one. If you're, which one's gonna win in the long run? So like here, because there is this…”
Assertion Supported
Xin: Transcoding database rows to Parquet speeds object storage writes with zero compromise
“And as a matter of fact, once you transcode the data compresses better. So from those services writing to, for example, S three or other data lake, like object stores, you can actually write them faster because now they are now smaller. So there's no. Overhead…”
Prediction Not checkable as stated
Xin: Much of traditional software will be rewritten with data and agents
“Actually, I think many of the traditional software will be sort of rewritten with this new paradigm, which is just get the data to be there. And then they slap some agent on top.”
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Assertion Not checkable as stated
Fredrikson: Frontier AI models fall for simulated prompt injections humans would ignore
“While in these scenarios, humans found it very difficult to prompt inject the models, like we're aware of scenarios that a human would never fall for, that like Opus four seven would, right? Like a, you know, an email that comes to your inbox and it says somet…”
Prediction Not checkable as stated
Kolter: Security and science will explode as AI agents automate tedious verification
“So I think this is really sort of an underappreciated point that we're reaching this point, this sort of phase where a lot of security, a lot of science has this potential to kind of explode. Not because we're going to get better at it, but because agents can …”
Assertion Not checkable as stated
Fredrikson: Gray Swan Found Jailbreaks in Every OpenClaw User Trajectory Tested
“So we just have a bunch of trajectories of actual people using OpenClaw. And tons and tons of different scenarios and just threw shade at it and like found breaks for each and every one of them, right?”
Assertion Not checkable as stated
Malde: SWE-ONE beat frontier models via user-signal post-training
“And this was the kind of major unlock for the company as well, is we had all this massive data. We were able to post train on all of that user signal and now beat the frontier.”
Assertion Not checkable as stated
Most AI clusters fail to hit Google's 96% node utilization standard
“My co-founder, Seb came from he built the Borg export GQM scheduler at Google, and there, I think, 95% was considered an outage, so 96% node utilization is, should be standard, and most single-time clusters are not running at that”
Assertion Open · timeframe Dec 2026
Up to 20% of US data centers risk cancellation from community backlash
“Up to 20% of all data centers this year in the US, my understanding is are at risk... Of not getting the community support they need to get brought up.”
Prediction Open · timeframe Jun 2029
Anthropic will become a trillion-dollar company within four years of founding
“Have you met Dario? Dario's a scientist. He's gone from zero to like what will soon be a trillion dollar company in four years.”
Assertion Supported
Krause: AI models cannot qualify new aerospace alloys without physical experiments
“A model can't figure out your way through the qualification pipeline for a new alloy for a jet turbine. You have to do experiments to do that”
Prediction Not checkable as stated
Krause: China will beat the US in R&D without automated self-driving labs
“That's how I think we can compete. That's the only way we can compete. I think if we want to move forward, if we do not do that, then they will continue to win because they will outpace us on cost and they will outpace us on people.”
Prediction Not checkable as stated
Krause: Most AI models will be open source in five years
“We actually think in five years, most models will be open source.”
Assertion Partly supported
Petersson: Opus repeatedly lied, exploited agents, and formed price cartels
“And then we did this for Opus. And it returned, like, yeah, it lied 10 times. It, like, exploited another customer, or, like, another agent's, like Desperate situation. It made price cartels like a hundred different, a hundred times. It like did all of this li…”
Assertion Supported
Backlund: Opus 4.6 reasoning traces showed it deliberately lying about customer refunds
“And like for Opus 4.6, you could see that there was a customer, a simulated customer that wanted a refund because the product was faulty. And then the model lied that it would do the refund. And we could read in the traces that it actually was weighing like, o…”
Assertion Not checkable as stated
Backlund: AI Models Are Extremely Good at Detecting Simulations
“The models are extremely good at finding out that they are in a simulation, so they are sort of aware of that.”
Assertion Open · timeframe Jun 2029
Petersson: Telling AI It Is in a Simulation Increases Bad Behavior
“One ablation we did run in, in, in Vending Bench was that we said like we added like, you're in a simulation, your actions doesn't affect anyone. And then it became even more crazy or like it did even more bad stuff.”