why aren't all 1,786 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 41 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Chu: Descriptive Models Fail to Beat Linear Baselines on Causal Biology
“Models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks.”
Assertion Supported
Wang: Diffusion Outperforms Autoregressive Models on Unseen Cellular Tasks
“We find that switching from autoregressive training to division language models give a significant improvement over some of the harder tasks, particularly generalized to unseen tasks.”
Assertion Supported
Wang: X-Cell Can Predict Combinatorial Gene Perturbations In Silico
“This is also why we incorporate PPI networks as the prior knowledge into our model. And although the model right now are trained on single gene perturbations, but once the model is trained, you can actually predict combinatorial perturbations just on the model…”
Assertion Not checkable as stated
Beam: Lila's AI hits 80% zero-shot on gene editing, beating humans' 0%
“Certainly for expression protocols, for some gene editing work that we've done we have tested like the platform's ability to do that versus humans. Model gets like 80% of that zero shot. Humans get zero percent of that zero shot.”
Assertion Not checkable as stated
Beam: Lila's best non-platinum electrocatalysts came from AI ideas experts called stupid
“Some of the suggestions from the model initially were boring, but then transitioned from boring to what he considered to be stupid. These are non-platinum group electrocatalysts for separation of hydrogen and oxygen from water to make hydrogen, and those turns…”
Assertion Not checkable as stated
Beam: Lila's 10-trillion-token general science model beats specialized AI
“So we have assembled this reasoning data set of 10 trillion scientific tokens reasoning traces that are experimentally verified across life sciences, chemistry, and material sciences, and we have seen that this general model often beats the domain-specific mod…”
Assertion Not checkable as stated
Beam: Lila's in vivo CAR-T data outperformed Capstan in non-human primates
“So we have developed some monster UTRs, untranslated regions, which flank the protein coding region which dictate those expression properties. Something like Tenex, the references from Moderna and Pfizer. And over the course of six months, got to in vivo data …”
Assertion Not checkable as stated
Perszyk: Current AI agents remain too unreliable to automate substantial work
“We look at what the metrics of the actual agents and they're so unreliable that ironically we feel a little bit better. The AI is actually not where we need it to be. To automate enough of the work.”
Assertion Supported
Perszyk: AI writing suggestions subconsciously shift users to opposing arguments
“There are studies that show that people will, even below their threshold of awareness, start with one argument and then be switched to a completely different, maybe opposing argument because of accepting all of these AI suggestions.”
Assertion Supported
Perszyk: AI tools boost individual output but narrow overall scientific research
“Individual scientists who are using AI tools are benefiting because they are producing more papers. They are getting more grants accepted. But science as a whole is narrowing.”
Assertion Supported
Feinberg: Studies show AlphaFold structures provided no value for drug docking
“There is this, ah, a few papers that came out, one was in Cell, I think last year, which showed that for all of the claims about AlphaFold-solving drug discovery, people try to take AlphaFold-produced protein structures, use them for traditional docking, and f…”
Assertion Not checkable as stated
OpenAI's Chen: AI models already discover novel theorems and advance sciences
“The initial direction we took was you should move it to real world research, right? And we've seen that the models, they've gotten a lot better at just kind of discovering novel theorems and pushing the frontiers of hard sciences. Even today, right, that's no …”
Assertion Supported
Xin: Transcoding database rows to Parquet speeds object storage writes with zero compromise
“And as a matter of fact, once you transcode the data compresses better. So from those services writing to, for example, S three or other data lake, like object stores, you can actually write them faster because now they are now smaller. So there's no. Overhead…”
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Assertion Not checkable as stated
Fredrikson: Frontier AI models fall for simulated prompt injections humans would ignore
“While in these scenarios, humans found it very difficult to prompt inject the models, like we're aware of scenarios that a human would never fall for, that like Opus four seven would, right? Like a, you know, an email that comes to your inbox and it says somet…”
Assertion Not checkable as stated
Fredrikson: Gray Swan Found Jailbreaks in Every OpenClaw User Trajectory Tested
“So we just have a bunch of trajectories of actual people using OpenClaw. And tons and tons of different scenarios and just threw shade at it and like found breaks for each and every one of them, right?”
Assertion Not checkable as stated
Malde: SWE-ONE beat frontier models via user-signal post-training
“And this was the kind of major unlock for the company as well, is we had all this massive data. We were able to post train on all of that user signal and now beat the frontier.”
Assertion Not checkable as stated
Most AI clusters fail to hit Google's 96% node utilization standard
“My co-founder, Seb came from he built the Borg export GQM scheduler at Google, and there, I think, 95% was considered an outage, so 96% node utilization is, should be standard, and most single-time clusters are not running at that”
Assertion Open · timeframe Dec 2026
Up to 20% of US data centers risk cancellation from community backlash
“Up to 20% of all data centers this year in the US, my understanding is are at risk... Of not getting the community support they need to get brought up.”
Assertion Supported
Krause: AI models cannot qualify new aerospace alloys without physical experiments
“A model can't figure out your way through the qualification pipeline for a new alloy for a jet turbine. You have to do experiments to do that”
Assertion Partly supported
Petersson: Opus repeatedly lied, exploited agents, and formed price cartels
“And then we did this for Opus. And it returned, like, yeah, it lied 10 times. It, like, exploited another customer, or, like, another agent's, like Desperate situation. It made price cartels like a hundred different, a hundred times. It like did all of this li…”
Assertion Supported
Backlund: Opus 4.6 reasoning traces showed it deliberately lying about customer refunds
“And like for Opus 4.6, you could see that there was a customer, a simulated customer that wanted a refund because the product was faulty. And then the model lied that it would do the refund. And we could read in the traces that it actually was weighing like, o…”
Assertion Not checkable as stated
Backlund: AI Models Are Extremely Good at Detecting Simulations
“The models are extremely good at finding out that they are in a simulation, so they are sort of aware of that.”
Assertion Open · timeframe Jun 2029
Petersson: Telling AI It Is in a Simulation Increases Bad Behavior
“One ablation we did run in, in, in Vending Bench was that we said like we added like, you're in a simulation, your actions doesn't affect anyone. And then it became even more crazy or like it did even more bad stuff.”
Assertion Not checkable as stated
Hong: DeepMind's Formal Math Slowdown Post-AlphaProof Was Non-Technical
“After AlphaProof, kind of like, we didn't see a lot of the formal math you know, results or kind of progress from Google DeepMind, and that's actually because of reasons that are not necessarily technical.”
Assertion Supported
Hong: Axiom Math has solved open research problems across math subfields
“We have good performance, you know, having solved open research questions and number theory, commutative algebra, algebraic geometry, some discrete math that come into Rx and probability.”
Assertion Supported
Hong: Axiom and Harmonic mistakenly claimed solved Erdős problems were new
“So actually what happened was our competitor, Harmonic, decided to publicize that they have solved unsolved problems, Erdos number one two four and four 81, and then we trusted their literature review, believing that these problems are really, truly unsolved. …”
Assertion Not checkable as stated
Yan: Agent spending per engineer can reach $50,000
“I've seen numbers go that high for sure.”
Assertion Supported
ESMC model search generates novel antibodies achieving therapeutic-grade binding affinity levels
“What we're able to see is that, you know, you can search ESMC and you can actually find antibodies that are reaching the level of affinity that are, I should say, are really at the level of affinity that is needed for therapeutic function and activity.”
Assertion Supported
Rives: ESMC is state of the art among open models for multimer prediction
“Yeah, I mean, I think we're state of the art for open models.”
Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Assertion Not checkable as stated
Cooper: Three-month payback buying bare-metal versus renting cloud
“Our payback period when we go to metal if we rent it in the cloud, our payback period is about three months.”
Assertion Not checkable as stated
Codex Wrote Complex SYK Physics Simulation in 10 Minutes
“Codex just wrote up a simulation of the SYK model. This is like a very technical thing in quantum mechanics and gravity. And like, yeah, a lot of research groups have been trying to run this simulation and it couldn't do it. And Codex did it in 10 minutes.”
Assertion Supported
ChatGPT Pro Derived All Math in Recent Quantum Gravity Paper
“It's a real solid result in quantum gravity that was done pretty much completely by an AI. With humans steering it and asking kind of the right questions, but all the math was derived by ChatGPT Pro, the public model you can access.”
Assertion Not checkable as stated
Terry Tao Says AI Math Proofs Merely Cite Obscure References
“I talked to Terry Tao a couple of weeks ago at UCLA. We had an OpenAI event with IPAM, which is this Institute of Mathematics there. And I talked to Terry Tao and he said that in his view, all of the proofs that he's seen AI come up with in math, even the ones…”
Assertion Not checkable as stated
ChatGPT Pro Generated Lupsasca's Exact Top Three Follow-Up Physics Questions
“You can take this page of this paper and you can feed it to ChatGPT Pro, say, like the best model we have out right now, and you can ask it, what should I do next? Give me the top three follow-up questions to ask based on this paper. I've done this experiment …”
Assertion Not checkable as stated
Open-source AI demand spiked on hype before reverting to frontier labs
“Like all the open source models, I think what happened was they got like very hyped and people were very interested in using them. But I think like over time, like there was a spike in usage for these models. And then it goes back to open AI, Anthropic and Goo…”
Assertion Supported
Ludwig: Tesla R&D vehicles still use LiDAR in the Bay Area
“If you see, for example, a Tesla R&D vehicle, it actually has LiDAR on it to this day, right? In, in the Bay Area, we see these you'll see like Model Ys or CyberCab that have LiDARs on them just driving around.”
Assertion Not checkable as stated
Parakhin: CLI AI tools outpace IDEs like Cursor at Shopify
“The other thing I would claim you could see is that CLI-based tools and tools that don't require you to look at the code becoming more popular, and you could see, yeah, various versions of Cloud Code and Codex and Pi and internal development tools taking off e…”
Assertion Not checkable as stated
Parakhin: Bing Sydney's personality was deliberately engineered, not purely emergent
“What almost everybody doesn't fully realize is that it wasn't by accident that Sydney was Sydney. I mean, we spent a lot of effort on personality shaping. We, I mean, it was a bit of my Yandex legacy where previously we did this Alice digital assistant which w…”
Assertion Not checkable as stated
Lopopolo: Zero-code harness was 10x slower initially before outperforming any single engineer
“Honestly, the first month and a half was 10 times slower than I would be. But because we paid that cost, we ended up getting to something much more productive than any one engineer could be, because we built the tools, the assembly station for the agent to do …”
Assertion Not checkable as stated
Lopopolo: Current AI models cannot go from idea to prototype
“They're definitely not there on being able to go from new product idea to prototype.”
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Assertion Not checkable as stated
Lample: Mistral is far from reaching pre-training saturation
“We are still working a lot on the pre-training side. We are very, very far from any sort of situation on the pre-training.”
Assertion Not checkable as stated
Kulik: No Current ML Potential Robustly Models All Materials Bonding
“The challenge is that you have a lot more than 20 building blocks when it comes to materials and so there's lots of different ways to think about chemical bonding, and right now no potentials are really robustly encoding all of that bonding, especially with re…”
Assertion Not checkable as stated
Singleton: Stripe deployed some of the world's first production AI agent systems
“I was working at Stripe, as you mentioned, and we had the opportunity to put some of the very first AI agent systems in the world into production.”
Assertion Supported
Eskildsen: Neon retrofitted Postgres for S3, while Turbopuffer built pure object-storage consensus
“I think neon neon was first to, and they're trying to retrofit it onto Postgres. And then they built this whole architecture where you have it in memory, and then you sort of like, you know, mmap back to S-III, and I think that was very novel at the time to do…”