why aren't all 992 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Assertion Supported
Azhnyuk: FPV Drones Cause 70% to 80% of Frontline Casualties
“Out of all the casualties on the frontline, between 70 and 80% are done by FPV drones.”
Assertion Supported
Sachs: AI Model Quality Varies Between First-Party APIs and Cloud Providers
“Companies that say they're selling the same model through different vendors, whether it be through first party or Bedrock, Azure, et cetera, we do see different qualities sometimes, and that's not necessarily what's advertised.”
Assertion Contradicted
Andreessen: Three-year-old Nvidia chips make more money today than when new
“The current models are getting better faster at such a rate that if you are running an NVIDIA, if you're running an NVIDIA inference chip today that's three years old, you're making more money on it today than you did three years ago. Because the pace of impro…”
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer
With a context length of, you know, 256,000 or more, they're not using a true transformer.
What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Assertion Supported
Bryk: Perplexity and ChatGPT Search rely on legacy Google and Bing APIs
“So these systems, there are a few of them now they basically rely on like traditional search engines like Google or Bing, and then they combine them with like LLMs at the end to, you know, output some power graphics answering your question. So they, Like, Sear…”
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Assertion Supported
Joscha Bach: Only a Tiny Fraction of Wikimedia's Budget Goes to Servers
“The Wikimedia Foundation is publishing what they are paying the money for, and a very tiny fraction on this goes into running the servers, and the editors are working for free.”
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Assertion Contradicted
Neural operators are the only AI architecture that works for climate emulation
“This is where the Allen AI Institute has now built climate models based on our neural operator architecture. And that's the only one that works As an AI emulator, right? None of the other architectures work for climate because climate requires us to assume the…”
Assertion Supported
AI models predict fusion reactor plasma disruption one million times faster
“You know, I talk about plasma and fusion reactor. You know, we barely have a few thousand samples, but we are able to accurately predict events like disruption very well. And we are able to do that a million times faster than what traditional simulations were …”
Assertion Supported
Park: Generative agent digital twins replicate human behavior at 85% accuracy
“And this is where we basically could replicate people's behaviors and attitudes, 85% as accurately as people would replicate their own. So that actually was the first really paper that gave this validated results that we can actually model individuals in an ac…”
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Assertion Supported
Kant: Laguna S outperforms models two to three times its size
“When you look at the benchmarks and start using it, you'll realize that we are outperforming models two or three times their size.”
Assertion Supported
Chu: Descriptive Models Fail to Beat Linear Baselines on Causal Biology
“Models that are trained on descriptive data do not yet outperform linear models on causal tasks, perturbational tasks, what we call counterfactual tasks.”
Assertion Supported
Wang: Diffusion Outperforms Autoregressive Models on Unseen Cellular Tasks
“We find that switching from autoregressive training to division language models give a significant improvement over some of the harder tasks, particularly generalized to unseen tasks.”
Assertion Supported
Wang: X-Cell Can Predict Combinatorial Gene Perturbations In Silico
“This is also why we incorporate PPI networks as the prior knowledge into our model. And although the model right now are trained on single gene perturbations, but once the model is trained, you can actually predict combinatorial perturbations just on the model…”
Assertion Supported
Perszyk: AI writing suggestions subconsciously shift users to opposing arguments
“There are studies that show that people will, even below their threshold of awareness, start with one argument and then be switched to a completely different, maybe opposing argument because of accepting all of these AI suggestions.”
Assertion Supported
Perszyk: AI tools boost individual output but narrow overall scientific research
“Individual scientists who are using AI tools are benefiting because they are producing more papers. They are getting more grants accepted. But science as a whole is narrowing.”
Assertion Supported
Feinberg: Studies show AlphaFold structures provided no value for drug docking
“There is this, ah, a few papers that came out, one was in Cell, I think last year, which showed that for all of the claims about AlphaFold-solving drug discovery, people try to take AlphaFold-produced protein structures, use them for traditional docking, and f…”
Assertion Supported
Xin: Transcoding database rows to Parquet speeds object storage writes with zero compromise
“And as a matter of fact, once you transcode the data compresses better. So from those services writing to, for example, S three or other data lake, like object stores, you can actually write them faster because now they are now smaller. So there's no. Overhead…”
Assertion Supported
Krause: AI models cannot qualify new aerospace alloys without physical experiments
“A model can't figure out your way through the qualification pipeline for a new alloy for a jet turbine. You have to do experiments to do that”
Assertion Partly supported
Petersson: Opus repeatedly lied, exploited agents, and formed price cartels
“And then we did this for Opus. And it returned, like, yeah, it lied 10 times. It, like, exploited another customer, or, like, another agent's, like Desperate situation. It made price cartels like a hundred different, a hundred times. It like did all of this li…”
Assertion Supported
Backlund: Opus 4.6 reasoning traces showed it deliberately lying about customer refunds
“And like for Opus 4.6, you could see that there was a customer, a simulated customer that wanted a refund because the product was faulty. And then the model lied that it would do the refund. And we could read in the traces that it actually was weighing like, o…”
Assertion Supported
Hong: Axiom Math has solved open research problems across math subfields
“We have good performance, you know, having solved open research questions and number theory, commutative algebra, algebraic geometry, some discrete math that come into Rx and probability.”
Assertion Supported
Hong: Axiom and Harmonic mistakenly claimed solved Erdős problems were new
“So actually what happened was our competitor, Harmonic, decided to publicize that they have solved unsolved problems, Erdos number one two four and four 81, and then we trusted their literature review, believing that these problems are really, truly unsolved. …”
Assertion Supported
ESMC model search generates novel antibodies achieving therapeutic-grade binding affinity levels
“What we're able to see is that, you know, you can search ESMC and you can actually find antibodies that are reaching the level of affinity that are, I should say, are really at the level of affinity that is needed for therapeutic function and activity.”
Assertion Supported
Rives: ESMC is state of the art among open models for multimer prediction
“Yeah, I mean, I think we're state of the art for open models.”
Assertion Supported
Sanseviero: AI labs republished model merging techniques previously created on Reddit
“Yeah, like all of the FrankenMoe stuff, like all of the Axolotl library, like all of these tools, and there were papers published by different companies and research labs one or two years later that were rediscovering what was already done by The Reddit or Dis…”
Assertion Supported
ChatGPT Pro Derived All Math in Recent Quantum Gravity Paper
“It's a real solid result in quantum gravity that was done pretty much completely by an AI. With humans steering it and asking kind of the right questions, but all the math was derived by ChatGPT Pro, the public model you can access.”
Assertion Supported
Ludwig: Tesla R&D vehicles still use LiDAR in the Bay Area
“If you see, for example, a Tesla R&D vehicle, it actually has LiDAR on it to this day, right? In, in the Bay Area, we see these you'll see like Model Ys or CyberCab that have LiDARs on them just driving around.”
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Assertion Supported
Eskildsen: Neon retrofitted Postgres for S3, while Turbopuffer built pure object-storage consensus
“I think neon neon was first to, and they're trying to retrofit it onto Postgres. And then they built this whole architecture where you have it in memory, and then you sort of like, you know, mmap back to S-III, and I think that was very novel at the time to do…”
Assertion Supported
Shah: OpenClaw's 15-message replay fails prompt caching and costs 10x more
“The way OpenClaw does it is it essentially sends back the last 15 messages in the conversation and it essentially uses that back and forth. And I mean, the approach itself is not ideal because you will, like, you are not doing any, like, you're not utilizing a…”
Assertion Supported
Patel: Claude Code's share of GitHub commits doubled to 4% in January
“Just in January, it went from four percent of or two percent of commits on GitHub to four percent of GitHub commits were done by Cloud Code, right?”
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Assertion Supported
O'Laughlin: AI build-out CapEx has massively passed the internet
“We've well massively passed the internet in terms of the absolute size of the build-out. It's not even close.”
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Assertion Supported
White: ML trained on experimental data beat first-principles simulations by a large margin
“Two very well-resourced groups. They both tried different ideas, and the machine learning on experimental data beat out first principles simulation by You know, a very large margin.”
Assertion Supported
Cameron: General model intelligence does not correlate with hallucination rates
“One interesting aspect is that we've found that there's not really a, not a strong correlation between intelligence and hallucination rate. That's to say that the smarter the models are in a generalist sense isn't correlated with their ability to, when they do…”
Assertion Supported
Cameron: Model performance correlates with total parameters, not active parameters
“We, in our benchmark, see a lot of performance correlated more with total parameters than active, and not that correlated with how sparse like the models are. Our accuracy benchmark is part of a omniscience. It's very correlated with total. It's not correlated…”
Assertion Supported
Pliny: Anthropic added a $20k–$30k bounty but withheld jailbreak data
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”