why aren't all 851 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Rumbelow: Leap Labs' Discovery Engine Automates Novel Scientific Discovery via Interpretability
“Discovery Engine is an end-to-end system, takes in arbitrary scientific data set, automatically trains a bunch of neural networks on it, and then We systematically, with our interpretability methods, which is the real secret extract the patterns that have been…”
Assertion Supported
Corbitt: OpenPipe Beat Frontier Models Using a Qwen 32B Judge
“One of the results we published was we used Quen 2.5 14 B as the model we're training, and as the judge we used Quen 2.5 32 B, which is, like, Not, I mean, it's fine, but it's like not a, it's much worse than any frontier model. Right. And even with that combi…”
Assertion Supported
Feldman: Cerebras provides 2,625x more memory bandwidth than traditional GPUs
“And we have 2625 times more memory bandwidth than the GPU does.”
Assertion Supported
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Prediction Held up
Taskaya: Training a state-of-the-art image model costs under $1M
“Like right now, like if you look, if you want to train a Sota image model, I don't think it's going to cost more than a million dollars. It's extremely cheap. It's like a matter of data engineering effort, cleaning. It's, I think it's a function of data set.”
Assertion Supported
Morcos: Soft inductive biases become harmful past 1M data points in vision
“Turns out in the small data regime, and when I say small data here, I mean, say less than 500,000 data points. And this was in the context of image self-supervised learning. So in that small data regime, this is super helpful. And where this paper's actually b…”
Assertion Supported
Morcos: Kaplan and Chinchilla scaling laws incorrectly assume all data is equal
“And even if you go and you look at the scaling laws work from Kaplan and Chinchilla and all these other things, they all assume IID data which is insane. We know that all data are not created equal, that garbage in garbage out is like the oldest adage in compu…”
Assertion Supported
Morcos: Proper data curation can bend neural scaling laws
“And what that paper showed was that if you use your data correctly, you can actually bend the scaling laws themselves.”
Prediction Held up
Morcos: Training a specialized frontier model will cost under $1M very soon
“I believe that getting to a frontier model should cost a million dollars or less for most organizations, at least in a specialized domain, right?
And when you think about what enterprises need, that's generally what they need.
They don't need a model that can …”
Assertion Supported
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Assertion Supported
Palazzolo: Claude Code leads stayed at Cursor only two weeks
“We know that they went there, they were there for, I think, about two weeks, and they came back.”
Assertion Supported
Ermon: Diffusion LLMs Pareto-dominate autoregressive models on inference efficiency
“On the inference side, what we're seeing is that diffusion models are much more efficient. We're actually able to Pareto dominate autoregressive models. If you think about the typical trade-off between throughput versus latency, which you kind of like cannot, …”
Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Assertion Supported
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Assertion Supported
McCloy: ChatGPT Does Not Index or Retrieve llms.txt by Default
“I knew that there's debate about this, but I'd say the evidence is like ChatTriPT is not indexing and it's not retrieving content from LMS.txt by default.”
Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Assertion Supported
Cherny: Anthropic is currently bordering on AI Safety Level 3 capabilities
“Yeah, we're kind of bordering on three right now.”
Assertion Supported
Factorio benchmark results show reasoning models underperform expectations on extended planning
“One thing we have found in preliminary results is that the reasoning models don't seem to do as well as you'd expect in this setting. And I think that's probably because the way we set this up, it's a bit like we're already making it do reasoning traces over a…”
Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Assertion Supported
Cursor Composer solves Convex benchmarks but fails on alternative backends
“We did notice that I mean, with convex, it pretty much autonomously just solves the first two tasks. It has a few round trips on like some errors that are only show up and playing with the front end. And then it's able to complete this files task and kind of g…”
Assertion Supported
AI models struggle debugging Supabase RLS recursion compared to procedural code
“The particular example was like RLS rules and Supabase where debugging like an infinite loop for infinite recursion for the RLS rules was something that the models just really struggled with in a way that we didn't see for procedural code.”
Prediction Held up
Roucher: AI agents will reach a 90% GAIA score by 2026
“So I think if we solve Gaia, that's like 90% score. That means mostly we double productivity of every task done in front of a computer. And if you take the trend line of the scores so far this should be crossed in 2026 or something.”
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Assertion Supported
Zhang: XGrammar outperforms Outlines and is integrated into TensorRT-LLM
“And I think Xgrammar's performance is better than the outline's, and also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrates Xgrammar as the backend for the constructed coding.”
Assertion Supported
Neubig: SWE-bench Scores Are Inflated By Training Data Contamination
“Sweebench is on popular open source repos and all of these popular open source repos were included in the training data for all of the language models. And so, the language models already know these repos. In some cases, the language models already know the in…”
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Assertion Supported
Reddy: Chai Discovery's open-source Chai-1 model outperforms Google's AlphaFold 3
“We're lucky to work with the folks at Chai Discovery who just released Chai One, which is open source model that outperforms Alpha Fold Three.”
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Assertion Supported
Friedman: AlphaCodium boosts OpenAI o1, proving o1 lacks true System 2
“We took their all one preview with Alpha Codium and did better. Like it just shows like, and there is a big difference between the preview and the IOI. It shows, like, that these models are not still system two thinkers, and there's a big difference.”
Assertion Supported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Assertion Supported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Prediction Held up
Altman: 10-million-token fast context windows are coming within months
“Even getting to the, like, Ten million tokens of very fast and accurate context, which I expect to measure in, like, months, something like that.”
Assertion Supported
Schulhoff: LLMs Rely More on Prompt Structure Than Exemplar Labels
“There are a number of papers which have found that the label of the exemplar doesn't really matter, and the model reads the exemplars and cares more about structure than label.”
Assertion Supported
Jamil: Writing in the Margins solves the lost-in-the-middle problem
“It improves the ability of any language model to extract relevant information, so solving the lost in the middle problem”
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Assertion Supported
Tay: Zero-shot benchmark scores at 1B model scale are random chance
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance …”
Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Assertion Supported
Albrecht: AgentBench paper's appendix examples are actually incorrect solutions
“Like we were looking at the agent bench paper, I think just last week for our paper club. And one of the things that we noticed is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not…”
Prediction Held up
Bach: Smaller, more powerful models will ensure unconstrained AI remains accessible
“Yes, but there will also be better jailbroken models or models that have never been jailed before, because we find out how to make smaller models that are more powerful.”
Assertion Supported
Murphy: Five-minute voice calls cost 6.5 cents on Deepgram versus ElevenLabs.
“And then on the text-to-speech side, and doing something like this with an 11 labs would be about maybe a dollar 20. And just to give you an idea of comparison. So you can do a five minute call here for about six and a half cents.”