why aren't all 2,445 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 100 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Prediction Not checkable as stated
AI labs will increasingly co-design models alongside proprietary custom inference ASICs
“I definitely think things are moving towards you design a model. It's sized the way that fits well on hardware that you can design. You immediately start creating an inference hardware that is custom made to fit the sizes of your model. And then you can also s…”
Prediction Not checkable as stated
Companies pushing AI code volume metrics will drown in unmaintainable slop
“I think actually there's a lot of big tech examples where they're kind of pushing really hard for their teams to use more code. They're being evaluated on how much code they're using. And they're just kind of getting more and more slop that nobody understands.…”
Assertion Supported
Rumbelow: Leap Labs' Discovery Engine Automates Novel Scientific Discovery via Interpretability
“Discovery Engine is an end-to-end system, takes in arbitrary scientific data set, automatically trains a bunch of neural networks on it, and then We systematically, with our interpretability methods, which is the real secret extract the patterns that have been…”
Assertion Not checkable as stated
Rumbelow: arXiv Is Filling With Plausible but Unverified AI Papers
“I'm actually really worried about this because I think we're already seeing archive and other online repositories and... Submissions too, full of these very, very plausible papers. That may or may not be true. And like at that point, what good is the, is our s…”
Prediction Not checkable as stated
Future app platforms must extract auth and authorization completely
“Auth cannot be part of the app, because they're not going to get that right, right? So, Auth has to be extracted from the app. In fact, which data you can see, they also cannot be under control of the app, because again, you're going to get it wrong, right? So…”
Prediction Not checkable as stated
Sands: AI Agents Will Make Fraud Decisions Within Six Months
“Now you can think about, okay, actually foundation model, text alignment, like human readable description of like why we're worried about this charge. And then today, a human tomorrow, an agent sitting on top of that and decisioning, like reasoning over The mo…”
Prediction Not checkable as stated
Sands: Static dashboards will be obsolete within nine months as agents take over
“I think the value of near real time, high quality, well documented data is about to skyrocket because I'm pretty sure that nine months from now, no one is going to want to go and like look at a even like static dashboard and click around. They're going to want…”
Assertion Not checkable as stated
Sands: Top 100 AI startups reach ARR milestones 2-3x faster than SaaS
“One cohort that we looked at was the hundred highest grossing AI companies on Stripe. And you kind of need a reference point. And so we were like, let's compare them to the hundred highest grossing SaaS companies from five years prior. And we looked at things …”
Assertion Not checkable as stated
Sands: Vertical AI wrapper companies are building healthy unit economics
“Increasingly what we're seeing from the AI companies on Stripe is they do want to have healthy Unidec. I mean, let's not talk about like the big labs that are pouring crazy money into research, but if you're talking about like the vertical kind of wrappers, wh…”
Prediction Not checkable as stated
Sands: Most businesses target AI employee efficiencies for 2027 and 2028
“I don't think it's going to show up Next year. I don't think most businesses are targeting employee efficiencies next year, but I think every business is targeting employee efficiencies for 27 and for 28, which is suggesting more efficiency.”
Prediction Not checkable as stated
Webster: Meaningful AI red teaming will require internal tracing and observability
“I think especially where, where things are headed, like with more complex rags and agents and so forth, you're going to have to have some type of observability or like internal tracing in order to have, to do meaningful automated red teaming.”
Prediction Not checkable as stated
Swix: Frontier model sizes have likely peaked around 2 trillion parameters
“I wonder if we've hit the peak big model craze because now I do expect, you know, 10 trillion model releases, you know, 100 trillion model releases. Probably not. I think we might have peaked at two.”
Assertion Not checkable as stated
Bakouch: Novel optimizer speedups are exaggerated due to undertuned AdamW baselines
“And what they find is that the speed up is greatly, greatly exaggerated. And mostly because often people like undertone the Adam W baseline.”
Assertion Contradicted
Swix: Every frontier lab now distills dense models into MoEs
“I think like, I think this is the pattern for every frontier lab now.”
Assertion Not checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Prediction Not checkable as stated
Merrill: AI evals will shift to observing real jobs over 2-3 years
“I think like the future of evals does look much more like this is like observing people who are doing their real jobs and then translating those real jobs into a format that allows you to evaluate language models and harnesses on them. And it's probably going …”
Prediction Not checkable as stated
Merrill: Frontier AI labs will center operations around vertical products
“And with the Claude codes and the codec CLIs and the deep researchers, researchers, you starting to see some evidence that the products are going to be a much more central part of how these frontier labs operate.”
Prediction Not checkable as stated
Corbitt: 55-60% chance RL becomes the standard pattern for deploying scale agents
“I think that the chances that like everyone should be, or, you know, everyone who's deploying an agent at scale should be doing RL with it, either as part of sort of like a, you know, like pre-deployment or even like continuously as it's deployed, that that's …”
Assertion Supported
Corbitt: OpenPipe Beat Frontier Models Using a Qwen 32B Judge
“One of the results we published was we used Quen 2.5 14 B as the model we're training, and as the judge we used Quen 2.5 32 B, which is, like, Not, I mean, it's fine, but it's like not a, it's much worse than any frontier model. Right. And even with that combi…”
Assertion Not checkable as stated
Lenz: Local smartphone AI requires hybrid models due to KV cache limits
“So if you wanted to do something local on your phone to search your images, as an example, you can't do that without a hybrid architecture or without doing drastically changes because the model plus KVCache won't fit.”
Prediction Not checkable as stated
Agarwal: AI coding assistants will cause uninterpretable outages and endless firefighting
“And so it's pretty clear to us that, you know, and we were starting to use it ourselves and sometimes we didn't understand what the code was doing, but you know, we shipped it. And so it's like, well, if this is clearly, this is going to happen a lot more. And…”
Prediction Not checkable as stated
Dwivedi: AI self-healing for complex incidents is 6-12 months away
“Now for, then there is this level of 30 to 40% of the incidents or issues where you need to involve you know, a senior engineer for sanity checking. I think that healing will appear in, I don't know, six months to a year that will be comfortably, the technolog…”
Assertion Not checkable as stated
Howard: Answer.AI runs fully in-house stack without AWS or Google Cloud
“This group of, which has averaged about 10 to 12 people, currently nine, I think, have built a pretty Transformational and complex piece of software, which we can do a quick demo of later if you're interested. Using a complete web application development platf…”
Prediction Not checkable as stated
Field: Prompting will be remembered as the MS-DOS era of AI
“I think we'll look back on this era as like the MS-DOS era of AI, and the prompting and natural language that everyone's doing today, I think is just sort of like the start of how we're going to create interfaces to explore it in space.”
Prediction Not checkable as stated
Field: AI code generation will force developers to rely on visual abstractions
“I also think that it's going to be something that as we move forward in time with more Asians writing more parts of your code base, you will also be less familiar with the code. And so then you might want a different abstraction where you're able to work on th…”
Prediction Not checkable as stated
Field: Autonomous AI agents building complex software like Figma is a long way off
“I'm not saying, okay, go build Figma, and you, agent, are just gonna go figure out all the complexities of Figma. I think that's just not something I see happening in any near term future, even as longer range running agents start to occur and we've got better…”
Assertion Supported
Feldman: Cerebras provides 2,625x more memory bandwidth than traditional GPUs
“And we have 2625 times more memory bandwidth than the GPU does.”
Assertion Supported
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Assertion Not checkable as stated
Feldman: AI startups are replacing closed-source models with fine-tuned open-source
“I think for sort of AI companies like Cognition, like all your competitors, like AlphaSense, like dozens of others, they are trying to replace closed source models with very, very fast open source models, and they're trying to drive the open source Accuracy dr…”
Assertion Not checkable as stated
Slack: No AI developer tool has retained developers past 12 months
“There's not been any tool that has stuck with devs for more than, Six or 12 months or something.”
Prediction Not checkable as stated
Slack: Asynchronous background agents will dominate AI inference and output
“That's going to be blown up with async agents when they're running 24, seven concurrently in the background, then you can have 10 or a hundred times as many, and that's going to dominate inference. That's going to dominate the output you get.”
Prediction Not checkable as stated
Ball: Foundation models will become background implementation details in AI tools
“So I think we're going towards a future where the model will become an implementation detail to some sense, and we will end up on a different abstraction layer.”
Assertion Not checkable as stated
Fanelli: Cursor Default Model Switch Cost Anthropic $200M in ARR
“When cursors switch from Sonnet to GPT-V as like the default model that was like, you know, Two hundred million our revenue for Anthropic that kind of went away and like moved on to GPT-V.”
Assertion Not checkable as stated
Slack: Competitors discounted products up to 100% to win deals against Amp
“So we've had one head to head loss with Amp where we lost against the usual players. And the reason why is one of them discounted their other product a hundred percent for two years. The other one discounted at 85% for two years, which is just crazy.”
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Assertion Open · timeframe Sep 2026
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”
Assertion Open · timeframe Sep 2028
Bachman: Power Retention models match original base model performance
“They'll come out with a nice shiny new, a power retention architecture that has the same performance on whatever data set they want as the original base model did.”
Assertion Not checkable as stated
Taskaya: NSFW content makes up less than 1% of Fal traffic
“Moderation is optional to a level where, like, illegal content is moderated, and we also track, like, the non-illegal content NSFW moderation, and, like, we haven't seen, like, we haven't seen more than one percent.”
Prediction Not checkable as stated
Taskaya: 80% of promotional video content will be AI-generated within 12 months
“Like 12 months, I think like 80% of this is going to be generated.”
Prediction Held up
Taskaya: Training a state-of-the-art image model costs under $1M
“Like right now, like if you look, if you want to train a Sota image model, I don't think it's going to cost more than a million dollars. It's extremely cheap. It's like a matter of data engineering effort, cleaning. It's, I think it's a function of data set.”
Assertion Supported
Morcos: Soft inductive biases become harmful past 1M data points in vision
“Turns out in the small data regime, and when I say small data here, I mean, say less than 500,000 data points. And this was in the context of image self-supervised learning. So in that small data regime, this is super helpful. And where this paper's actually b…”
Assertion Supported
Morcos: Kaplan and Chinchilla scaling laws incorrectly assume all data is equal
“And even if you go and you look at the scaling laws work from Kaplan and Chinchilla and all these other things, they all assume IID data which is insane. We know that all data are not created equal, that garbage in garbage out is like the oldest adage in compu…”
Assertion Contradicted
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Assertion Supported
Morcos: Proper data curation can bend neural scaling laws
“And what that paper showed was that if you use your data correctly, you can actually bend the scaling laws themselves.”
Assertion Open · timeframe Aug 2028
Morcos: Datology matches DCLM performance 12x faster with under 10% tokens
“We're able to now get to the same performance as DCLM about 12 x faster. So, you know, in fewer than 10% of the tokens we can match What you get from training to convergence.”
Prediction Not checkable as stated
Morcos: Data curation still has at least 100x in performance gains ahead
“You know, we've already been able to get 10 X gains. I think there's at least another hundred X behind this that are still to be done.”
Prediction Held up
Morcos: Training a specialized frontier model will cost under $1M very soon
“I believe that getting to a frontier model should cost a million dollars or less for most organizations, at least in a specialized domain, right?
And when you think about what enterprises need, that's generally what they need.
They don't need a model that can …”