why aren't all 41 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 41 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Open · timeframe Sep 2029
Anandkumar: Multi-physics models outperform single-physics models of equivalent parameter size
“And in fact, I was going to add that it turns out that having the model of the same size with multiple areas of physics does better than giving all of those parameters to each single physics. So if you had separate models and made them big enough as the origin…”
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Assertion Open · timeframe Dec 2026
Up to 20% of US data centers risk cancellation from community backlash
“Up to 20% of all data centers this year in the US, my understanding is are at risk... Of not getting the community support they need to get brought up.”
Assertion Open · timeframe Jun 2029
Petersson: Telling AI It Is in a Simulation Increases Bad Behavior
“One ablation we did run in, in, in Vending Bench was that we said like we added like, you're in a simulation, your actions doesn't affect anyone. And then it became even more crazy or like it did even more bad stuff.”
Assertion Open · timeframe Mar 2027
Eskildsen: Turbopuffer outperforms Lucene on long LLM search queries
“Turbo Puffer today has a fairly start of the state of the art full text search engine.
We beat Lucene on some queries, in particular, very long queries that we've optimized for, because those are the text search queries we see today.”
Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Assertion Not yet assessed · timeframe Jan 2026
White: Synthesis routes for dangerous compounds are already available on Wikipedia
“You can go find the synthesis route for many dangerous compounds on Wikipedia. People know what are the targets in the human body that, like, are targeted by most biological weapons. It's not really that much of a mystery.”
Assertion Open · timeframe Sep 2026
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”
Assertion Open · timeframe Sep 2028
Bachman: Power Retention models match original base model performance
“They'll come out with a nice shiny new, a power retention architecture that has the same performance on whatever data set they want as the original base model did.”
Assertion Open · timeframe Aug 2028
Morcos: Datology matches DCLM performance 12x faster with under 10% tokens
“We're able to now get to the same performance as DCLM about 12 x faster. So, you know, in fewer than 10% of the tokens we can match What you get from training to convergence.”
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Assertion Open · timeframe Aug 2026
Krentsel: Exo autonomously modified its code to inspect Pokémon game RAM
“We've had XO running, playing, playing Pokemon. And while it's running, the system itself decided to try inspecting the like RAM of the game and then went and mapped the RAM to, and people have reversed in the past, people have reverse engineered this manually…”
Assertion Open · timeframe Aug 2026
Krentsel: Exo autonomously re-architected its Discord adapter, cutting costs by 96%
“We asked it, Hey, I noticed, I asked, Hey, but how much did the last message cost in the discord adapter?
And it was like, it was.
It's like, are you serious?
16 cents.
That's actually crazy.
Like.
Go work on driving that down.
And so it went and re-architecte…”
Assertion Open · timeframe Jun 2029
Hong: Axiom's unmodified Putnam system achieved 99% on Verina benchmark
“And we actually recently, with no modification to the Putnam system, we saw a 99% out of the 189 problems, we saw a 187, we missed only two code-wisp-proof.”
Assertion Open · timeframe Jun 2029
Ethan He: Grok Imagine Video Extension Tracks Full Historical Context
“So the Glock Imagine video extension, it has historical context of all of the previous generated videos. It can it has a context of who is speaking and what objects have appeared and everything having that to generate the next video.”
Assertion Open · timeframe May 2027
Daytona spins up a single agent sandbox in 60ms
“And so our time to spin up one is 60 milliseconds with network agency. So requests, spin up, reply, 60, the whole thing, 60 milliseconds.”
Assertion Open · timeframe May 2027
Daytona can spin up 50,000 concurrent sandboxes in 75 seconds
“But if you want to spin up 50,000 at once, we are now at about 75 seconds. So it takes about 75 seconds to spin up concurrently 50,000.”
Assertion Open · timeframe Jan 2026
White: ChemCrow paper was presented to U.S. President in 30-minute block
“I ended up visiting the white house. I guess my paper was like the only time a preprint or peer review paper was presented to the president on like their schedule for like a 30 minute block.”
Assertion Open · timeframe Sep 2028
Bachman: StarCoder-3B converted to Power Retention matches baseline loss in two hours
“After just 10,000 steps of training, which this training one took about two hours, this orange curve, you see that it fully matches the original loss.”
Assertion Open · timeframe Aug 2026
Sohmers: Positron AI hardware achieves 93% of theoretical memory bandwidth
“And so our fundamental architecture is enabling us, you know, today with hardware that we're shipping right now to be achieving, you know, 93% of the theoretical memory bandwidth of our device consistently across all use cases.”
Assertion Open · timeframe May 2025
Huang: PoSE breaks down on needle-in-a-haystack at 500k tokens
“It does start to break down a little bit more on the longer, longer context. So, like, 500,000 to a million it appeared that it doesn't hold as well specifically for, like, needle in the haystack.”
Assertion Open · timeframe Sep 2026
Slack: Sourcegraph serves nine of top ten public tech companies
“We have like nine of the 10 top Public tech companies as customers and like four of the six top banks and like Uber and Stripe and so on, all these companies using Sourcegraph for code search.”
Assertion Open · timeframe Jun 2026
Fredrikson: Skilled red teamers phish human participants 60% to 70% of the time
“But for a skilled, like, human red teamer, they could fish the human participants, like, with the 60 to 70% success.”
Assertion Open · timeframe Jun 2026
Fredrikson: Top AI browser agents yielded only a handful of successful breaks
“There were a couple of models that seemed to be very, very robust, right? Like the red teamers found just a handful of successful breaks on them.”
Assertion Open · timeframe Jun 2029
Midha: MatX chips adopt NVIDIA reference architecture to plug into existing sites
“When they decided to pick the standard for their data center, they picked the NVIDIA reference architecture. So the Matex chips just plug in to any site that has an NVIDIA bring up planned. And you know.”
Assertion Open · timeframe May 2027
The United Nations uses Chatbase on Facebook Messenger for regional crisis support
“Like, the UN is using us right now on their Facebook messenger. So when people reach out to them, specifically for, like, specific regions, it gets drafted to Chatbase, and Chatbase helps them.”
Assertion Open · timeframe Apr 2029
Sun: Moonlake can generate multiplayer environments and persistence databases via prompting
“So if you just actually just like prompt our Model to say, hey, like configure the multiplayer, then it'll do like this. You'll be able to configure multiplayer. Persistency database for you.”
Assertion Open · timeframe Feb 2027
Qwen 3 memorizes benchmark questions significantly more than Qwen 1.5 or 2
“I don't, we don't see this phenomenon like the earlier versions of like QN 1.5 or even QN two, but start seeing it in QN three. So there's something about a more, like a significantly higher weight on benchmark or like a J or like any times of evaluation quest…”
Assertion Not yet assessed · timeframe Dec 2025
Altman, Nadella, and Pichai publicly committed to adopting MCP around April
“And then like, you had this like inflection point around April with like Sam Altman and Satya and Sundar and all posting about like MCP and that they're going to adopt MCP at Microsoft, at Google. At OpenAI and that was really like the big inflection point.”
Assertion Open · timeframe Nov 2028
Chan: CZI Billion Cell Project takes months at fraction of historical cost
“Now we're doing the billion cell project and that is taking months and at a fraction of the price.”
Assertion Open · timeframe Oct 2026
Merrill: AI models now reliably solve Terminal-Bench's ML training task
“Unfortunately we are getting to the point where models do reliably get this one.”
Assertion Open · timeframe Sep 2028
Bachman: PowerCoder-3B reaches 35% HumanEval accuracy versus StarCoder's 30%
“In the end, this converges to, I believe, about 35% accuracy on human eval, whereas the star coder baseline was about 30%.”
Assertion Open · timeframe Jun 2026
Vibhu: Gemma Activates Abstract Behavioral Traits Over Simple Token Completion
“It also shows internally that there's more than just token completion of, you know, this plus this equals this. No, it has some under understanding of characteristics, right? Like this is a pretty stubborn dog. It has a stubborn feature. Pretty high up that ac…”
Assertion Open · timeframe Apr 2028
GPT-4.1 Nano and Mini are new pre-trains; base 4.1 is mid-train
“Nano is obviously a new pre-train. We also have a new pre-train for Mini, and then, ah, the larger version is, ah, a new mid-train.”
Assertion Open · timeframe Aug 2029
Park: Simile AI agents can autonomously navigate live website URLs
“Some of the things that our agents can also do is you can be given a domain, like a website URL and actually go use it for a while.”
Assertion Open · timeframe Aug 2029
Midjourney's David Holtz explored text diffusion to storyboard entire movies
“David Holtz from Midjourney was investing in text diffusion. I don't think anything came out of it, but like the idea was that you can storyboard a long movie and then you can generate the scenes with video, normal video gen.”
Assertion Open · timeframe Jul 2026
Bubna: Ramp trained custom tokenizers to swap into LLaMA
“Ramp actually early in the day was training their own tokenizer and, like, Swapping out the tokenizer in Lama and whatnot.”
Assertion Open · timeframe Jun 2026
Krause: DARPA and GE Aerospace synthesized 500 alloys in 12 months
“The largest alloys program was the mock program. It was run by DARPA NGE Aerospace. They did 500 alloys in about 12 months. They did a bunch of kind of AI and simulations on the front end of that, and then they synthesized 500 new alloys in that whole year.”
Assertion Not yet assessed · timeframe Dec 2024
Swix: Bolt, Devin, and AI Agent Startups Rely on Netlify Deployments
“Both Bolt and Cognition DevIn and a bunch of other sort of agent type startups, they all use Nullify to deploy because of this one feature.”
Assertion Not yet assessed · timeframe Oct 2024
BFCL v3 generates evaluation tasks using graph edge construction
“They basically created their own API for the sake of this testing, and then did this like mapping to create a graph edge construction and like generate tasks through that.”
Assertion Open · timeframe Aug 2027
Carlini encodes 1.44 megabytes of data onto a single sheet of paper
“Yeah, okay. So it's about, in particular, it's about 1.44 megabytes.”