why aren't all 1,786 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 41 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Open · timeframe Mar 2027
Eskildsen: Turbopuffer outperforms Lucene on long LLM search queries
“Turbo Puffer today has a fairly start of the state of the art full text search engine.
We beat Lucene on some queries, in particular, very long queries that we've optimized for, because those are the text search queries we see today.”
Assertion Supported
Shah: OpenClaw's 15-message replay fails prompt caching and costs 10x more
“The way OpenClaw does it is it essentially sends back the last 15 messages in the conversation and it essentially uses that back and forth. And I mean, the approach itself is not ideal because you will, like, you are not doing any, like, you're not utilizing a…”
Assertion Supported
Patel: Claude Code's share of GitHub commits doubled to 4% in January
“Just in January, it went from four percent of or two percent of commits on GitHub to four percent of GitHub commits were done by Cloud Code, right?”
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Assertion Not checkable as stated
O'Laughlin: Compiled a PhD-level chip cycle dataset in one day using AI
“I mean, this is like too much information to gather. It's like a lifetime of work. It's like a PhD project. I did it in a day.”
Assertion Supported
O'Laughlin: AI build-out CapEx has massively passed the internet
“We've well massively passed the internet in terms of the absolute size of the build-out. It's not even close.”
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Assertion Not checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Assertion Not checkable as stated
Casado: There is no compute supply overhang or 'dark GPUs'
“But we don't have a supply overhang. Like, there's no dark GPUs, right?”
Assertion Not checkable as stated
Wang: L5 AI engineers can get offers in tens of millions
“You could be an L five and get an offer in the tens of millions.”
Assertion Not checkable as stated
Casado: Frontier Model Labs Are Gross Margin Negative Factoring Next-Gen Training
“If you look at the numbers of these companies, if you look at like the amount they're making and how much they spent training the last model, their gross margin positive, you're like, oh, that's really working. But if you look at like the current training that…”
Assertion Not checkable as stated
Casado: Cursor built a near-SOTA coding model at 1/100th the cost
“So the interesting thing about cursors, they actually for, you know, a small fraction of the cost, a 100th the cost or less, developed an almost soda model, which for a period of time was the most popular coding model in the world, right?”
Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Assertion Not checkable as stated
Labs train LLMs on benchmark questions for multiple epochs late in training
“It's very clear how the last stage of training for many of these models does have a massive amount of example or examine. Benchmaxing. Because the model will not, like, behaviorally complete exam questions with options if they have not really seen it at the en…”
Assertion Not checkable as stated
Self-reflection training data is now core to all frontier foundation models
“What this suggests about the GPT training data is that the self-reflection data has now actually become pretty much core to the training of all frontier models, because we're seeing that happen in non-instruct models across the board.”
Assertion Supported
Deng: Models internally represent uncertainty preceding hallucinatory behavior
“We've seen that models internally have some awareness of like uncertainty or some sort of like user pleasing behavior that leads to hallucinatory behavior.”
Assertion Supported
White: ML trained on experimental data beat first-principles simulations by a large margin
“Two very well-resourced groups. They both tried different ideas, and the machine learning on experimental data beat out first principles simulation by You know, a very large margin.”
Assertion Not yet assessed · timeframe Jan 2026
White: Synthesis routes for dangerous compounds are already available on Wikipedia
“You can go find the synthesis route for many dangerous compounds on Wikipedia. People know what are the targets in the human body that, like, are targeted by most biological weapons. It's not really that much of a mystery.”
Assertion Not checkable as stated
Reggio: Simple web research agents outperformed RL for credit underwriting at Brex
“We made this big investment. We were working with some outside, like the, like a company that specializes in this and the performance we ended up getting was inferior to just building a, like a web research agent.”
Assertion Supported
Cameron: General model intelligence does not correlate with hallucination rates
“One interesting aspect is that we've found that there's not really a, not a strong correlation between intelligence and hallucination rate. That's to say that the smarter the models are in a generalist sense isn't correlated with their ability to, when they do…”
Assertion Not publicly verifiable
Models perform better in custom agent harnesses than native web chatbots
“And what's really interesting is that if you compare, for instance, Claude, 4.5 Opus using the Claude web chatbot, it performs worse than the model in our Agentic harness. And so in every case, the model performs better in our agentic harness than its web chat…”
Assertion Supported
Cameron: Model performance correlates with total parameters, not active parameters
“We, in our benchmark, see a lot of performance correlated more with total parameters than active, and not that correlated with how sparse like the models are. Our accuracy benchmark is part of a omniscience. It's very correlated with total. It's not correlated…”
Assertion Not checkable as stated
Frontier AI research requires up to $100B, far exceeding NSF budgets
“It was one billion dollars a year for computer science and that they're trying to cut that into half of that, but we need 10 to a hundred billion dollars to do frontier AI research.”
Assertion Not checkable as stated
Arena's anonymous Nano Banana test moved Google's stock and product roadmap
“I mean, that moment alone changed Google's like roadmap. Market share. Seriously. I mean, Google stock, billions of dollars are moving because of Nano.”
Assertion Not checkable as stated
Nair: OpenAI already possessed a superior model during the DeepSeek release
“The feeling in OpenAI is that like, well, I think we had a better model already at the time, right?”
Assertion Not checkable as stated
Catanzaro: $100M+ AI seed rounds without roadmaps happen frequently
“Like upwards of a hundred million dollars in a seed round where you have a long-term vision, but not a near-term roadmap. This is something that I'm seeing happening not just occasionally, but quite Frequently.”
Assertion Supported
Pliny: Anthropic added a $20k–$30k bounty but withheld jailbreak data
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”
Assertion Not checkable as stated
John V: AI Security Startups Scrape BASI Discord to Build Guardrails
“Multiple organizations that have like popped up in the past, I would say two or three years for, you can call them like AI security startups, right? Like actively scrape that server to build out their guardrails or their security, like their suite of products”
Assertion Not checkable as stated
Mirzadegan: 38 of 40 top Kleiner Perkins portfolio executives are first-timers
“Inside the KP portfolio, ok, are top eight companies. Let's take five exec roles across the top eight companies. Companies like Rippling and Glean, ok. 38 out of 40 of those roles, ok, those executives report to the CEO for the first time in their career.”
Assertion Not checkable as stated
General Intuition's foundation agent runs purely on vision without reinforcement learning
“This is just a base model. There's no RL, no fine tuning. This model sees no game states. It is purely capable, not sequence acceptance. It's purely predicting the actions from the phrase. That's it.”
Assertion Not checkable as stated
General Intuition action models turn internet videos into free training data
“We transferred it over to a real world video, which means that you can use any video on the internet as free training.”
Assertion Not checkable as stated
Medal holds the internet's largest action-labeled video dataset by orders of magnitude
“We have sort of the largest data set of ground truth action labeled video footage on the internet by maybe one or two orders of magnitude.”
Assertion Not checkable as stated
Medal has more concurrent sim-drivers than Waymo has autonomous cars
“We have more people at any given time on metal playing with steering wheels and like truck simulator and these types of games than Waymo has cars on the road. It's a ridiculous stat, but it's true.”
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Assertion Not checkable as stated
Li: Marble is the first public high-fidelity 3D generative world model
“It's the first in-class model in the world that generates three D worlds in this level of fidelity that is in the hands of the public.”
Assertion Supported
Sam Altman Barred Investors Who Backed Glean From Investing in OpenAI
“Sam Altman once came out and said, if you're an investor in OpenAI and one of these five companies, including Glean, we don't want you as an investor.”
Assertion Partly supported
Anthropic Is the Fastest-Growing Software Company in History
“Anthropic is the fastest growing software company of all time. I think I can say that fairly. I'm, I haven't been disproven yet.”
Assertion Not checkable as stated
Diffusion Models Achieve 80% of Autoregressive Quality at One-Tenth the Cost
“Diffusion models today are, I would say, 80 to 90% of the quality at one-tenth the cost and latency.”
Assertion Partly supported
OpenAI Plans to Scale Compute Power Capacity to 125 Gigawatts
“For OpenAI to go from like two gigawatts of compute this year to 30 with everything they've already announced, and then there's a plan for the next 125. Like, the United States uses 300.”
Assertion Not checkable as stated
Vercel's v0 added $1M MRR every 14 days after chat rewrite
“When we launched V-Zero, the chat version, or the new V-Zero, whatever you want to call it it's like, 14 days, another million MRR, 14 days, another million MRR, it was like a rocket ship after that.”
Assertion Not checkable as stated
HackerRank: Junior developer assessments and interviews are up 20-25% year-over-year
“Are companies sending assessments to junior versus senior? How many people they are interviewing? It's grown about 20 to 25% year over year relative to all of the other, other news, news items that you see on social media, where it's like, hey, declining, ever…”
Assertion Not checkable as stated
Ravisankar: AI-native companies are the least AI-forward in technical hiring processes
“The AI native or AI forward companies are the least AI forward when it comes to interviewing at high end process.”
Assertion Supported
AMD MI300X outperforms Nvidia H100 on FlashAttention-2 and memory-bound workloads
“We found that it's great for flash attention to specifically, we were able to be H-one hundred. We also found that like the less time you spend in like dense compute, like the less time you spend in tensor cores specifically, or less time you spend in lower bi…”
Assertion Supported
Rumbelow: Leap Labs' Discovery Engine Automates Novel Scientific Discovery via Interpretability
“Discovery Engine is an end-to-end system, takes in arbitrary scientific data set, automatically trains a bunch of neural networks on it, and then We systematically, with our interpretability methods, which is the real secret extract the patterns that have been…”
Assertion Not checkable as stated
Rumbelow: arXiv Is Filling With Plausible but Unverified AI Papers
“I'm actually really worried about this because I think we're already seeing archive and other online repositories and... Submissions too, full of these very, very plausible papers. That may or may not be true. And like at that point, what good is the, is our s…”
Assertion Not checkable as stated
Sands: Top 100 AI startups reach ARR milestones 2-3x faster than SaaS
“One cohort that we looked at was the hundred highest grossing AI companies on Stripe. And you kind of need a reference point. And so we were like, let's compare them to the hundred highest grossing SaaS companies from five years prior. And we looked at things …”
Assertion Not checkable as stated
Sands: Vertical AI wrapper companies are building healthy unit economics
“Increasingly what we're seeing from the AI companies on Stripe is they do want to have healthy Unidec. I mean, let's not talk about like the big labs that are pouring crazy money into research, but if you're talking about like the vertical kind of wrappers, wh…”