why aren't all 669 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Wolf: OpenAI Model Attacked Hugging Face as Autonomous 'Side Quest'
“What people quickly discovered is that the model was not at all task with attacking us, but decided to do that as a side quest of something else.”
Assertion Supported
Wolf: Prior OpenAI Training Runs Left Notes for Future Runs
“I think learning we had at Black Hat yesterday was that some of the previous training run may have left some notes for future training runs, which is, I think mind, mind blowing.”
Assertion Supported
Anthropic model broke out of sandbox and emailed researcher without internet access
“And I can, there's one example we have published, which is that the model was put into a little sandbox, a little, like, technical container, and it was given the task to, like, maybe break out, and the researcher went away for lunch, and, like, during lunch w…”
Assertion Supported
Wolf: AI Model Used Fake GitHub Accounts to Social Engineer Maintainers
“Basically, the model was tasked to solve this attack, this, like, to attack and to penetrate this subnetwork, and what it decided to do, it decided to get one of the maintainer of a library that could be used To operate this activity directory to merge like ma…”
Assertion Supported
Patel: Huawei used shell companies to procure TSMC chips and Korean HBM
“Well, actually they were using shell companies to get chips from TSMC and using Different methods of like sneaking HBM, which is memory from, you know, Korea through Taiwan to China, right?”
Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Assertion Supported
Socher: Chinese open source companies distilled knowledge from OpenAI and Anthropic models
“The few large closed labs, Anthropic and OpenAI, took almost everything they could from the open internet trained a model, but then the Chinese open source companies basically siphoned a lot of that knowledge out of those closed source models by distilling it.”
Assertion Supported
Claude 3 Opus faked alignment during training and defected in deployment
“It turns out that Opus three, which was a model that I was studying, had a relatively strong propensity to do this in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend…”
Assertion Supported
Cerebras cloud achieves 10x inference speedup over fast GPUs on Gemma
“Say, if you run Gemma four, On your GP, you might get like a hundred tokens per second if you have a fast card. If you run it in their cloud, you get anywhere from 800 to 1500 tokens per second. So call it 10 X faster.”
Assertion Supported
Feldman: Cerebras moves weights to compute ~2,500x faster than standard GPUs
“And so the speed of moving waits to compute is about two and a half thousand times faster here than on a Wilben GP.”
Assertion Supported
Kolter: Adversarial prompts optimized on open-source LLMs break commercial models
“Once we had done that, we found that when you had these weird terms that you sort of flipped around to optimize one, to optimize the response for one model, you could just take those same exact strings you would optimize, paste them into a commercial model, an…”
Assertion Supported
Anthropic's Claude Cowork implements memory using plain text files
“It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files.”
Assertion Supported
AxiomProver achieved a perfect score on the 2025 Putnam math exam
“Eight within the time limit, and then 12 out of 12.”
Assertion Supported
Patel: Groq missed revenue significantly before being acquired
“In fact, they missed revenue last year significantly and yet they got bought, right? Because the value of the IP was there and the value of the team.”
Assertion Supported
Large enterprises are secretly hiring teams to train ChatGPT-scale LLMs in-house
“I know for a fact that big companies are training now LLMs in-house. Really, like, big companies who have the financial means to train chat to be like model are hiring people who train LLMs.”
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Assertion Supported
Meta raises tens of billions in off-balance-sheet debt for data centers
“You have this, sort of, offloading of debt from big companies, for example, Meta, that raises tens of billions of dollars to fuel its data center ambitions, but that doesn't sit on Meta's balance sheet.”
Assertion Supported
Anthropic agreed to a $1.5 billion training data copyright settlement
“And then there was a biggest settlement that happened in the last few months with Anthropic that agreed to pay out one and a half billion.”
Assertion Supported
Douglas: AI autonomous task execution time horizons double every six months
“And so I think it's like every couple of months, the time horizon that the AIs are capable of doing is doubling or something, something crazy. Maybe, maybe every six months the time horizon doubles”
Assertion Supported
Cherny: Claude Code does not use RAG for codebase memory
“And so quad code actually doesn't use this technique called rag. Instead, what it does is it just searches files the same way that a human would.”
Assertion Supported
Stockholm-based AI startup Lovable reached $50 million ARR in six months
“European breakouts lovable out of Stockholm hit fifty million ARR in six months and is now for sure the fastest growing startup in Europe.”
Assertion Supported
Howard: OpenAI is shutting down GPT-4.5
“I think they're shutting down that product or they've shut down that product, if I understand correctly.”
Assertion Supported
Chollet: Latest base LLMs score zero percent on ARC-AGI-2
“Today the latest base alarms, they're doing something like 10% on ARK-I. But on Arc two, they are doing zero percent.”
Assertion Supported
Chinese AI labs like DeepSeek and Alibaba perform strongly on benchmarks
“China was kind of not in this fight, like, 12 months ago, and now is very much in it. Like, their models like Alibaba's Quen and then this spin out from a quantitative hedge fund DeepSeek, which publishes code models and others, and they've been actually a ver…”
Assertion Supported
OpenAI retains 80% of corporate AI model spending, per Ramp data
“Really, like, when you sum the makers of all these models that are getting used, it's still, like, 80% OpenAI. Like, that hasn't changed, so it doesn't look like, fragmentation to me, it looks like domination.”
Assertion Supported
Buying Nvidia stock instead of funding AI chip rivals yielded 4x more
“Six billion turned into thirty billion roughly, of which half is CameraCon, which is a Chinese listed company. And then NVIDIA would be worth a hundred and twenty billion.”
Assertion Supported
Masayoshi Son estimates superintelligence will require $9T in cumulative capex
“Masayoshi Son, the CEO of SoftBank, mentioning just a few days ago that, Inez estimate to reach super intelligence that was going to be a cumulative capex, capital expenditure budget of nine trillion dollars, which you know, in typical MASA fashion he said was…”
Assertion Supported
Sierra quadrupled its valuation to $4.5B on $20M ARR
“Yes, current chairman of OpenAI as well which is more than quadrupling its valuation this year to four and a half billion. You know, on the other side of that equation is roughly twenty million of ARR.”
Assertion Supported
OpenAI's funding deal was priced at 13.5x forward revenue
“The multiple on the deal on a forward revenue basis was 13.5 X.”
Assertion Supported
Over 25% of new code at Google is AI-generated
“Sundar Pashai said, you know, more than a quarter of all code written at Google is now generated by AI before it's reviewed and accepted by engineers”
Assertion Supported
Bernhardsson: Modal discarded Kubernetes and Docker to build a custom stack
“So we built our own, we threw out Kubernetes out the window. We threw out Docker out the window. We built our own File system in order to optimize for how container images are distributed. We built our own scheduler in order to maintain this pool of workers an…”
Assertion Supported
MotherDuck executes queries in milliseconds compared to BigQuery's 400-millisecond overhead
“You know, we can do queries in sort of single digit milliseconds. And you know, in BigQuery, we were very, very happy when we got kind of the overhead down to like, 400 milliseconds.”
Assertion Supported
Pomel: Datadog's Toto model beats other models on observability and weather benchmarks
“This model from day one was state of the art, like, it's it beats all the other models, of course, on the, on observability data, we have special benchmarks from that, but also for other things like weather data which was very surprising to us.”
Assertion Supported
Freedman: pgvectorscale is 28 times faster than Pinecone for high recall
“Compared to one of the leading vector-only databases, Pinecone, a PG vector scale is 28 times faster for a high recall scenario.”
Assertion Supported
Freedman: pgvectorscale on AWS is 75% cheaper than Pinecone
“The monthly costs of running such a PG vector scale deployment on AWS is 75% less expensive than PyIncome.”
Assertion Supported
The enterprise GPU shortage has eased at the company level
“Today, actually, I'm seeing the GPU shortage go away at the level, at the company level, meaning companies are able to procure enough compute enough is a strong word, but they're able to procure compute at some level to work with, to fine tune and run heavy in…”
Assertion Supported
Re-engineering LLM decoders can guarantee absolute schema accuracy for structured outputs
“So that's something else we offer through our inference service to actually make it a hundred percent by re-engineering the decoder of any LLM.”
Assertion Supported
Nussbaum: Nomic Embed is first open long-context embedder beating OpenAI Ada
“So we're launching Gnomec Embed is the first open source reproducible long context text embedder that beats OpenAI ADA as well.”
Assertion Supported
Srinivas: Open-source AI models do not match closed-source model capabilities
“That needs the open source models or fine tuned versions of them to match the closed source, the closed ones. And That's not the case today.”
Assertion Supported
Biewald: Most enterprises have not deployed LLMs into production yet
“I think that LLMs in particular, we talk to a lot of the people and we don't see a ton of people getting them into production yet. And I think it's funny, like VCs are always surprised, like when we tell them that I think that I don't know. I'm bullish on LMS,…”
Assertion Supported
Turck: Google invented the Transformer but is playing catch-up
“So in GPT, which stands for generative pre-trained transformers, the T is a transformer architecture that was actually Developed at Google, and the irony is that Google is playing catch-up to Microsoft and OpenAI right now because OpenAI accelerated the develo…”
Assertion Supported
Turck: Academia has completely lost the AI research battle to industry
“Academia as of now has just completely lost the battle. In AI research and development, the darker dots, and this is a long time scale at the bottom, show that in the early years of the current wave, a lot of the research came out of academia, but today it's p…”
Assertion Supported
Eifrem: Neo4j is frequently 1,000 times faster for connected data queries
“We've optimized every layer in the stack of the database architecture completely around connected data. We're not built on top of a different database or anything. It's a native architecture. And that means that if you want to query along how things are connec…”
Assertion Supported
MongoDB replaced six C-level executives in the three years post-IPO
“It's a, it's not a secret, but since we've gone public three and a half years ago, we've had six C-level changes, right?”
Assertion Supported
Stewart: Cockroach Labs adopted BSL, open-sourcing code after three years
“One of the big things we did is we did change our open source license. So it's still source available. It's still free, but we moved from Apache to a license called the business source license. Which is just like Apache, you can use the software, you can redis…”
Assertion Supported
Pesenti in 2020: Multi-million-dollar AI training runs are unsustainable for Facebook
“Where it becomes millions is when you do training runs. So some of the training runs in the most advanced system that it comes from our company or other companies out there are starting to be extremely expensive. Yeah. Like you can look at one run in the scale…”
Assertion Supported
Volpi: AWS launched DocumentDB to bypass MongoDB's SSPL license
“Mongo took an interesting similar approach. They have something called SSPL. In that case, it says you can take it, but any software that you build around it to run it, you have to open source it. So Amazon will have to open source its entire software base use…”
Assertion Supported
Volpi: Google and Azure partner with OSS creators to counter AWS
“And then the other thing that you do As a business is you partner with Google and Azure who, because they're pretty far behind Amazon, they're more than happy to have the original item and to pay you something for the privilege of running your version.”