why aren't all 107 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year.
And my money is on, they have to do a coin.
Like it's, I'm not a crypto guy at all, but like, y…”
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Assertion Partly supported
OpenAI Plans to Scale Compute Power Capacity to 125 Gigawatts
“For OpenAI to go from like two gigawatts of compute this year to 30 with everything they've already announced, and then there's a plan for the next 125. Like, the United States uses 300.”
Assertion Supported
Sam Altman Barred Investors Who Backed Glean From Investing in OpenAI
“Sam Altman once came out and said, if you're an investor in OpenAI and one of these five companies, including Glean, we don't want you as an investor.”
Assertion Partly supported
The Information: OpenAI hit $12B ARR as burn rose to $8B
“We had a story yesterday about open AI and how, like, I think they've reached about twelve billion ARR and yeah, but their burn went from like They projected, like, one billion to, like, eight billion or something.”
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Assertion Supported
Hong: All OpenAI formal math researchers have left the company
“No, no, they all left.”
Assertion Contradicted
All major US AI labs stopped publishing research after OpenAI closed
“Whereas in the United States, since OpenAI closed their doors and stopped publishing, so did all the other labs.”
Prediction Held up
Chen: AI agents will master GUI-based computer use by 2026
“And I can continue just by sort of like saying that that's definitely going to be something I think is going to be something that we'll be capable of in 20, 26.”
Assertion Supported
Fioca: Codex Max manages its own context window to run indefinitely
“Codex Max manages its own context window. And so it can run basically forever without you having to worry about it while it's inside of the Codex harness.”
Assertion Contradicted
OpenAI Spent $7 Billion on Compute, With $5 Billion for R&D
“This year, OpenAI spent seven billion dollars on compute. Only two of that was for all of their inference. The remaining five was R&D. So all of ChatGPT, all eight hundred million users, all of Sora, all of like, all, all the sort of like API volume, two billi…”
Assertion Supported
Glean Operates at a Several Hundred Million Dollar Revenue Scale
“Look at the revenue of Anthropic and OpenAI right now. These are billion dollar revenue scale businesses. Glean is several hundred million dollar revenue scale business.”
Assertion Contradicted
Martin: OpenDeep Research is the top-ranked open-source Deep Research agent
“OpenDeep Research is a deep research agent that I've been working on for about a year, and it's now, according to Deep Research Spence, the best performing Deep Research agent at least on that particular benchmark. So it's pretty good. Listen, it's not as good…”
Assertion Supported
Brockman: OpenAI's robotics team pivoted to build GitHub Copilot
“And we've been through times where, for example, robotics was one in 2018, where we had a great result, but we kind of realized that actually, like, that we can move so much faster in a different domain, right? That, that actually, you know, we had this great …”
Assertion Partly supported
Early GPT-5 testers report noticeable gains across coding, science, and writing
“And the story that we wrote, we kind of talked about how at least the people that we've talked to who tested it so far have been pretty impressed. They seem to think that it's been, you know, there's been improvements in a number of domains and both like scien…”
Prediction Held up
OpenAI's technology will surpass o3 within six months
“I think that Oh, three is not where the technology will be in six months.”
Assertion Supported
Fanelli: Anthropic scrapes 6,000 pages per referral, compared to OpenAI's 250
“Google would be a two to one crawl to referral ratio, so for every two pages, they will read, they will send you one visitor. He said OpenAI is 250 to one, so they'll read 250 of your pages and send you one person. And Anthropic was like 6000 to one. So they'l…”
Assertion Partly supported
Swyx: OpenAI production market share dropped from 95% to 50–75%
“Basically over the course of 23, going into 24, OpenAI has gone from 95 market share to reasonably somewhere between 50 to 75 market share.”
Assertion Supported
Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60%
“And the opening I spend at the beginning, at the end of last year in November of 23 was close to 90% of total volume. And today, less than a year later, it's closer to 60% of total volume.”
Assertion Partly supported
Polu: GPT-4 was ready internally at OpenAI months before September 2022
“I had seen GPT-IV internally at the time. It was September, 20, 22. So it was pre-chat GPT, but GPT-IV was ready since, I mean, I'd been ready for a few months internally.”
Assertion Supported
Angelopoulos: OpenAI o1 crushed Chatbot Arena, proving the benchmark isn't saturated
“So there's this model and it crushed the benchmark. You know, it's just like really like a big gap. And what that's telling us is that it's not saturated yet. And so it's still measuring some signal that was encouraging point.”
Assertion Supported
Hu: OpenAI o1-preview surpasses human Kaggle Grandmasters with seven gold medals
“Since a grandmaster requires five gold medals and oh, and preview gets an average of eight or sorry, seven gold medals. They're out competing even capital grandmasters.”
Assertion Supported
Schulhoff: Preamble Discovered Prompt Injection Before Riley Goodside
“Preamble is the company that first discovered Prompt Injection, even before Riley, and they, like, responsibly disclosed it, kind of, internally to OpenAI”
Assertion Supported
Carlini extracted production models from Google and OpenAI with legal permission
“We ran the attack that let us, yeah, stole several of OpenAI's models. With their permission... We notified everyone who was vulnerable to this attack. Some Google models were vulnerable. Some open AM models were vulnerable. There were one or two other people …”
Assertion Supported
Reddit makes over 200 million dollars in AI data licensing deals
“Yeah, the, I guess the winner in all of this is Reddit, which is making over two hundred million just in data licensing to OpenAI and some of the other AI providers.”
Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Assertion Supported
Patel: Google, OpenAI, and Anthropic are developing multi-datacenter training
“One of the big bottlenecks is how much power and how many chips you can get into a single data center. So, like, A, Google and OpenAI and Anthropic are working on this, right?”
Assertion Contradicted
No commercial products augmented GPT models with custom data pre-ChatGPT
“Like I saw some people doing demos, but like in like a CLI or something like that, but there was no product doing like this model, but with additional data on top of it.”
Assertion Supported
Parakhin: Bing Sydney first launched in India using Megatron, not OpenAI
“The funny thing, I mean, the most interesting anecdote is that Sydney was first shipped in India for and it was not noticed for a long time. And first implementation of Sydney didn't even have open AI model under it. It was during Megatron. Microsoft and the N…”
Assertion Supported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Prediction Didn’t hold up
Swyx: OpenAI will always release both general and Codex model variants
“I'm pretty, like, have pretty high confidence that basically OpenAI will always release a GPT-V and a GPT-V codex.”
Assertion Supported
Pre-December 2024 OpenAI models did not exhibit seahorse emoji self-correction loops
“And so I like ran the OpenAI API across like models released from 23 to 25, and you would see like all the models until twenty-twenty-four December had very Terce and short responses to the question. Is there a seahorse emoji? They would either say that there …”
Assertion Supported
White: OpenAI reached out to red team new models after reading his chemistry paper
“And then opening eye, some people there Lama was there. She saw this paper, and they reached out, like, hey, we're building this new model, and we think it'd be great to red team it to see, like, what could happen with these models if they're applied to chemis…”
Assertion Supported
Weil: OpenAI's Prism allows unlimited collaborators for free
“I think most other tools in the space have hard limits and charge you money and other things. In Prism, it's as many collaborators as you want for free.”
Assertion Partly supported
McGrath: GPT-5.1 dramatically reduced token usage over GPT-5 while boosting evals
“Yeah, and so you can see, like, from five to 5.1, our overall evals, you know, we bumped some. But if you look at a two D plot of how many tokens it takes for us to get that, it went way down.”
Assertion Supported
McGrath: GPT-5 Thinking Matches or Beats Deep Research on Published Evals
“I mean, I think if you look at our published evals, they're, they look, like, basically on par if it's not better, so, like, I mean, that's personally what I do.”
Assertion Supported
Anthropic and OpenAI are collaborating to build a unified AI UI standard
“And now one thing we just announced three weeks ago on the MCP blog is that we're actually working with all, all two of them together to build like a common standard.”
Assertion Supported
Google, Microsoft, Amazon, OpenAI, and Anthropic joined AAIF as platinum members
“You have Google, Microsoft, Amazon Block, Bloomberg, Cloudflare, OpenAI, Anthropic. Just a platinum member, create a foundation.”
Assertion Supported
Fioca: Codex Max can run continuously for 24 hours or more
“Max can run for a really long time. We can go 24 hours or more. I've actually, like, sort of had it gone for more than that”
Assertion Supported
Fioca: GPT-5 matches Codex coding capability but adds step-by-step preambles
“With the five series, because it's more general, and it's just about as good as coding as codex for a lot of things. We've taught it to be more communicative. And so it has preambles before tool calls. It'll say things like, I'm about to go look for this.”
Assertion Supported
OpenAI adds third-party model support to its evals product
“One of the things that we launched today with evals too is ability to use, like, third-party models as well and kind of bring that into one place”
Assertion Supported
OpenAI's API throughput has surpassed six billion tokens per minute
“We actually zoomed past that.”
Assertion Supported
OpenAI integrates with OpenRouter for multi-provider evals
“We have a really cool setup with Open Router, where we're working with them, and then you can bring your Open Router setup. And then with that, you can actually, you know, you write your evals using our data sets tool, or use our data set tool to create a bunc…”