why aren't all 183 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 2 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Disclosure
Tworek: OpenAI's high and low reasoning modes use the exact same model
“Where you can have like a high rezoning model and the low rezoning models. And this is like in the end, the same model. You just, we just tweak the parameter, which says we want you to think longer or shorter.”
Disclosure
OpenAI VP Jerry Tworek pays $200 per month for ChatGPT Pro
“I think I am pretty heavy user of ChatGPT right now, happily paying like 200 dollars a month for it”
Assertion Not checkable as stated
Tworek: OpenAI internal models can currently reason for hours on tasks
“The models I can think for, like, 30 minutes, hour, two hours these days on certain, certain types of Tasks and problems like even, even, even longer than that.”
Disclosure
Tworek: OpenAI's RL algorithm is not GRPO but shares similar components
“Like what we, what OpenAI is doing is not exactly GRPO. It is slightly different in many different ways, but like some parts are definitely similar.”
Insight
Tworek: AI alignment is a never-ending pursuit as human goals evolve
“And it's I think it's a never ending pursuit because like, even, even for humans, it's not super easy to define what's, what do we consider a light? And I think as our civilization will evolve, it will, the notion of alignment and the goals of humanity will, K…”
Disclosure
Tworek: OpenAI focuses research on only three or four large-scale projects
“We all work on a very few projects total. There are not that many projects. OpenAI is not trying to do everything. We are not trying to like have portfolio. We are trying to have like multiple different bets. Always the idea is we do a few core things really, …”
Assertion Not checkable as stated
Tworek: Ilya Sutskever set OpenAI's current RL research roadmap in 2019
“And what he said at the beginning of 2019 was to train large generative model on all data we can and then do reinforcement learning on it. That was the OpenAI research Plan at the beginning of 2019. And this is exactly what we are doing today.”
Assertion Not checkable as stated
Douglas: OpenAI's o1 established test-time compute and RL as a scaling axis
“And I think OpenAI deserves a lot of credit for you know, releasing the first, like, serious RL plus LLMs release with O-one. And I think this really kicked off a pretty, you know, substantial change because it opened up a new axis of scaling, right? There was…”
Disclosure
Douglas: Anthropic intentionally deprioritized math reasoning unlike OpenAI and DeepMind
“You know, one thing that Anthropik, like, You noticeably hasn't focused on compared to DeepMind and to OpenAI is is mathematical reasoning, right? DeepMind and OpenAI have been pursuing mathematical reasoning because of the implications for science and for sci…”
Prediction Not checkable as stated
Pedregal: OpenAI will try to do everything for everyone
“OpenAI is going to try to do everything to everyone.”
Disclosure
Rauch: Vercel Fluid Compute only bills active CPU cycles during LLM wait times
“If you're waiting for 15 minutes for OpenAI to respond, and you do a tiny bit of compute, maybe to like transform it into HTML or transform it into an MCP response or, you know, call other systems. We're only going to charge you for the actual CPU cycles that …”
Assertion Supported
Rauch: Vercel AI SDK is second largest AI module behind OpenAI
“It's the second largest AI module behind OpenAI.”
Insight
Evans: Memory features in AI models create user stickiness, not network effects
“I think that stickiness, I don't think it's a network effect.”
Assertion Not checkable as stated
Evans: Foundational LLMs are similar, but OpenAI holds consumer mindshare
“The models are all sort of the same, but OpenAI is the only one that anyone uses, that, that, that has consumer mindshare.”
Opinion
Evans: Sam Altman's role is mostly fundraising, politics, and promotion
“Like a lot of Sam Altman's role at the moment, it's like you could split his role into capital raising, politics, like internal tech politics and promotion.”
Assertion Not checkable as stated
Chollet: OpenAI used about 75% of ARC training tasks to adapt o3
“So they told us that they were using a significant fraction, I think they said something like 75%, of the training tasks to, you know, to adapt the model in some way.”
Assertion Not checkable as stated
Misra: Microsoft Azure OpenAI outperformed direct OpenAI on uptime and latency
“Microsoft OpenAI did much better in terms of reliability. It just worked perfectly for latency and never went down, which is what you want.”
Assertion Partly supported
OpenAI's $6.6B round required a $250M minimum ticket
“In the OpenAI round that we're talking about, you know, the 6.6 billion dollar round. The minimum ticket size was two hundred and fifty million.”
Assertion Supported
OpenAI's $6.6B round was the largest VC round ever
“We just had the biggest venture capital round of all time with OpenAI, which raised 6.6 billion at one hundred and fifty seven billion post-money valuation.”
Assertion Supported
Safe Superintelligence raised a $1B seed round pre-product
“There's also kind of the biggest seed round ever, right, that we saw. Safe superintelligence helmed by Former OpenAI chief scientist and co-founder Ilya Sutskover raising roughly a billion dollars in cash at a five billion dollar valuation pre-basically anythi…”
Assertion Partly supported
Delangue: OpenAI uses open source and is a Hugging Face customer
“As a matter of fact, open AI is using open source as a customer of Fergingface”
Opinion
Pomel: OpenAI's latest models remain early stage for reasoning capabilities
“There's more we can do maybe on the reasoning side, though I would say even with the latest releases from OpenAI, it's still fairly early in terms of the quality of the models and what can be done there.”
Opinion
Turck calls Sam Altman the Steve Jobs of this generation
“It seems that Sam Altman, who's you know, the Steve Jobs of our generation, we're actually joking the, you know, Steve Jobs of the TikTok generation when you know, unlike Steve Jobs who was fired and came back years later, like he left and came back within a, …”
Assertion Supported
Mistral and Poolside funding is a trickle compared to OpenAI, says Polu
“It's already awesome that Mistral was able to raise that much, that Toolside is able to raise that much, but it's a trickle compared to what Open Air is raising, compared to what Anthropik is raising.”
Insight
Polu: Compute allocation naturally aligns AI researchers with company goals
“There is a way to Orion's an organization, a research organizations, the way you allocate computes. Which means that as a researcher, it's often the case that you are free to work on whatever the things you want to work on. Right. But if the things you're work…”
Assertion Not checkable as stated
Biewald: OpenAI is a W&B customer with a small number of production models
“OpenAI has been, like, a longtime customer. I mean, I consider them, like, extraordinarily sophisticated, and they have a pretty small number of models in, in production, so.”
Assertion Partly supported
Google, Facebook, Microsoft, and OpenAI used Reddit data for AI models
“Google, Facebook, Microsoft, and OpenAI all used Reddit's data to train their conversational AI models.”
Assertion Supported
Katti: OpenAI's Abilene data center is operational and training latest models
“It's up and running. It's being used for training the last two models more actually.”
Assertion Supported
OpenAI taped out its custom Jalapeño chip in just nine months
“Yes, it was incredibly quick. Nine months is very, very fast.”
Assertion Not checkable as stated
OpenAI's Jalapeño chip optimizes tokens produced per watt
“So the key metric that Jalapeno is optimizing is maximizing the number of tokens you can produce per watt.”
Assertion Supported
Katti: Many OpenAI chip team members previously designed Google TPUs
“They have, many, many of the team have designed TPU chips at Google in the past.”
Assertion Supported
Katti: Turbine and transformer manufacturers face multi-year lead times to expand capacity
“Those industries have historically have not added much capacity for the last decade or so ago, and they've suddenly experienced a demand shock. And it takes years before you can add capacity to produce more turbines and transformers.”
Disclosure
Stripe partners with Microsoft and OpenAI for Copilot commerce
“Microsoft and open AI, we're doing something similar with them, like helping businesses make their products discoverable inside copilot and chat GPT.”
Assertion Supported
Prince: OpenAI is a major Cloudflare customer powering its mobile app
“OpenAI is, you know, big customer of ours you know, all their mobile app is built, you know, on, on, on us.”
Prediction Not checkable as stated
OpenAI will release reinforcement learning products for consulting, banking, and legal
“I definitely think OpenAI will have amazing products that will be relevant in those domains, and some amount of RL will play a role in there.”
Disclosure
OpenAI publishes most math research results using informal language settings
“Most of our results that we publicize, as far as I can think, are all in the informal setting.”
Assertion Not checkable as stated
Dubois: Final rehearsal for OpenAI's GPT-5 live demo failed
“Right before we did that, like the last rehearsal, it did not work.”
Insight
Dubois: Quantifying AI improvements is as important as training models
“Finding issues and, like, making sure that we can quantify improvements is just as important, if not more important, but there's always this, like, cultural gap.”
Insight
Dubois: Creating AI evaluations inherently creates methods for building training datasets
“Every time you build an eval, you actually build a way to build training data sets.”
Insight
Dubois: AI performance in new verticals depends on domain expert focus
“And in general, I would say the performance of the model really depends on, like, the number of people who care about the final output of the model and who are looking at that model. So if they start looking more on specific verticals, like, these verticals wi…”
Assertion Not checkable as stated
OpenAI differentiated early on by prioritizing model scale over new methods
“About opening up early on is that they always had this bet on scale. In a time where I think that was looked upon very suspiciously that, oh, if you, the thought somehow that we had all the methods already, and all you had to do was scale them up that mindset …”
Insight
Patel: AI model performance is a lagging indicator of prior hardware CapEx
“Ultimately the capex that Microsoft spent in 2024 for OpenAI is what results in 2025 for OpenAI or CoreWeaver or whoever is what results in their models being so good this year. Same with Anthropic and Amazon Google and their models now being so good now is th…”
Disclosure
Kaiser: Projects, not individual researchers, compete for GPU access at OpenAI
“I don't think it's so much people that compete. I think it's more projects that Compete for GPU access.”
Assertion Supported
Kaiser: Reinforcement learning causes AI models to self-correct mistakes
“Even for math and coding, you start seeing that the models start correcting their own mistakes, right? Earlier, if the model made a mistake, it generally just tell you what it did and insist that the mistake was right or something like that. With the thinking,…”
Disclosure
Kaiser: OpenAI began working on reasoning models around three years ago
“So we started working on it maybe three years ago”
Disclosure
Kaiser: OpenAI model names are now detached from technical milestones
“Now the naming is by capability, right? GPT-Five is a capable model. 5.1 is a more capable model. Mini is the smaller model that's slightly less capable, but faster and cheaper. And the thinking models are the ones that do more research, right? In that sense, …”
Assertion Not checkable as stated
Kaiser: RL is a major component in post-training tone steering
“I don't work on post-training and it certainly has a lot of quirks, but I think the main part is, is indeed RL where you say, okay, is this response cynical? Is this response like that? And you say, okay, if you were told to be cynical, this is how you should …”
Assertion Supported
OpenAI and DeepMind achieved Math Olympiad gold medals
“And, you know, we saw gold medals on the International Math Olympiad by a couple of labs, including OpenAI and DeepMind.”