why aren't all 107 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
OpenAI launches Agent Kit to build, deploy, and optimize agents
“We launched Agent Kit today. Full set of solutions to build, deploy, and optimize agents.”
Assertion Supported
Wu: OpenAI was first to launch stateful responses API
“Obviously we were the first one to launch responses API, but like a couple of other people have kind of adopted, I think Grok has it in their API. I think I saw LMSYS just did something”
Assertion Supported
Feldman: Sam Altman and Ilya Sutskever invested in Cerebras' early rounds
“In 2016, we met with Sam Altman and Ilya Suskovard at OpenAI and they were an idea and we were PowerPoint, right? That's amazing. And what AI was doing was identifying cats in pictures. And I think they ended up investing in us, both of them and many of their …”
Prediction Held up
Brockman: Most AI compute will shift from training to inference
“We're going to move from a world where most of the compute is training the model as we've deployed these models more, you know, more of the compute goes to inferencing them and actually using them.”
Assertion Supported
DeepMind and OpenAI eliminated formal Lean translation for 2025 IMO solutions
“What surprised me is this time they don't use formal language, but instead they just use LM. And so last year when they tried to do the IMO, they need like a like a human to kind of translate the natural language. Problems to Lean, and then they use Lean to ki…”
Assertion Supported
OpenAI increases prompt caching discount from 50% to 75% on GPT-4.1
“We've increased our prompt caching discount from 50% to 75% on these models.”
Assertion Supported
GPT-4.1 reduces extraneous edit rate to 2%, down from GPT-4o's 9%
“And we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot. 4.1 is at two percent, so it's a pretty big improvement.”
Assertion Supported
OpenAI launches GPT-4.1 model lineup featuring 1M-token context window
“Yeah, I'll just say we released three new models today, GPT-Fort.one, GPT-Fort.one mini, and GPT-Fort.one data, and the real focus on these were just making the models that were great for developers so we improved instruction following, coding, and shipped our…”
Assertion Partly supported
Swyx: Microsoft and OpenAI account for 77% of CoreWeave revenue
“Which are together, 77% of the revenue of CoreWeave.”
Assertion Supported
Swix: Google's entire SigLIP vision team left to join OpenAI
“I think the most recent notable move, I think the entire vision team from Google Lucas Beyer and all the other authors of Siglip left Google to join OpenAI”
Assertion Supported
Swyx: OpenAI Assistants API target sunset is H1 2026
“And assistance API we've has a target sunset date of first half of 26.”
Assertion Supported
Nikunj Handa: OpenAI distilled o-series models into GPT-4o search
“They use, like, synthetic data techniques. They've done, like, O-series model distillation to, like, make these four or fine tunes really good.”
Assertion Supported
Colvin: AI providers are centralizing around OpenAI's API standard
“I think the truth is that everyone is centralizing around OpenAI's SD API as the one to do. So DeepSeek support that. Grok with a K support that. Olama also does it. Well, I mean, if there is that library right now, it's more or less the OpenAI SDK.”
Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Assertion Supported
Soldani: Content owners blanket block crawling due to closed AI models
“What they found is, as a reaction to, like, the close like, of the existence of closed models, like OpenAI or Cloud GPT or Cloud a lot of content owners have blanket blocked any type of crawling to their website.”
Assertion Supported
Weil: ChatGPT supports over 200 million weekly active users
“As we, you know, we support over two hundred million people every week on ChatGPT.”
Assertion Supported
Swyx: Distilling GPT-4 to Mini cut costs 15x with 2% hit
“Yeah, I sat in the distillation session just now, and they showed how they distilled from four to four mini, and it was like only like a two percent hit in the performance, and 15 X cheaper.”
Assertion Supported
Structured response format is limited to GPT-4o and GPT-4o mini
“Actually, the new response format is only available on two models. It's Foro Mini and the new Foro. So the old Foro doesn't have the new response format. However, for function calling, we were able to enable it for all models that support function calling, and…”
Assertion Supported
OpenAI Structured Outputs enforces schemas in one shot without retries
“We are not retrying, you know, we're doing it in one shot and this is how you save on latency and cost.”
Assertion Supported
OpenAI's seed parameter is best-effort and not fully deterministic
“Yeah, the seed parameter is not fully deterministic, and it's kind of a best effort thing. So you'll notice there's more determinism in the first few tokens. That's kind of the current implementation.”
Assertion Supported
Lambert: GPT-4 Turbo Showed a Noticeable Jump on LMSYS Chatbot Arena
“GPT-IV Turbo is also notably ahead of the other GPT-IVs, which it kind of showed up immediately once they added it to the leaderboard, or to the arena, and I was like, all the GPT-IV memes aside, it seems like this is effectively a bump in the model.”
Assertion Supported
McPartlon: OpenAI co-led Chai Discovery's seed funding round
“Actually OpenAI co-led our seed round.”
Assertion Supported
Nathan: ChatGPT Work includes persistent computer environments across sessions
“In ChatGPT work in web and mobile, like, you get access to those, like, persistent computer environment where, you know, you can store files, and those files stay around between sessions.”
Assertion Supported
Lopopolo: GPT OSS Safeguard supports custom enterprise safety specs
“The GPT OSS Safeguard model, for example.
One thing that's really cool about it is it ships the ability to interface with a safety spec.
Safety specs are things that are bespoke to enterprises.
We owe it to these folks to figure out ways for them to instrument…”
Assertion Supported
Lopopolo: Codex can run background builds while concurrently reviewing code
“It basically means that Codex is able to spawn commands in the background and then go continue to work while it waits for them to finish. So it can spawn an expensive build and then continue reviewing the code, for example.”
Assertion Supported
Watkins: OpenAI hired nearly 100 engineers to curate 500 SWE-bench tasks
“So folks at OpenAI did a pretty extensive human data campaign, hiring like almost a hundred real-world software engineers to go through the problems and figure out, like, are the tasks well-specified? Are the tests actually fair and kind of created a curated s…”
Assertion Supported
Hill-Smith: OpenAI Was Untouchable for Well Over a Year
“If we go back even a little bit before then, we're in the era where, when you look at this chart, like, OpenAI was untouchable for well over a year.”
Assertion Partly supported
McGrath: OpenAI 10xed Effective Context Window for GPT-4.1
“I worked on long context, that was why I was on last, was for 4.1, where we, you know, I think, tenxed the effective context window for 4.1”
Assertion Supported
Fioca: GPT-5.1 allows disabling preambles, unlike the reasoning-dependent Codex model
“So Five One, you can turn that off, you can prompt it not to do that, but the Codex model can't actually do that, and it relies on the reasoning summarizer to give you that update.”
Assertion Supported
OpenAI reports 4 million active developers at DevDay 2025
“Every year in Dev Day, you report the number of developers. This year is four million. I think last year was like three.”
Assertion Supported
OpenAI's Nick Cooper sits on Anthropic's MCP steering committee
“We actually have a member of our team, Nick Cooper, who is sitting on kind of like that, that steering committee for MCP as well.”
Assertion Supported
Brockman: OpenAI open-source models saw millions of downloads within days
“Now being used by, you know, there's been millions of downloads of that just over the past couple days.”
Assertion Contradicted
Brockman: OpenAI's Dota AI used only 300 million parameters
“And by the way, Dota was like a three hundred million parameter neural net. Tiny, tiny little insect brain, right?”
Assertion Supported
Brockman: OpenAI Dota used pure RL without human demonstrations
“If you rewind to even 2017, we were working on Dota, which was all reinforcement learning, no behavioral cloning from human demonstrations or anything. It was just From a randomly initialized neural net, you'd get these amazingly complicated, very sophisticate…”
Assertion Supported
DeepMind and OpenAI used different reduction methods to solve IMO Problem 1
“Google has one method and then OpenAI AI also have another method but both kind of works. Basically you just reduced any n to three, and then you just do case by case analysis.”
Assertion Supported
McCloy: ChatGPT personalization and custom preferences directly alter AI search retrieval and sources
“ChatGPT personalization, memories, just explicit preferences, if you set them up, do definitely affect the results you get. Now, Again, obviously, there's, they're still using traditional search, so the search index itself is not necessarily personalized, but …”
Assertion Supported
Marimo Surpasses 300K Monthly PyPI Downloads and Jupyter's GitHub Stars
“I think last I checked, over 300,000 monthly downloads on PyPy. More GitHub stars than Jupyter Notebook for whatever that's worth. And we're used at companies like OpenAI, Hugging Face, Cloudflare, BlackRock, universities like Stanford and Berkeley.”
Assertion Supported
OpenAI stealth-tested GPT-4.1 models on OpenRouter before official release
“Yeah yeah, we really wanted to get as much developer feedback as possible on this model to make sure it worked well in the real world, and so we tested it kind of through Open Router and it was super cool to see people latch on to the names and get the theorie…”
Assertion Supported
OpenAI currently restricts reinforcement fine-tuning exclusively to its reasoning models
“No, that's reinforcement fine tuning is only for reasoning models.”
Assertion Supported
OpenAI will cover inference costs for developers who share custom evaluation data
“You can upload an eval and opt in such that we'll pay for the inference, inference costs if we can also use the eval.”
Assertion Supported
Handa: OpenAI Responses API stores conversation state for 30 days free
“Yeah, it's free. We store your state for 30 days. You can turn it off. But yeah, it's free.”
Assertion Partly supported
Alessio Fanelli: GPT-4o Search jumps to 90% accuracy on simple QA
“On simple QA, GPT four O is 30% accuracy. Four O search is 90%.”
Assertion Supported
OpenAI API throws errors when users exceed a 128-tool limit
“Which OpenAI would be figured out at some point is, like, oh, there is an upper limit of, like, a 128 Tools you can have. And then the API basically gives you an error. So we ran into some of those issues as well.”
Assertion Supported
Reddy: Flagship OpenAI API costs fell 80% to 85% in roughly 18 months
“This is a graph of flagship OpenAI model costs, where the cost of the API has come down roughly 80, 85%, and call it the last year, year and a half which is pretty remarkable.”
Assertion Supported
Schluntz: SWE-bench Verified was created in partnership with OpenAI
“SweetBench Verified was actually made in partnership with OpenAI, and they hired humans to go review all these tasks and pick out a subset to try to remove any obstacle like this that would make the tasks impossible.”
Assertion Supported
Stabilization Techniques Help Continuous Consistency Models Outperform Discrete Models
“And so like when you stack all of these things together, then you're able to train much more effectively and continuous time does much better. Then these discrete, this n is the number of discrete steps that your model is taking, and, you know, maybe one inter…”
Assertion Supported
Wang: OpenAI is opening an office in Singapore
“You know, OpenAI is opening an office in Singapore, and how we can grow more AI engineers in Singapore as well, because I think, I do think that that is something that people are interested in, whether or not it's for their own careers or to hire out in Singap…”
Assertion Supported
Hu: OpenAI o1-preview achieves bronze medals in 17% of MLE-bench competitions
“Their final results with a one preview and this a scaffolding from a different company was that they got a bronze medal. I don't think I've ever achieved once but I haven't competed that much in. 17% of competitions.”