why aren't all 476 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 11 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Yegge: OpenAI sees a 10x productivity gap between AI adopters and non-adopters
“Anecdotally, they're sharing that performance, the performance differences like 10 X By any way that you measure it. So lines of code, commits, business impact, whatever. And it's so stark and pronounced that the people who aren't adopting it are now 10 times …”
What-if
Google would have crushed OpenAI by giving Noam Shazeer half its TPUs
“That muscle did not exist during my time at Google. And I think had they had it, what they would have done would be say, hey, Noam Shazir, you're a brilliant guy. You know how to scale these things up? Like, here's half of all of our TPUs. And then I think the…”
Assertion Not publicly verifiable
Hotz: GPT-4 is an 8-way mixture model with 220B parameters per head
“GPT-IV is two hundred twenty billion in each head, and then it's an eight-way mixture model.”
Opinion
Hotz: Meta attracts researchers who want to publish while OpenAI keeps ideologues
“OpenAI can keep ideologues who, you know, believe ideological stuff, and Facebook can keep every researcher who's like, dude, I just want to build AI and publish it.”
Opinion
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Disclosure
Chen: OpenAI's three-year goal is models conducting end-to-end research
“When we look at our kind of three-year roadmap, right the end goal that we want to reach is one where You know, the models are just doing end-to-end research, and I think a part of that problem is just being able to have the model come up with good taste.”
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Assertion Not checkable as stated
GPT-5 Reproduced Lupsasca's Best Physics Paper in 30 Minutes
“Then when GPT-V came out. It was able to reproduce one of my best papers that took me a very long time to come up with, in like, 30 minutes.”
Assertion Not checkable as stated
Internal OpenAI Model Proved Gluon Amplitude Formula in 12 Hours
“We had this Internal model that could think for a very long time and was extra strong in physics. So we gave it the whole problem from scratch without actually giving it this. We just formulated the problem in a very sharp way and asked the model to solve, to …”
Assertion Not checkable as stated
ChatGPT Solved Open Physics Problem Before Collaborator's Flight Landed
“We decided to start working on it using AI a little bit before Andy was scheduled to come, like the week before. And in fact, using ChatGPT, we solved the problem before he even got off the plane.”
What-if
Parakhin: Liquid AI could beat frontier models with equal compute
“I think if they if they had similar level of compute, they would be very competitive and maybe even beat the largest models, at least from what I've seen.”
Opinion
Lopopolo: Bearish on MCP due to forced token injection and compaction issues
“MCPs I'm pretty bearish on because the harness forcibly injects all those tokens in the context and I don't really get a say over it. They mess with auto compaction. The agent can forget how to use the tool. There's probably only like, what, three calls in Pla…”
Opinion
Lopopolo: Coding models have largely solved all tasks except hard and new
“And I think things that are hard and new is still something that the models need humans. Yeah. Drive. Yeah. But I think those other quadrants are largely solved, given the right scaffold and the right thing that's going to drive the agent to completion.”
Opinion
Lopopolo: Coding models and harnesses are now isomorphic to human engineering capability
“The models are there enough. The harnesses are there enough where they're isomorphic to me and capability and the ability to do the job.”
Opinion
OpenAI Has Effectively Won the Consumer AI Market
“Now that means we're at a point in consumer where, maybe this is too early to say, but OpenAI has kind of won, right? Like, How do you catch up to something where model quality is not going to be differentiated? You already have the users, you already have the…”
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year.
And my money is on, they have to do a coin.
Like it's, I'm not a crypto guy at all, but like, y…”
Opinion
Lenz: Model providers should not dictate enterprise AI policies
“Right now, if you're using a model, you're taking in their own policy. Even if I want to use GPT-OSS, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave, etc. And I don't thi…”
Opinion
Gorkem: Hosting LLMs is a bad business due to Google search competition
“Language models, hosting language models is not a good business. At the time we thought, okay, we are going to be competing against OpenAI and Anthropic and all these labs. Turned, turned out that it was even worse because the killer application of language mo…”
Insight
Embiricos: Scaffolding-heavy AI agents are limited by developers' mental capacity
“A lot of, like, agents that I see are really impressive, but it's basically, like, part of what's impressive is it's like a bunch of developers building this, like, really bespoke state machine around a bunch of, like, short model calls, and so then the upper …”
Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And
Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Prediction Not checkable as stated
Goyal: OpenAI o1 will make agentic frameworks obsolete
“And I think O-one is going to do that to agentic frameworks as well. Hey, I think To me, it seems very unlikely that the, you know, you and me sort of like sipping an espresso and thinking about how, like, different personified roles of people should interact …”
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Prediction Open · timeframe Dec 2026
Patel: OpenAI and Microsoft partnership will likely collapse within years
“Yeah, I expect in the next few years that the OpenAI and Microsoft probably falls apart too.”
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Opinion
Chen: Meta poaching has calmed down and OpenAI came out on top
“I think that met us calmed down a little bit. I think we came out on top”
Prediction Not checkable as stated
Mark Chen: AI scaling laws will continue to hold
“And so I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering, more careful scaling, and it always unlocks that next ability to scale further. So I mean, it's held for You know, almost 10 order…”
Opinion
Chen: Pre-training is not dead and remains underrated in AI research
“Well, I think if you still have a pre-training is dead view of the world I think pre-training is definitely yeah, yeah, not, not dead. It's underrated.”
Assertion Not checkable as stated
OpenAI's Chen: AI models already discover novel theorems and advance sciences
“The initial direction we took was you should move it to real world research, right? And we've seen that the models, they've gotten a lot better at just kind of discovering novel theorems and pushing the frontiers of hard sciences. Even today, right, that's no …”
Disclosure
Chen: OpenAI favors unifying modalities in as few architectures as possible
“For a research lab, I think there are a lot of advantages for it to being under one. So you just have to maintain one infrastructure stack, for instance. I think the cost to, like, maintaining and scaling many infrastructure stacks at once I think that's somet…”
Assertion Not checkable as stated
ChatGPT Pro Generated Lupsasca's Exact Top Three Follow-Up Physics Questions
“You can take this page of this paper and you can feed it to ChatGPT Pro, say, like the best model we have out right now, and you can ask it, what should I do next? Give me the top three follow-up questions to ask based on this paper. I've done this experiment …”
Assertion Not checkable as stated
Terry Tao Says AI Math Proofs Merely Cite Obscure References
“I talked to Terry Tao a couple of weeks ago at UCLA. We had an OpenAI event with IPAM, which is this Institute of Mathematics there. And I talked to Terry Tao and he said that in his view, all of the proofs that he's seen AI come up with in math, even the ones…”
Assertion Not checkable as stated
Open-source AI demand spiked on hype before reverting to frontier labs
“Like all the open source models, I think what happened was they got like very hyped and people were very interested in using them. But I think like over time, like there was a spike in usage for these models. And then it goes back to open AI, Anthropic and Goo…”
Disclosure
Lopopolo: OpenAI Frontier team operates with post-merge or zero human code review
“You know, we, we've moved beyond even the humans reviewing the code as well. Most of the human review is post merge at this point, but it's not even reviewed.”
Insight
Lopopolo: AI-amplified small teams require extreme package decomposition and strict boundaries
“The structure of the repository is like, 500 NPM packages. It's like architecture to the access for what you would consider, I think, normal for a seven person team. But if every person is actually, like, 10 to 50. Then the, like, numbers on, like, being super…”
Insight
Lopopolo: Converting UI Images to ASCII Art Improves AI Agent Layout Perception
“If we want to actually, like, make it see the layout, it's almost easier to rasterize that image to ASCII arc and feed it in to the agent.”
Insight
Lopopolo: AI models can in-house 2,000-line dependencies in an afternoon
“The level of complexity of the dependencies that we can internalize is I would say low medium right now, right? Just based on model capability. What is medium? I would say like a couple thousand line dependency is a thing that we could in house no problem in a…”
Disclosure
Lopopolo: Symphony discards failed PRs entirely to regenerate from scratch
“In Symfony, there's this like rework state where once the PR is proposed and it's escalated to the human for review, it should be a cheap review, right? It is either mergeable or it is not. And if it's not, you move it to rework. The Elixir service will comple…”
Assertion Not checkable as stated
Lopopolo: Zero-code harness was 10x slower initially before outperforming any single engineer
“Honestly, the first month and a half was 10 times slower than I would be. But because we paid that cost, we ended up getting to something much more productive than any one engineer could be, because we built the tools, the assembly station for the agent to do …”
Insight
Lopopolo: Autonomous Coding Removes Human Language Familiarity Constraints
“No humans in the loop here. So like my, Own personal ability to write or not write Elixir doesn't really have to bias us away from using the right tool for the job, which is just wild.”
Opinion
Manning: OpenAI's Sora cannot produce compelling gameplay or persistent mechanics
“Don't think you can take Sora and produce compelling gameplay, right? If you want to have a world that you can wander around in a bit, you're good, but what are your abilities to have gameplay mechanics implemented the way you'd like them to be, and to have th…”
Assertion Supported
Reddy: Voxtral speech model is much stronger than Whisper
“And I think a big people, I think there's a big rich ecosystem of people finding whisper and people want the same thing with Voxer. It's much stronger than whisper.”
Disclosure
Pydantic built a VIP issue scorer after closing an OpenAI founder's ticket
“Basically this started off because one of the OpenAI co-founders created an issue on Pydantic. And we just closed it and said it was wrong. And so we have this that, like, injects itself and tries to summarize someone and it gives them a, like, brutal score of…”
Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Assertion Not checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”