why aren't all 39 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Insight
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Assertion Not checkable as stated
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Insight
Kaiser: Pre-training science is plateauing, but compute scaling still improves loss
“Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put More compute. You will get better losses if you do things right, which is extremely hard, and that's valuable.”
Assertion Not checkable as stated
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Insight
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Assertion Not checkable as stated
Kaiser: No frontier AI model can solve a specific first-grade math exercise
“I took one exercise from this math book and none of the frontier models is able to solve it.”
Insight
Łukasz Kaiser: AI pre-training expands stored knowledge rather than generalization
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”
Prediction Not checkable as stated
Kaiser: OpenAI aims to create an AI intern by late 2026
“I think that's what OpenAI says is they say, you know, we say we'd like an AI intern by the end of next year.”
Assertion Not checkable as stated
Kaiser: AI progress has been a smooth exponential increase in capabilities
“Fundamentally, if you look at AI progress, it's been a very smooth exponential increase in capabilities.”
Insight
Kaiser: Reasoning yields far greater AI capability gains per dollar than pre-training
“With the new paradigm of reasoning, you can get much more gains for the same amount of money because it's on this like lower and like, there are just discoveries to be made and these discoveries unlock insane capabilities.”
Insight
Łukasz Kaiser: Reasoning models require verifiable data, excelling in math and coding
“So currently, and current for at least the Most basic ways we use it currently, it needs to be fairly verifiable. So there is an, is your answer correct or not? You prepare data for that. You can do that in mathematics, coding very well. You can do this in sci…”
Prediction Not checkable as stated
Łukasz Kaiser: Next-gen reinforcement learning will operate on general data
“I do believe the era of tomorrow will be broader. It will work on general data and maybe then it will expand to like domains that, that go beyond where, where it shines today.”
Insight
Kaiser: Math reasoning in AI models transfers to generic web searching
“If you learn to think for math, you can, you will sometimes do some, you know, some strategies are the transfer very much like look up on the web and see what they say and use that information. So some of these things are very generic and they start to transfe…”
Opinion
Kaiser: AI reasoning in visual domains is currently very undertrained
“I think, especially thinking in the visual domains is very under trained, I believe.”
Assertion Not publicly verifiable
Kaiser: ChatGPT uses a secondary model to summarize raw reasoning steps
“So in the current chat GPT, you will see a summary of the chain of thought on the side. So there is another model that takes the full chain of thought and shows you a summary because the full ones are usually not very nice to read.”
Assertion Not checkable as stated
Kaiser: The eight Transformer paper co-authors were never in one room
“I don't think all eight of us were ever in the same physical room.”
Insight
Kaiser: AI tech labs are more similar than people think
“I think in, in general, the tech Labs are more similar to each other than people think. There are some differences, but I think if I look at it from the world, you know, from the university in France, the difference between this university and any of the tech …”
Insight
Kaiser: Reinforcement learning reasoning works better on larger pre-trained models
“Pre-training has always worked. And the beautiful thing is it even stacks with RL. So if you run this thinking RL process on top of a better model, it works even better. Than if you run it on top of a smaller model.”
Insight
Kaiser: Interpretability of large AI models faces fundamental complexity limits
“So the understanding of what the models are doing on a higher level has progressed a lot, but then it's still an understanding of what smaller models do, not the biggest ones. But it's not so much that these patterns don't apply to bigger models. They do. It's…”
Assertion Not checkable as stated
Łukasz Kaiser: AI models still struggle with multimodal and sequential reasoning
“The models are just, they're starting, like you see the first example they manage, so they've clearly made some progress, but they have not yet learned to do good reasoning in multimodal domains, and they have not yet learned to use one reasoning in context to…”
Assertion Supported
Kaiser: Translation industry grew and translators earn more post-Transformers
“The translation industry has grown considerably since then. It has not shrunk. There's more translations to be done. Translators are paid more.”
Disclosure
Kaiser: OpenAI began working on reasoning models around three years ago
“So we started working on it maybe three years ago”
Assertion Not checkable as stated
Kaiser: AI coding tools recently became how many programmers work
“But I think it's the recent few months when the transition happened from, you know, people using it sometimes, but rarely, to now basically this being how a lot of people work in coding.”
Assertion Not checkable as stated
Kaiser: Multimodal AI capabilities lag behind text performance
“The multimodal part still lags behind the text part to a large extent.”
Insight
Łukasz Kaiser: Early RLHF was brittle but crucial for chatbot development
“So it was a bit of a brittle technique, but it was a bit of RL that was extremely crucial to making the models chat.”
Assertion Supported
Kaiser: Reinforcement learning causes AI models to self-correct mistakes
“Even for math and coding, you start seeing that the models start correcting their own mistakes, right? Earlier, if the model made a mistake, it generally just tell you what it did and insist that the mistake was right or something like that. With the thinking,…”
Insight
Kaiser: Making deep learning ideas work is harder than generating them
“In deep learning, people laugh that ideas are cheap. Making them work is, is the hard part.”
Disclosure
Kaiser: Projects, not individual researchers, compete for GPU access at OpenAI
“I don't think it's so much people that compete. I think it's more projects that Compete for GPU access.”
Assertion Not checkable as stated
Kaiser: RL is a major component in post-training tone steering
“I don't work on post-training and it certainly has a lot of quirks, but I think the main part is, is indeed RL where you say, okay, is this response cynical? Is this response like that? And you say, okay, if you were told to be cynical, this is how you should …”
Disclosure
Kaiser: OpenAI model names are now detached from technical milestones
“Now the naming is by capability, right? GPT-Five is a capable model. 5.1 is a more capable model. Mini is the smaller model that's slightly less capable, but faster and cheaper. And the thinking models are the ones that do more research, right? In that sense, …”
Insight
Kaiser: Engineering complexity is the main bottleneck for experimental AI research
“The engineering part is the biggest bottleneck. I mean, GPUs are a bottleneck too, when you scale really up. But implementing something that's larger than one machine, it's an experimental research project, so you don't have a team to do that.”
Insight
Kaiser: Unlimited model connection to the real world carries severe safety risks
“This thing, like how do models connect with the external world? It's a fundamentally very hard problem because, you know, when you connect in an unlimited way, you can break things in the real world.”
Insight
Łukasz Kaiser: Robotics limitations reveal gaps in multimodal AI reasoning
“Robotics is probably just an illustration that we are not doing that well in multimodal and that we're not doing that well in general reasoning yet.”
Disclosure
Łukasz Kaiser says Ray Kurzweil was his first manager at Google
“I came to Ray Kurzweil's group. He was my first manager.”
Assertion Not checkable as stated
Kaiser: Pre-training consumes the most GPUs of any AI development stage
“Currently, pre-training just uses the most GPUs of all the parts, so it needs the most GPUs, right?”
Assertion Not checkable as stated
Kaiser: OpenAI's frontier AI models do not yet have 100 trillion parameters
“Our models don't have a hundred trillion parameters yet.”
Insight
Kaiser: Distillation lets OpenAI combine research projects without long pre-training runs
“With distillation, you have the ability to put a number of projects into one model. It's kind of nice that you don't need to wait on all of them to complete at the same time. You can try to periodically put this together, actually make sure that as a product i…”
Assertion Supported
Kaiser: GPT-V Pro solves dot puzzles by running Python code loops
“The GPT-V Pro will run Python code to extract these dots from an image, and then it will count them in a loop.”