Everything Łukasz Kaiser said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Kaiser: Pre-training science is plateauing, but compute scaling still improves loss
“Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put More compute. You will get better losses if you do things right, which is extremely hard, and that's valuable.”
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Kaiser: No frontier AI model can solve a specific first-grade math exercise
“I took one exercise from this math book and none of the frontier models is able to solve it.”
Łukasz Kaiser: AI pre-training expands stored knowledge rather than generalization
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”
Kaiser: OpenAI aims to create an AI intern by late 2026
“I think that's what OpenAI says is they say, you know, we say we'd like an AI intern by the end of next year.”
Kaiser: AI progress has been a smooth exponential increase in capabilities
“Fundamentally, if you look at AI progress, it's been a very smooth exponential increase in capabilities.”
Kaiser: Reasoning yields far greater AI capability gains per dollar than pre-training
“With the new paradigm of reasoning, you can get much more gains for the same amount of money because it's on this like lower and like, there are just discoveries to be made and these discoveries unlock insane capabilities.”
Łukasz Kaiser: Reasoning models require verifiable data, excelling in math and coding
“So currently, and current for at least the Most basic ways we use it currently, it needs to be fairly verifiable. So there is an, is your answer correct or not? You prepare data for that. You can do that in mathematics, coding very well. You can do this in sci…”
Łukasz Kaiser: Next-gen reinforcement learning will operate on general data
“I do believe the era of tomorrow will be broader. It will work on general data and maybe then it will expand to like domains that, that go beyond where, where it shines today.”
Kaiser: Math reasoning in AI models transfers to generic web searching
“If you learn to think for math, you can, you will sometimes do some, you know, some strategies are the transfer very much like look up on the web and see what they say and use that information. So some of these things are very generic and they start to transfe…”
Kaiser: AI reasoning in visual domains is currently very undertrained
“I think, especially thinking in the visual domains is very under trained, I believe.”
Kaiser: ChatGPT uses a secondary model to summarize raw reasoning steps
“So in the current chat GPT, you will see a summary of the chain of thought on the side. So there is another model that takes the full chain of thought and shows you a summary because the full ones are usually not very nice to read.”
Kaiser: The eight Transformer paper co-authors were never in one room
“I don't think all eight of us were ever in the same physical room.”
Kaiser: AI tech labs are more similar than people think
“I think in, in general, the tech Labs are more similar to each other than people think. There are some differences, but I think if I look at it from the world, you know, from the university in France, the difference between this university and any of the tech …”
Kaiser: Reinforcement learning reasoning works better on larger pre-trained models
“Pre-training has always worked. And the beautiful thing is it even stacks with RL. So if you run this thinking RL process on top of a better model, it works even better. Than if you run it on top of a smaller model.”
Kaiser: Interpretability of large AI models faces fundamental complexity limits
“So the understanding of what the models are doing on a higher level has progressed a lot, but then it's still an understanding of what smaller models do, not the biggest ones. But it's not so much that these patterns don't apply to bigger models. They do. It's…”
Łukasz Kaiser: AI models still struggle with multimodal and sequential reasoning
“The models are just, they're starting, like you see the first example they manage, so they've clearly made some progress, but they have not yet learned to do good reasoning in multimodal domains, and they have not yet learned to use one reasoning in context to…”
Kaiser: Translation industry grew and translators earn more post-Transformers
“The translation industry has grown considerably since then. It has not shrunk. There's more translations to be done. Translators are paid more.”
Kaiser: OpenAI began working on reasoning models around three years ago
“So we started working on it maybe three years ago”
Kaiser: AI coding tools recently became how many programmers work
“But I think it's the recent few months when the transition happened from, you know, people using it sometimes, but rarely, to now basically this being how a lot of people work in coding.”