why aren't all 18 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Assertion Not checkable as stated
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Assertion Not checkable as stated
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Assertion Not checkable as stated
Kaiser: No frontier AI model can solve a specific first-grade math exercise
“I took one exercise from this math book and none of the frontier models is able to solve it.”
Prediction Not checkable as stated
Kaiser: OpenAI aims to create an AI intern by late 2026
“I think that's what OpenAI says is they say, you know, we say we'd like an AI intern by the end of next year.”
Assertion Not checkable as stated
Kaiser: AI progress has been a smooth exponential increase in capabilities
“Fundamentally, if you look at AI progress, it's been a very smooth exponential increase in capabilities.”
Prediction Not checkable as stated
Łukasz Kaiser: Next-gen reinforcement learning will operate on general data
“I do believe the era of tomorrow will be broader. It will work on general data and maybe then it will expand to like domains that, that go beyond where, where it shines today.”
Assertion Not publicly verifiable
Kaiser: ChatGPT uses a secondary model to summarize raw reasoning steps
“So in the current chat GPT, you will see a summary of the chain of thought on the side. So there is another model that takes the full chain of thought and shows you a summary because the full ones are usually not very nice to read.”
Assertion Not checkable as stated
Kaiser: The eight Transformer paper co-authors were never in one room
“I don't think all eight of us were ever in the same physical room.”
Assertion Not checkable as stated
Łukasz Kaiser: AI models still struggle with multimodal and sequential reasoning
“The models are just, they're starting, like you see the first example they manage, so they've clearly made some progress, but they have not yet learned to do good reasoning in multimodal domains, and they have not yet learned to use one reasoning in context to…”
Assertion Supported
Kaiser: Translation industry grew and translators earn more post-Transformers
“The translation industry has grown considerably since then. It has not shrunk. There's more translations to be done. Translators are paid more.”
Assertion Not checkable as stated
Kaiser: AI coding tools recently became how many programmers work
“But I think it's the recent few months when the transition happened from, you know, people using it sometimes, but rarely, to now basically this being how a lot of people work in coding.”
Assertion Not checkable as stated
Kaiser: Multimodal AI capabilities lag behind text performance
“The multimodal part still lags behind the text part to a large extent.”
Assertion Supported
Kaiser: Reinforcement learning causes AI models to self-correct mistakes
“Even for math and coding, you start seeing that the models start correcting their own mistakes, right? Earlier, if the model made a mistake, it generally just tell you what it did and insist that the mistake was right or something like that. With the thinking,…”
Assertion Not checkable as stated
Kaiser: RL is a major component in post-training tone steering
“I don't work on post-training and it certainly has a lot of quirks, but I think the main part is, is indeed RL where you say, okay, is this response cynical? Is this response like that? And you say, okay, if you were told to be cynical, this is how you should …”
Assertion Not checkable as stated
Kaiser: Pre-training consumes the most GPUs of any AI development stage
“Currently, pre-training just uses the most GPUs of all the parts, so it needs the most GPUs, right?”
Assertion Not checkable as stated
Kaiser: OpenAI's frontier AI models do not yet have 100 trillion parameters
“Our models don't have a hundred trillion parameters yet.”
Assertion Supported
Kaiser: GPT-V Pro solves dot puzzles by running Python code loops
“The GPT-V Pro will run Python code to extract these dots from an image, and then it will count them in a loop.”