why aren't all 6,166 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 36 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Assertion Not checkable as stated
Kaiser: Pre-training scaling laws still hold across OpenAI and Google
“What scaling clause says is that your loss will log linearly decrease with your compute. We totally see that and clearly Google sees that and all other labs.”
Insight
Kaiser: Pre-training science is plateauing, but compute scaling still improves loss
“Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put More compute. You will get better losses if you do things right, which is extremely hard, and that's valuable.”
Assertion Not checkable as stated
Kaiser: Model hallucinations are dramatically lower than two years ago
“There was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago.”
Insight
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Assertion Not checkable as stated
Kaiser: No frontier AI model can solve a specific first-grade math exercise
“I took one exercise from this math book and none of the frontier models is able to solve it.”
Insight
Łukasz Kaiser: AI pre-training expands stored knowledge rather than generalization
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”
Prediction Not checkable as stated
Kaiser: OpenAI aims to create an AI intern by late 2026
“I think that's what OpenAI says is they say, you know, we say we'd like an AI intern by the end of next year.”
Assertion Not checkable as stated
Lambert: OLMo 3 32B base model matches Qwen 2.5 32B quality
“This base model is similar in quality to the best available, which is like Quinn's 2.5, 32 B is, was still the best base model.”
Assertion Not checkable as stated
Lambert: OLMo 3 7B outperforms Meta's Llama 3.1 8B in internal tests
“And I just think of this cause like Lama 3.1 AP is one of the most used models and hugging base of all time. And this should be better. We're, In our measurements, we see it as being better than Llama.”
Assertion Partly supported
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Assertion Not checkable as stated
Lambert: Chinese companies with $1B+ valuations routinely pirate SaaS software
“Mediumly large, like billion dollar plus valuation companies in China will just like pirate SaaS software.”
Assertion Not checkable as stated
Lambert: Best open-license AI models near the frontier in 2025 were Chinese
“The models that are from closest to the frontier in performance with good license all happened to be Chinese models throughout the year for this case.”
Prediction Not checkable as stated
Lambert: Big tech will realize 95-98% of LLM potential by 2030
“I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech E…”
Assertion Not checkable as stated
Kant: Next-token pre-training gains hit a sigmoidal curve and slowed down
“The first paradigm of kind of pre-training of predicting the next token on the web was becoming sigmoidal and was slowing down in terms of the gains that it had.”
Insight
Kant: If AI intelligence commoditizes, only scale and delivery cost matter
“And within this world, if you think that intelligence is going to become less distinguishable between the companies building it, And becomes a commodity probably more like oil or cloud compute than like bread at the bakery, is there's two things that matter, y…”
Prediction Not checkable as stated
Kant: AI is on track to rewrite $29T of global knowledge work
“I don't think anyone has any doubts anymore that we're now on track to reach human level capabilities and intelligence. And in that world, 29 trillion dollars of knowledge work rewrites itself, right?”
Prediction Not checkable as stated
Kant: AI foundation model gross margins will settle near 40%
“I'm not yet convinced that gross margins in our industry will look like a SaaS company, like 80% plus. I think when we're talking about a commodity as intelligence, that we build value added services on top. Right. It's a foundation model company. On one hand,…”
Insight
Kant: Full-stack AI infrastructure ownership cuts token costs 20-40%
“So when you take all those margins out, all of a sudden you can start seeing that you can serve your tokens, 2030, 40% cheaper than someone else.”
Insight
Kant: Internet data lacks the thought processes needed for AGI
“But the reason that never got us to AGI is because, well, the internet never actually included the data set of the thoughts and actions that created it. It's the final piece of code, it's the final article you write, but not the thoughts that you had and actio…”
Prediction Not checkable as stated
Kant: RL2L will enable model reasoning much earlier in training
“We think we'll push models to a level of reasoning and thought That will happen far earlier in their training than it does today.”
Prediction Not checkable as stated
Kant: AI training will evolve into continuous learning from real-world agent experiences
“And then the fourth stage of training over time increasingly will become learning, continuous learning from real world experiences of these agents. And so those are kind of the four stages that we think training will go to.”
Prediction Not checkable as stated
Kant: External software and agent frameworks will collapse into base models
“But we have a phrase at poolside, which is over time, everything collapses into the models.”
Prediction Not checkable as stated
Kant: AI models will equal top human knowledge workers within 36 months
“I think we have a hard time holding the point of view that models will reach the same level of intelligence and capabilities that the world's most capable people in every field have. And when you take a step back, and don't take the next 12 month view, but jus…”
Prediction Not checkable as stated
Humanoid robotics will mirror self-driving car long-tail struggles
“My personal bet is I think it's gonna be, the humanoid space is gonna look much more like self-driving, where we have some very good isolated demos, but the long tail will kill you. And so we're gonna go through many false starts, and I think this is just the …”
Assertion Not checkable as stated
Top AI serving companies achieve 70% to 90% gross margins
“There are companies here that are making very, very good margins on serving their AI systems, like, 7080, sometimes 90%, depending on the modality and so, like with everything, the average number sucks but, like, when you look at the best companies, it's reall…”
Assertion Supported
Meta raises tens of billions in off-balance-sheet debt for data centers
“You have this, sort of, offloading of debt from big companies, for example, Meta, that raises tens of billions of dollars to fuel its data center ambitions, but that doesn't sit on Meta's balance sheet.”
Assertion Not checkable as stated
Major tech companies abandon green commitments to secure AI power
“Well, a year or two ago, big companies did make commitments to be green as of, you know, as of, you know, as soon as they started inking deals with you know, nuclear companies and, Various energy providers for data centers, all those commitments basically got,…”
Opinion
Sovereign AI initiatives are mostly marketing and sovereignty washing
“I think it's more marketing than it is like a real policy, because at the end of the day, if you buy your stock from the U S and you're not an ally of the U S at some point, then they'll just switch it off. And so part of this is like sovereignty washing, I th…”
Assertion Supported
Anthropic agreed to a $1.5 billion training data copyright settlement
“And then there was a biggest settlement that happened in the last few months with Anthropic that agreed to pay out one and a half billion.”
Assertion Not checkable as stated
Google Search survives because ChatGPT relies heavily on referencing it
“People say, oh, Google search is dead. I think that's like probably completely wrong because ChatGPT references Google a ton.”
Prediction Not checkable as stated
Data center NIMBYism will feature prominently in 2026 political campaigns
“We predicted this kind of nimbyism, not, not in your backyard will kind of take precedence in in major political campaigns in 20, 26.”
Insight
Software solving boring problems survives because AI developers hate boredom
“What's not going to be dead is the problems that like these AI people don't want to work on because it's so boring to build that software.”
Prediction Open · timeframe Oct 2030
Some countries will abandon AI sovereignty to declare AI neutrality
“Some countries will basically abandon their efforts to achieve AI sovereignty and declare AI neutrality.”
Assertion Not checkable as stated
AI autonomous task duration doubles every three to four months
“We are seeing this very consistent improvement over many, many years where every say like, you know, three, four months is able to like do a task that is twice as long as before completely on its own.”
Opinion
Schrittwieser: Wider AI ecosystem may face bubble while frontier labs thrive
“There may simultaneously be like some sort of bubble in, you know, the wider ecosystem, while at the same time, the frontier labs on a very solid trajectory, having a lot of revenue, making a lot of money.”
Prediction Not checkable as stated
Schrittwieser: Current AI paradigm likely to achieve human-level performance in productivity tasks
“I think if you're thinking of, oh, we want some kind of system that can perform at roughly human level in basically all tasks that we care about. Productivity wise. Then I think, yeah, it's extremely likely that the current approach, pre-training RL, you know,…”
Insight
Pre-training aids AI alignment by implicitly instilling human values
“I definitely think we would keep using pre-training data, not just from an efficiency point of view as well, but also I think there is interesting safety angles, because by pre-training and, you know, all this human knowledge, we're implicitly creating an agen…”
Insight
Using chain-of-thought as an RL reward destroys model interpretability
“If you're not careful with RL, you can make interpretability harder. For example, one Common thing with modern models is they do reasoning with the chain of thought. You could look at the chain of thoughts to, you know, see what are the model internal thoughts…”
Insight
Tworek: Calling LLMs strictly next-token predictors is inaccurate in RL era
“Language models do on their own, like fundamental level is they are often called as next token prediction machines. And that's not completely accurate in the age of reinforcement learning, but they still operate on mostly on tokens that are mostly text.”
Opinion
Tworek: OpenAI's o1 was mostly a tech demo for solving puzzles
“O-one like, to be perfectly honest, it was really mostly good at solving puzzles and like maybe a few kind of thinking problems here and there, but it wasn't like, it wasn't a very useful model. It was almost more like a technology demonstration.”
Assertion Not checkable as stated
Tworek: Coding agents are the first successful agentic AI products
“Like, coding agents are at the moment the first, like, pretty successful agentic products built on top of AI.”
Opinion
Tworek: The landmark 2012 ImageNet AI results were not that significant
“From my perspective, and again, this is just how my brain works, that the 20 12 ImageNet results, like, weren't that significant.”
Insight
Tworek: Uninformed researchers pose a greater risk than IP leaks
“It is like, yeah, it is some like risk of losing IP, but I think the risk of not doing the right thing and of people not being informed about research and not being able to do the best research is much higher in my personal opinion and how, how I approach thos…”
Assertion Not checkable as stated
Tworek: OpenAI's o1 release caught US AI labs unprepared for RL
“As far as I know, like our O-one release mostly caught a lot of us labs by surprise. They didn't have like similarly advanced RL research program to my knowledge, basically no one.”
Insight
Jerry Tworek: Pre-training AI models is mathematically simple compared to RL
“The first thing that is important to know and understand, RL is hard. Like, conceptually, if you think about it, and there's still a lot of depth to it, but very conceptually, mathematically speaking, pre-training is dead simple.”
Prediction Not checkable as stated
Tworek: Pre-training and RL are necessary for AGI, but not sufficient
“I generally think something that we are doing, like, pre-training today is necessary. I think something that, like, we are doing RL today is necessary, and there will surely be a few things more, and like, we have a lot of, Very ambitious research programs on …”
Insight
Tworek: Reinforcement learning and pre-training require each other to succeed
“And like, I don't like in terms of a pure RL, I don't think like really pure RL makes sense. RL needs Pre-training to be successful. And I think pre-training, as I said before, needs RL to be successful as well.”