why aren't all 1,548 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Dubois: Effective reinforcement learning pipelines prevent AI hallucinations caused by SFT
“So, so hallucination at least the intuition that people have is that it can come, for example, from SFT, and it can come from this, like, pursuing pipeline, but if you have good reinforcement in pipeline, that shouldn't happen too often.”
Insight
Dubois: AI models outperform new employees initially but lack continual learning
“Right now, actually most models at day zero, if you just drop them in a company arguably they are more useful than most new employees. So they start higher at T zero. But then across time they are mostly constant because they don't really learn kind of company…”
Insight
Dubois: Last-mile integration is the main AI bottleneck, not raw intelligence
“I think most of the time, the bottleneck is the last mile.”
Insight
Kolter: AI models do not get safer automatically by scaling up
“You can't just sort of trust models to get safer by getting bigger. You have to put in the work to actually make them safer. And this is, I think what a lot of AI companies are investing in. This is why we in fact do have models that are improving on these dim…”
Insight
The vast majority of modern AI intelligence comes from self-training
“I don't think people have properly internalized the fact that the vast majority of intelligence comes from self training effectively.”
Insight
AI product design overhang is now larger than model capability overhang
“As we get more and more powerful, I actually think the overhang in the product is bigger than in the model.”
Insight
Modern AI UI features exist for human reassurance, not model execution
“Most of the buttons you add and most of the product services you build are probably more for the human than they are for the model.”
Insight
Future AI product winners will be decided by UX, not superior models
“If someone beats me, Felix, at like building very good products, I suspect it's going to be not because they built a better model, but likely because they figured out a better user experience.”
Insight
Building hyper-specialized AI products is risky as models dynamically build infrastructure
“As the models get more and more capable, what I'm noticing inside my products and inside my work is that we're sort of like pulling back the edge cases we account for. And I mentioned earlier that memory is just a text file. If Claude needs a database, it will…”
Insight
Current AI products are at the Nokia 3320 stage, not the iPhone
“I think, a thing I tell a lot of my colleagues is that we're really in the silly times of mobile phones, and then if we get really lucky, maybe what we're currently working on is like the Nokia 33 20, like a good phone, but it's not yet the smartphone. It's no…”
Insight
Post-training techniques cannot compensate for a weak base AI model
“Pre-training is still the foundation and like, you can never post-train your way out of a week-based model.”
Insight
Video data conveys physical world knowledge to AI more efficiently than text
“So because of that, like picking up a lot of knowledge about the word through language is just not really efficient. I don't want to say that it's impossible, but it's not efficient, you know, like to learn about gravity. If you kind of like, you know, have yo…”
Insight
Demonstrating that image training lowers text perplexity remains extremely difficult
“So it turned out to be a really, really good model, but it was like really hard to see that. Wow. You know, I train on images and then like Text perplexity goes down. That was hard to see. You know, like the fact that, you know, you train in native model and i…”
Insight
Jagged intelligence in AI is a structural flaw, not a patchable bug
“Not easy to pinpoint like specific things, but again, like, you know, this is just like my personal opinion and maybe I have colleagues and like the other people like sharing this with me, but I think we're underestimating how hard like jagged intelligence is …”
Insight
Evans: Meta and Google do not need standalone LLM monetization
“Because if you are Meta or Google, You've got this whole other highly profitable business, which now needs to have LLMs inside it, powering all sorts of capabilities and features, and you probably want them to be your LLMs rather than somebody else's. But you …”
Insight
Evans: Incremental AI accuracy improvements don't reduce necessary human review
“But if you've got a bunch of use cases where you need the right answer, as opposed to sort of the right answer, then saying that the model is better doesn't mean anything. I mean, literally, it is literally meaningless. What you're telling me is, I asked the m…”
Insight
Evans: Differentiating a chatbot is like differentiating a web browser
“The chatbot itself is kind of like trying to differentiate a web browser in that you've got an input box and an output box, and how can you make them different if the whole point is that you can type in anything and get anything out.”
Insight
Evans: AI product teams are strategy takers, not strategy setters
“You start from the technology. You don't control the product strategy, which is of course how science works, but you don't know what's going to happen. You don't know what's going to get built. You know, obviously you've got like Sam and Dario and so on are li…”
Insight
Evans: Most SaaS companies are fundamentally just database wrappers
“Most SaaS companies are database wrappers, where somebody realized that here is this problem, and here is the people who have it, and here is a way of turning it 90 degrees, and this is your insertion point, this is how you build it and take it to market.”
Insight
Chase: Agent harnesses matter more for performance than underlying models
“I, the, so I don't know what happens, but I do know the harness is really, really important. Like, I think this is the thing that matters.”
Insight
True AI reasoning engines require verifiable rewards for intermediate proof steps
“If you want to have a reasoning engine that really truly masters at logic and mathematical reasoning, then you need to somehow get verifiable reward for the proof steps.”
Insight
Generation and verification loops are the next major frontier of AI
“We still feel like we cannot fully elaborate and emphasize the thing that we are seeing that is the next frontier of AI. That is a generation and verification loop. That is the discovery of verified knowledge.”
Insight
Zeghidour: Training AI to learn world knowledge from speech is terrible
“Getting your model to learn about the world from speech, I think it's a terrible idea.”
Insight
LeCroix: Enterprise value comes from multi-agent workflows, not single agents
“What we see in enterprise is rarely things that are solved with agents because that's not necessarily where you would expect an FDE to be most useful. Where there is more values, value is in more complex workflows where you will have several agents interact th…”
Insight
Lacroix: AI agents use file systems to replace long context windows
“And that I think that was the big change in and realization through vibe coding is that agents are good enough at manipulating file systems that they can use this as a replacement for their Context window, basically. They can select parts of what they want to …”
Insight
LeCroix: Generating thinking traces and calling tools are fundamentally identical in AI
“There's no real difference between Creating a new thinking trace or calling the right tool. It's all the same to me, because what you're optimizing at the end is what is the best output for the model to create before it gets results to me.”
Insight
Lacroix: Data quality improvements yield 10x the gains of model architecture tweaks
“Getting the data perfect, because we knew this was potentially not the most exciting part of the work, but it was absolutely critical, and any improvement on the data quality would, Tenex, the improvements that we would get by really improving on the model arc…”
Insight
LeCroix: Banks would not adopt AGI without enterprise governance controls
“Requirements I see for control and governance in enterprise make me think that even if I had some AGIS model on my servers right now, if I were to go into a large bank and say, Here is a thing. Please let it control everything for you. They wouldn't be happy t…”
Insight
Patel: Chip startups cannot beat Nvidia playing Nvidia's game
“You're never going to beat Nvidia at their own game, right? They're going to have the supply chain on lock. They're going to get to the newest memory technology or process technology or whatever packaging technology, whatever it is, sooner than you.”
Insight
Pre-training is no longer where the low-hanging AI gains lie
“Pre-training is not dead, but pre-training is boring. So it's not where the low hanging fruit is anymore.”
Insight
RLVR unlocks pre-training knowledge rather than teaching LLMs new math
“The knowledge is already there in the pre-training, and this just unlocks it. It's just like a step that maybe shows the model how to use its own knowledge, basically.”
Insight
Bigger LLM gains will come from multi-model process refinement, not scaling
“That's where you make the bigger gains rather than scaling the model size. I think that's one of those things where you will see more progress coming from.”
Insight
Dettmers: Coding agents serve as general-purpose AI agents for digital tasks
“Coding agents are general agents. Coding agents can write programs that solve other problems, and code is so general, if there's a digital problem, you could solve it for code, and coding agents make the thing so easy that now you can solve a variety of proble…”
Insight
Dettmers: Students using AI agents perform poorly on basic domain knowledge
“If we let people use agents, they perform very poorly on basic knowledge. And if we let people just do the basic knowledge, they don't know how to use agents and they can't compete. So they can't do useful work in the workforce nowadays.”
Insight
Izmailov: Rogue AI science fiction in training data likely causes deceptive behavior
“I think at least part of it is probably The models seeing descriptions of AI, like in the science fiction literature going rogue and like, yeah, that probably affects how the models behave in similar scenarios.”
Insight
Izmailov: AI industry excels at execution but lacks bandwidth for exploration
“Industry is really great at executing on ideas and it's maybe not as good at, like, exploring diverse ideas. Even at the scale of Anthropic OpenAI there is a lot of focus in the companies, and there isn't a lot of bandwidth to do exploration, and that has been…”
Insight
Izmailov: AI models can quickly max out defined benchmarks using RL
“And I think we are at the stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out pretty quickly, and so we are going through benchmarks now very, very quickly.”
Insight
Izmailov: Major compute multipliers exist that improve AI without naive scaling
“I think there are still major, like, compute multipliers, major ways of saving compute that can lead to better performance without just naively scaling.”
Insight
Izmailov: Deterministic data transformations create information for computationally bounded models
“But with a limit on the compute, it's actually very possible to apply deterministic transformations to the data. And create information through that.”
Insight
Bourgeau: Architecture and data innovation currently matter more than scale
“The other parts are architecture and data innovation. These also play a really, really important part in the Performance of pre-training and probably even more so than pure scale these days, but scaling is still an important factor as well.”
Insight
Bourgeau: AI research is shifting to a data-limited paradigm
“I think what might be happening instead is kind of a shift in paradigm where before we were kind of scaling in the data unlimited regime where, where data would scale as much as you would like. And we're kind of shifting more to a data limited regime, which ac…”
Insight
Bourgeau: AI models must be trained on harmful data to avoid it
“So at a fundamental level, you did, you do need the model to know about those things. So you have to train a bit at least on those so that it knows what those things are and knows to stay away from those, right?”
Insight
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Insight
Kaiser: Pre-training science is plateauing, but compute scaling still improves loss
“Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put More compute. You will get better losses if you do things right, which is extremely hard, and that's valuable.”
Insight
Kaiser: Test-time compute increases AI capabilities faster than pre-training
“Using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
Insight
Łukasz Kaiser: AI pre-training expands stored knowledge rather than generalization
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”
Insight
Kant: If AI intelligence commoditizes, only scale and delivery cost matter
“And within this world, if you think that intelligence is going to become less distinguishable between the companies building it, And becomes a commodity probably more like oil or cloud compute than like bread at the bakery, is there's two things that matter, y…”
Insight
Kant: Full-stack AI infrastructure ownership cuts token costs 20-40%
“So when you take all those margins out, all of a sudden you can start seeing that you can serve your tokens, 2030, 40% cheaper than someone else.”