why aren't all 1,824 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Patel: Tech moats are shallower than ever due to massive capex
“I think moats are as shallow as they've ever been, right? Because how fast things are moving, the size of the numbers that are being thrown around now, right? It's hundreds of billions of dollars for each major hyperscaler. The size of the numbers are so large…”
Insight
O'Laughlin: If AI automated coding, finance work will be automated too
“I think coding is a little harder, if I'm being honest with you, and you're telling me the hard one got automated. Why can't the easy one get automated? So I started to ask myself, how much can we do? And the answer is, it feels like a skill issue.”
Insight
O'Laughlin: Reviewing AI agents requires past hands-on manual domain experience
“If you didn't pay any like human cognition to get there, I don't think you're going to be a great reviewer. One of the reasons why, you know, what makes that, that human feet, that loop well is because once upon a time you did that and you could make the three…”
Insight
Casado: AI breaks the mythical man-month by turning money directly into capability
“Another thing that is very different this time than in the history of computer science is, is in the past, if you raised money, then you basically had to wait for engineering to catch up, which famously doesn't scale. Like, the mythical man must take a very lo…”
Insight
Casado: Most hardware robotics startups inevitably become vertical sector companies
“When it comes to hardware most companies will end up verticalizing. Like, if you're
If you're investing in a robot company for an app, for agriculture, you're investing in an ag company, because that's the competition and that's the pricing and that's the sup…”
Insight
Casado: A $1B model training run economically justifies a custom ASIC
“No, no, no, a billion dollar training run, a one billion dollar training run, it makes sense to actually do a custom ASIC if you can do it in time. The question now is timeline. Not money. Cause just rough math. If it's a billion-dollar training run, then the …”
Insight
Casado: Dedicated AI coding models fail since programming requires general intelligence
“And I think one of the conclusions is, is like there's no such thing as a coding model. You know, like, that's not a thing. Like, you're talking to another human being, and it's good at coding, but like, it's got to be good at everything.”
Insight
Swyx: Agent labs will capture higher margins than model labs
“I think what I've been calling agent labs, which are people who build on top of all the other models. We'll probably have a better time with the margins because they price against the end user hours spent or like human labor. Whereas models get commodity price…”
Insight
Jeff Dean: Analog computing loses power advantages at digital boundaries
“I mean, I think there's still a, there's also sort of the more exotic things like analog based computing substrates as opposed to digital ones.
I'm, you know, I think those are super interesting cause they can be potentially low power.
but I think you often …”
Insight
AI Excels at Structure Prediction but Fails to Model Physical Folding
“Folding is the more complex process of actually understanding, like, how it goes from, like, this disordered state into, like, a structured, like, state, and that I don't think we've made that much progress on, but the idea of, like, yeah, going straight to th…”
Insight
Core AI capabilities must be built during pre-training, not just fine-tuned
“If there is a core capability that you actually care about, that capability should be part of the foundation and not a fine-tuned artifact.”
Insight
Small specialized pre-trained models can match capabilities of larger fine-tuned models
“When you think of the fact that by doing specialized pre-training, you can train a smaller model, which is as capable as a much larger model when fine-tuned.”
Insight
Bissell: Activation Steering and In-Context Learning Are Quantitatively Equivalent
“He actually has a paper that, as well as some, you know, others from the team and elsewhere, that go into the essentially equivalence of activation steering and in-context learning, and how those are from a, he thinks of everything in a cognitive neuroscience …”
Insight
Garg: Context graphs for unstructured data will consist of small models
“And so where the unstructured data is such a large part of it and conversation data is such a large part, most likely the context graph will actually consist of small models.
The context graph itself is a model layer.
Like you use that data to train a set of m…”
Insight
White: AI Science Agents Do Not Require Bespoke Automated Labs
“We don't actually have to hold their hands so much anymore, or like they don't actually need to necessarily have an automated lab. They can like write an email to a CRO or something, or they can like tell you what experiment to do. And you can take a video of …”
Insight
White: Simulating biological systems hits fundamental limits; empirical lab measurement is unavoidable
“Basically the best model whatever, Opus seven or GPT 10, like it really can only propose the first experiment, maybe slightly more clever, but at a certain point you just need information, right? Like some little calculations you can do that, like, there's mor…”
Insight
White: RLHF fails on scientific hypotheses by ignoring impact and information gain
“We learned a lot about how bad our LHF is with people, just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what…”
Insight
White: Verifiers in the loop provide higher signal than hypothesis rankings
“I have a lot more faith in these, like verifier-in-the-loop kind of scenarios where you have either data analysis, literature search, or you're running a unit test, or whatever, you're going and running the experiment. Anything like that, I think, is going to …”
Insight
White: A world model serves as unifying glue for scientific agents
“In cosmos, we basically, we had all the pieces sitting around. We working on world models, we working on data analysis agent, working on literature agent. And then we're working on, you know, we built a platform for scientific agents. So we had things that can…”
Insight
Yi Tay: Gradient descent learning paradigm is AI's bottleneck, not architecture
“It's not architecture itself. That's, that there's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the to…”
Insight
Yi Tay: The 'Bitter Lesson' Is Overapplied; Architectural Ideas Fundamentally Matter
“The bitter lesson gets used too much in, like, too conveniently used around, but actually there's also a little bit of a, not a bit, there's also a sweet lesson where it's like, ideas matter.”
Insight
Reggio: MCP belongs on conventional systems, not inter-agent communication
“I think of like the MCP and tool usage as being like the interface to all of our conventional imperative systems, not at the AI space.”
Insight
Reggio: Agentic development amplifies bad architecture, yielding less net capacity increase
“I view agentic development as being something that amplifies all the good, just as much as it amplifies all the bad. And it amplifies sloppiness, poor architectural thinking misunderstanding of the requirements. Like there are, for all of the acceleration of g…”
Insight
Reggio: Constraining LLMs to deterministic DAGs undersells their planning power
“My intuition has been that trying to craft LLMs into deterministic workflows and DAGs is, is kind of underselling like the power that they have to actually plan and execute more in a more sophisticated, like fluid way.”
Insight
Hill-Smith: Widely tracked AI benchmarks improve without reflecting general intelligence gains
“Once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the …”
Insight
Cameron: Models perform better with minimal tools than rigid frameworks
“I think where we're getting to is that these models have gotten smart enough, they've gotten better, better tools that they can perform better when just given a minimalist set of tools and let them run, let the model Control the agentic workflow rather than us…”
Insight
Bissell: Probing internal model features matches LLM-as-a-judge quality at 500x lower cost
“If you ask that model, try to like use it as an LLM as a judge, it's not very good. But if you probe its mind and you sort of detect when the features related to personally identifiable information are firing, that gets you the highest recall of anything. It's…”
Insight
Eysenbach: 1,000-layer RL requires reward-free objectives, not just architectural tricks
“I think the main conclusion is that using big networks not only requires these architectural tricks, but also, as Kevin mentioned before, it requires using a different objective. This objective doesn't actually use rewards in it, and so there's another word in…”
Insight
Kevin Wang: Cross-entropy trajectory classification enables scalable deep reinforcement learning
“I think it's because we're fundamentally shifting the burden of learning from something like, Q-learning or, like, regressing to, like, TD errors, which we know is quite spurious and noisy and biased, to fundamentally, like, a classification problem. We're try…”
Insight
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR,
They're both policy gradient methods, but the, what's different is just like the input data.”
Insight
McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know…”
Insight
Nair: Academia rewards complex math over simple, generalizable solutions
“One of the pitfalls of academia is that it doesn't really reward, like, simple ideas that work, and instead kind of tends to reward, like, kind of mathier ideas. Those mathier ideas also give you these, like, kind of implicit knobs to tune that allow you to, l…”
Insight
Nair: RL on LLMs is peaky and fails to generalize beyond training
“RL, the way it's applied to LLMs right now, is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much. It generalizes to some extent, and generalizes in interesting ways, but It's like very peaky, right? Like …”
Insight
Nair: Context integration, not model intelligence, bottlenecks useful automation
“A big thing that needs to happen is, like, it's not, it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, …”
Insight
Nair: RLHF Is a Side Branch Because Compute Cannot Be Scaled
“I think human feedback is kind of like a bit of like a side branch, because you can't really pour that much compute Into it, right? It's like, you take the model, and you, like, elicit it to be a little bit better in terms of personality”
Insight
Yegge: IDEs should only run in background for AI agents
“Yeah, so you, all you do is leave IntelliJ running, but you shouldn't look in it. It's a tool for the AI now, right?”
Insight
Yegge: 10x AI productivity turns merging into architectural rewrites
“As soon as you get to the point where like every developer is 10 times as productive, merging their code becomes this incredibly complicated problem because I, you and I work at the same time for two or three hours. We make, you know, 30,000 line change each. …”
Insight
Fioca: AI models develop operational habits during training analogous to muscle memory
“This is one of the coolest things about, like, model training is literally, like, they develop habits. It's just like a person does. Like, if you're, like, working on some podcasting tool, right, you're really good at editing, and then somebody makes you use a…”
Insight
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Insight
AI will widen the performance gap between good and lazy engineers
“I do believe that AI will do only one thing. It will Separate faster the good engineers from the bad engineers. If you're a good engineer, and you're using AI well, you will be an amazing engineer. If you're a poor, lazy engineer, and you don't want to underst…”
Insight
Pure technical moats are gone; UX is now the only differentiator
“In the world where the technical moat is not that a moat anymore, because Like startups in two weeks, they can build something that is close to what you're building. The difference is like the, how you think about the user or the flow and all of that.”
Insight
Goyal: Publishing Public Benchmarks Is Marketing, Not Product Improvement
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's good marketing.”
Insight
Malte Ubl: When Vibes and Eval Data Disagree, Vibes Are Right
“I think that the common quip that if the vibes and the data disagree, the vibes are probably right. It's true, right? So you have to like, be honest with yourself, like, do they agree and kind of iterate On them over time.”
Insight
Goyal: Providing eval criteria and examples is more effective than writing specs
“In many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”
Insight
Proprietary data cannot be accurately valued without training models to evaluate capabilities
“I don't think you can value it unless you actually model it yourself and see what the capabilities are.”
Insight
Founders negotiating large AI data deals should demand equity, says de Witte
“I would recommend, if you're gonna do large data deals, like, just try to get like a large chunk of equity in the company that you're doing it with if you can.”
Insight
Making billions of gameplay clips playable bridges imitation learning to reinforcement learning
“Actually making every single clip on the platform playable at billions of clips scale is how we go from imitation learning to RL.”
Insight
Johnson: Pixels offer a more lossless world representation than tokenized text
“And then like you actually lose something if you translate to this like purely tokenized representations that we use in LLMs, right? Like you lose the font, you lose the line breaks, you lose sort of the two D arrangement on the page. And for a lot of cases, f…”