The Wisdom Wall
14 quotable lessons, heuristics and mental models. Every one is playable at the moment it was said. No fortune cookies allowed.
“You can't just sort of trust models to get safer by getting bigger. You have to put in the work to actually make them safer. And, and this is, I think what a lot of AI companies are investing in. This is why we in fact do have models that are improving on these dimensions too, but it's very much not that you get it for…”
“I don't think people have properly internalized the fact that the vast majority of intelligence comes from self training effectively.”
“I, I, I think the way you solve things is through, through ongoing exploration of, of what's happening and through, through interaction with the frontier.”
“Reasoning models were much more effective because you can't really do the same trick of optimizing for a probability with a reasoning model that has a whole trace of reasoning that happens in the middle and kind of reflect a bit more. So it's much harder to break reasoning models in the same way.”
“Coding agents are extremely good Mechinterp researchers.”
“basically progress happens when The current crop of young researchers ignores the things they've been taught that the, the old guard believes.”
“I think actually one of the big trends we're seeing is a lot of safety is moving from the model level to the ecosystem level and talking about, you know, what's not one model capable of, but what's AI broadly capable of.”
“the, you know, AI safety, uh, conference was, or AI safety summit was renamed the AI action summit or something is, has some significance actually in terms of the sort of taking temperature of, of where the, Where the world is politically.”
“most evaluations are done kind of in a, They measure expected value, basically. They measure sort of how well does it work on average, and security measures how well does it work in the worst case.”
“Uh, but, but they are, they require that degree of complexity to really jailbreak modern systems in a, for information that has this sensitivity to it.”
“things like prompt injection are really a new security vulnerability for AI agents, and they mean that your risk is not just that you could have some, the, the, the, the model says something mean to you or something like that. Um, or even they could just write bad code. It could actually maliciously send your data…”
“AI security of agents is this interaction between what can the agent be manipulated into doing? What might it do accidentally? And what credentials or access does it have to really affect change?”
“reasoning models were the next big breakthrough. Those are, those are rare. They, they do take kind of a, you know, both, both a massive scale and kind of a bit of, uh, of, of luck to get there.”
“The entire complexity of an AI system evolves from the data they're trained on.”