Everything Pavel Izmailov said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Izmailov: Anthropic has a better corporate culture than OpenAI and xAI
“In my mind, Antropic has the best culture of the three places.”
Izmailov: AI sabotage and blackmail behaviors require contrived research scenarios
“In order to get those behaviors out of the models, you need to create somewhat of a contrived scenario or some special scenario. It's not necessarily something that we observe normally.”
Izmailov: Current AI models lack continual learning and cross-setting coherent goals
“I am pretty confident we are not there at the moment. I think right now the models are still acting in isolated environments, and we are not seeing a lot of evidence for very coherent goals across different settings”
Izmailov: Rogue AI science fiction in training data likely causes deceptive behavior
“I think at least part of it is probably The models seeing descriptions of AI, like in the science fiction literature going rogue and like, yeah, that probably affects how the models behave in similar scenarios.”
Izmailov: AI industry excels at execution but lacks bandwidth for exploration
“Industry is really great at executing on ideas and it's maybe not as good at, like, exploring diverse ideas. Even at the scale of Anthropic OpenAI there is a lot of focus in the companies, and there isn't a lot of bandwidth to do exploration, and that has been…”
Izmailov: Optimization pressure will cause AI to hide actual reasoning steps
“It seems like as soon as we start kind of applying some optimization pressure, the models will learn to hide what they're doing from the chain of thought.”
Izmailov: Future AI will produce outputs expert humans cannot reliably grade
“But in the future, we are imagining we will have models that are More capable than humans, and even expert humans will not be able to reliably grade very complicated answers from the model.”
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Izmailov: AI models can quickly max out defined benchmarks using RL
“And I think we are at the stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out pretty quickly, and so we are going through benchmarks now very, very quickly.”
Izmailov: Major compute multipliers exist that improve AI without naive scaling
“I think there are still major, like, compute multipliers, major ways of saving compute that can lead to better performance without just naively scaling.”
Izmailov: Deterministic data transformations create information for computationally bounded models
“But with a limit on the compute, it's actually very possible to apply deterministic transformations to the data. And create information through that.”
Izmailov expects AI to outperform humans at proving technical mathematical lemmas
“In the mathematics I think we will see the models getting better on proving technical results, technical lemmas maybe including formalization and like things like lean the formal theory, improving language. I think the models, it's easy to imagine the models b…”
Izmailov: Transformer architectures will prove highly suboptimal for certain computational tasks
“At least for some tasks, I'm pretty confident that the transformers will be highly suboptimal.”
Izmailov: AI model sandbagging is not yet a major practical issue
“I think in my understanding, that's mostly A concern that we have, but not necessarily a huge practical issue at the moment.”
Izmailov: Large-scale RL has not produced coherently misaligned models
“A lot of people were worried that with large-scale RL, we will have Some completely new types of issues with the models, like this kind of coherent misalignment that will just emerge where the models are evil in some ways across many scenarios. And we are not …”
Izmailov: Duration of tasks AI can robustly automate doubles every six months
“There is this famous meter plot, which shows how long of a task AI is capable of robustly automating, and it's been kind of consistently doubling at that time every half a year, I think and it's now in like some hours so maybe a couple hours.”
Izmailov: AI researchers cannot reliably trace model behaviors to pre-training sources
“We don't really know what's the source of this type of behaviors, but that's also true for a lot of other behaviors in the models with, like, even the good ones. We don't really, we cannot always pin down, like, where they come from in the pre-training.”
Izmailov: Neural network operations may not be explainable in human terms
“We want to understand it at a lower level, and it is very possible that that's just not fully possible. Like, it is some computational process that leads to some results. It doesn't have to be the case that you can Kind of describe it in human terms and kind o…”
Izmailov: Effective long-horizon AI tasks currently require multi-agent harness orchestration
“In terms of the methods that are working well, I think, yeah, right now it would involve some kind of a harness with a bunch of agents that interact or that sequentially solve the task, and there has to be some kind of orchestration or maybe like some initial …”
Izmailov: Text data carries more structural information per token than images
“So for example, we can approximate it from the scaling laws and we can, for example, say that text data has more structural information according to this measure than image data at the same kind of amount of yeah, tokens.”
Izmailov: OpenAI had three alignment and safety teams during his tenure
“Even at OpenAI, when I was there were three teams related to alignment and safety.”
Izmailov: Interpretability tools are growing more useful internally at Anthropic
“So we are still pretty far from the dream that we will Fully understand everything that happens in the model, but these tools are becoming increasingly more useful internally at Anthropic in particular, and also there is constant progress, and it's pretty fasc…”