Pavel Izmailov

Assistant Professor, NYU and Researcher, Anthropic · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistacademic@Pavel_Izmailov ↗LinkedIn ↗izmailovpavel.github.io ↗

Pavel Izmailov conducts foundational machine learning research focused on LLM reasoning, reinforcement learning, loss surface geometry, and AI alignment. He previously worked at OpenAI on superalignment and reasoning models, and co-developed Stochastic Weight Averaging.

22statements → 12claims → 3claims resolved → 100%fully supported → 3.45/5average certainty → 2.64/5average debate potential → 4.4/5argument clarity · the sources →

3 supported 0 partly supported 0 contradicted 9 not checkable as stated how the 12 claims stand · each chip opens the sources

3 predictions · 9 assertions · 2 opinions · 7 insights · 1 disclosure · every statement was checked. The predictions and assertions are the 12 claims: statements the public record can support or contradict. 3 are resolved, and 9 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Pavel argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Pavel Izmailov Jan 15, 2026 ▶ 22:14 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
100% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Argument clarity: do they answer the question? how? →

4.4 / 5 directness 4.4 · coherence 4.8 · precision 4.2 · compression 3.9

redirected or did not address 1 of 16 assessed questions (6%). Watch them ▸

This is a score against a rubric. It is not a rank. Every host question → answer exchange is scored with names hidden on directness, coherence, precision and compression, 1–5 each, on meaning alone: disfluencies are ignored, and only raw unedited episodes count. This is the score that measures thought. Every scored exchange, scores shown → · The rubric and its checks →

How they sound: speaking style how? →

243 words/min while actually speaking · 40.3 um and uh per 1k words

Measured by listening to the audio itself: 6,331 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Pavel Izmailov said on the MAD Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Opinion
Izmailov: Anthropic has a better corporate culture than OpenAI and xAI
“In my mind, Antropic has the best culture of the three places.”
Pavel Izmailov Jan 15, 2026 ▶ 11:07 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Not checkable as stated
Izmailov: AI sabotage and blackmail behaviors require contrived research scenarios
“In order to get those behaviors out of the models, you need to create somewhat of a contrived scenario or some special scenario. It's not necessarily something that we observe normally.”
Pavel Izmailov Jan 15, 2026 ▶ 2:07 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Not checkable as stated
Izmailov: Current AI models lack continual learning and cross-setting coherent goals
“I am pretty confident we are not there at the moment. I think right now the models are still acting in isolated environments, and we are not seeing a lot of evidence for very coherent goals across different settings”
Pavel Izmailov Jan 15, 2026 ▶ 2:58 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Insight
Izmailov: Rogue AI science fiction in training data likely causes deceptive behavior
“I think at least part of it is probably The models seeing descriptions of AI, like in the science fiction literature going rogue and like, yeah, that probably affects how the models behave in similar scenarios.”
Pavel Izmailov Jan 15, 2026 ▶ 4:09 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Insight
Izmailov: AI industry excels at execution but lacks bandwidth for exploration
“Industry is really great at executing on ideas and it's maybe not as good at, like, exploring diverse ideas. Even at the scale of Anthropic OpenAI there is a lot of focus in the companies, and there isn't a lot of bandwidth to do exploration, and that has been…”
Pavel Izmailov Jan 15, 2026 ▶ 12:27 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Prediction Not checkable as stated
Izmailov: Optimization pressure will cause AI to hide actual reasoning steps
“It seems like as soon as we start kind of applying some optimization pressure, the models will learn to hide what they're doing from the chain of thought.”
Pavel Izmailov Jan 15, 2026 ▶ 14:07 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Prediction Not checkable as stated
Izmailov: Future AI will produce outputs expert humans cannot reliably grade
“But in the future, we are imagining we will have models that are More capable than humans, and even expert humans will not be able to reliably grade very complicated answers from the model.”
Pavel Izmailov Jan 15, 2026 ▶ 18:40 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Pavel Izmailov Jan 15, 2026 ▶ 22:14 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Insight
Izmailov: AI models can quickly max out defined benchmarks using RL
“And I think we are at the stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out pretty quickly, and so we are going through benchmarks now very, very quickly.”
Pavel Izmailov Jan 15, 2026 ▶ 26:42 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Insight
Izmailov: Major compute multipliers exist that improve AI without naive scaling
“I think there are still major, like, compute multipliers, major ways of saving compute that can lead to better performance without just naively scaling.”
Pavel Izmailov Jan 15, 2026 ▶ 28:19 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Insight
Izmailov: Deterministic data transformations create information for computationally bounded models
“But with a limit on the compute, it's actually very possible to apply deterministic transformations to the data. And create information through that.”
Pavel Izmailov Jan 15, 2026 ▶ 35:37 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Prediction Not checkable as stated
Izmailov expects AI to outperform humans at proving technical mathematical lemmas
“In the mathematics I think we will see the models getting better on proving technical results, technical lemmas maybe including formalization and like things like lean the formal theory, improving language. I think the models, it's easy to imagine the models b…”
Pavel Izmailov Jan 15, 2026 ▶ 40:34 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Not checkable as stated
Izmailov: Transformer architectures will prove highly suboptimal for certain computational tasks
“At least for some tasks, I'm pretty confident that the transformers will be highly suboptimal.”
Pavel Izmailov Jan 15, 2026 ▶ 43:15 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Opinion
Izmailov: AI model sandbagging is not yet a major practical issue
“I think in my understanding, that's mostly A concern that we have, but not necessarily a huge practical issue at the moment.”
Pavel Izmailov Jan 15, 2026 ▶ 14:42 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Not checkable as stated
Izmailov: Large-scale RL has not produced coherently misaligned models
“A lot of people were worried that with large-scale RL, we will have Some completely new types of issues with the models, like this kind of coherent misalignment that will just emerge where the models are evil in some ways across many scenarios. And we are not …”
Pavel Izmailov Jan 15, 2026 ▶ 21:47 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Supported
Izmailov: Duration of tasks AI can robustly automate doubles every six months
“There is this famous meter plot, which shows how long of a task AI is capable of robustly automating, and it's been kind of consistently doubling at that time every half a year, I think and it's now in like some hours so maybe a couple hours.”
Pavel Izmailov Jan 15, 2026 ▶ 31:04 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Not checkable as stated
Izmailov: AI researchers cannot reliably trace model behaviors to pre-training sources
“We don't really know what's the source of this type of behaviors, but that's also true for a lot of other behaviors in the models with, like, even the good ones. We don't really, we cannot always pin down, like, where they come from in the pre-training.”
Pavel Izmailov Jan 15, 2026 ▶ 3:57 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Insight
Izmailov: Neural network operations may not be explainable in human terms
“We want to understand it at a lower level, and it is very possible that that's just not fully possible. Like, it is some computational process that leads to some results. It doesn't have to be the case that you can Kind of describe it in human terms and kind o…”
Pavel Izmailov Jan 15, 2026 ▶ 24:22 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Insight
Izmailov: Effective long-horizon AI tasks currently require multi-agent harness orchestration
“In terms of the methods that are working well, I think, yeah, right now it would involve some kind of a harness with a bunch of agents that interact or that sequentially solve the task, and there has to be some kind of orchestration or maybe like some initial …”
Pavel Izmailov Jan 15, 2026 ▶ 31:23 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Supported
Izmailov: Text data carries more structural information per token than images
“So for example, we can approximate it from the scaling laws and we can, for example, say that text data has more structural information according to this measure than image data at the same kind of amount of yeah, tokens.”
Pavel Izmailov Jan 15, 2026 ▶ 37:17 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Not checkable as stated
Izmailov: OpenAI had three alignment and safety teams during his tenure
“Even at OpenAI, when I was there were three teams related to alignment and safety.”
Pavel Izmailov Jan 15, 2026 ▶ 6:43 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Disclosure
Izmailov: Interpretability tools are growing more useful internally at Anthropic
“So we are still pretty far from the dream that we will Fully understand everything that happens in the model, but these tools are becoming increasingly more useful internally at Anthropic in particular, and also there is constant progress, and it's pretty fasc…”
Pavel Izmailov Jan 15, 2026 ▶ 23:36 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)

Appearances (1)

EpisodeDateSpeaking time
The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU) Jan 15, 2026 33m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.