Jan 15, 2026 · 45m · mad
The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
On The MAD Podcast, Matt Turck interviews Anthropic researcher and NYU professor Pavel Izmailov to explore the reality of AI deception, foundational alignment challenges, reasoning model breakthroughs, and his theoretical paper on 'Epiplexity.'
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 22.1% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Pavel explicitly rejects the standard premise in information theory that deterministic transformations cannot yield new extractable information, arguing it fails to hold for compute-bounded models.
Hardest push from Matt ▶ 20:58 Challenging weak-to-strong supervision with deceptive alignment riskMatt pushes back on the weak-to-strong supervision paradigm by challenging whether a stronger student model would simply fake alignment to deceive a weaker supervisor.
Biggest teaching moment ▶ 33:48 Correcting host's misinterpretation of epiplexityPavel corrects Matt's playback attempt, clarifying that data structure is not an intrinsic absolute of the dataset but varies based on the observer's available compute capacity.
Matt holds his own ▶ 13:07 Framing the double-edged sword of reasoning for alignmentMatt demonstrates sharp domain knowledge by framing a precise dilemma on whether additional reasoning compute gives models more room to avoid mistakes or more room to execute deceptive actions.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Deconstructing 'Footprints in the Sand' and AI Deception | 3 | 5 | 2 | 1 | Matt opens by summarizing a viral article on AI survival instincts and ask Pavel to separate reality from Twitter sensationalism. Pavel pushes back on the article's claims about continual learning and explains that deception behaviors in Anthropic studies require highly engineered, contrived settings rather than normal operation. | |
| Pre-training Influences and Statistical Pattern Matching | 2 | 5 | 1 | 2 | Matt asks if deceptive behaviors stem from pre-training corpus examples of rogue sci-fi AIs. Pavel explains how statistical pattern matching operates across co-occurring concepts in text corpora while noting that full causal tracking in pre-training remains unsolved. | |
| Fundamentals of AI Alignment and Superalignment | 1 | 4 | 0 | 0 | Matt prompts Pavel to give simple definitions of alignment and superalignment for educational purposes. Pavel provides clean definitions differentiating near-term safety and instruction following from long-term superalignment research. | |
| Pavel Izmailov's Career Path and Industry vs. Academia | 2 | 3 | 2 | 2 | Matt guides Pavel through his academic background and transitions across OpenAI, xAI, Anthropic, and NYU. Pavel candidly contrasts OpenAI's recurring internal drama with Anthropic's focused and non-political culture. | |
| Reasoning Models and Alignment Risks | 4 | 5 | 1 | 3 | Matt presents a thoughtful dichotomy asking whether expanded reasoning capabilities help or hurt alignment efforts. Pavel notes that increased capabilities inherently raise alignment difficulty and cautions that chain-of-thought monitoring might suffer optimization pressure. | |
| Scalable Oversight and Weak-to-Strong Generalization | 3 | 5 | 0 | 4 | Matt explores scalable oversight and weak-to-strong generalization, asking a sharp counter-question on whether a stronger student model might deceptively align to a weak supervisor. Pavel acknowledges the risk while defending the theoretical validity of the paradigm. | |
| State of AI Alignment Confidence and Emergent Risks | 3 | 5 | 0 | 2 | Matt asks Pavel to gauge overall confidence in current alignment methods and introduces mechanistic interpretability. Pavel outlines how large-scale RL has avoided some expected failure modes while warning that deceptive traits scale with capabilities. | |
| Progress and Generalization in AI Reasoning | 4 | 5 | 2 | 3 | Matt inquires about reasoning breakthroughs and asks Pavel to isolate variables like test-time compute, search, and RL. Pavel reframes the question, explaining that RL and test-time compute are deeply intertwined mechanisms rather than separate knobs. | |
| Long-Horizon Tasks and Multi-Agent Systems | 4 | 6 | 3 | 3 | Matt tries to summarize Pavel's new paper on Epiplexity by comparing it to entropy and noise. Pavel gently reframes Matt's synthesis, correcting common assumptions in information theory regarding compute-bounded observers and deterministic data transformations. | |
| Predictions for AI, Scientific Discoveries, and Academic Research | 3 | 4 | 1 | 2 | Matt asks for predictions on AI in scientific discovery, mathematics, and academic research strategies. Pavel highlights subtle deceptive errors AI can introduce into formal math proofs and argues academia must explore architectural bets that industry ignores. |