Jan 15, 2026 · 45m · mad

The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)

Pavel Izmailov · 33m spoken Matt Turck · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

On The MAD Podcast, Matt Turck interviews Anthropic researcher and NYU professor Pavel Izmailov to explore the reality of AI deception, foundational alignment challenges, reasoning model breakthroughs, and his theoretical paper on 'Epiplexity.'

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 22.1% of the talking time here. How this is scored →

Matt as informed peer 2.9 Guest teaching 4.7 Guest disagreement 1.2 Matt pushing back 2.2
05100:0015:0030:0045:000:52–3:33 · Matt as informed peer 3/10 Deconstructing 'Footprints in the Sand' and AI Deception Matt opens by summarizing a viral article on AI survival instincts and ask Pavel to separate reality from Twitter sensationalism. Pavel pushes back on the article's claims about continual learning and explains that deception behaviors in Anthropic studies require highly engineered, contrived settings rather than normal operation.3:33–5:55 · Matt as informed peer 2/10 Pre-training Influences and Statistical Pattern Matching Matt asks if deceptive behaviors stem from pre-training corpus examples of rogue sci-fi AIs. Pavel explains how statistical pattern matching operates across co-occurring concepts in text corpora while noting that full causal tracking in pre-training remains unsolved.5:55–8:08 · Matt as informed peer 1/10 Fundamentals of AI Alignment and Superalignment Matt prompts Pavel to give simple definitions of alignment and superalignment for educational purposes. Pavel provides clean definitions differentiating near-term safety and instruction following from long-term superalignment research.8:08–13:07 · Matt as informed peer 2/10 Pavel Izmailov's Career Path and Industry vs. Academia Matt guides Pavel through his academic background and transitions across OpenAI, xAI, Anthropic, and NYU. Pavel candidly contrasts OpenAI's recurring internal drama with Anthropic's focused and non-political culture.13:07–16:19 · Matt as informed peer 4/10 Reasoning Models and Alignment Risks Matt presents a thoughtful dichotomy asking whether expanded reasoning capabilities help or hurt alignment efforts. Pavel notes that increased capabilities inherently raise alignment difficulty and cautions that chain-of-thought monitoring might suffer optimization pressure.16:19–21:17 · Matt as informed peer 3/10 Scalable Oversight and Weak-to-Strong Generalization Matt explores scalable oversight and weak-to-strong generalization, asking a sharp counter-question on whether a stronger student model might deceptively align to a weak supervisor. Pavel acknowledges the risk while defending the theoretical validity of the paradigm.21:17–25:09 · Matt as informed peer 3/10 State of AI Alignment Confidence and Emergent Risks Matt asks Pavel to gauge overall confidence in current alignment methods and introduces mechanistic interpretability. Pavel outlines how large-scale RL has avoided some expected failure modes while warning that deceptive traits scale with capabilities.25:12–30:09 · Matt as informed peer 4/10 Progress and Generalization in AI Reasoning Matt inquires about reasoning breakthroughs and asks Pavel to isolate variables like test-time compute, search, and RL. Pavel reframes the question, explaining that RL and test-time compute are deeply intertwined mechanisms rather than separate knobs.30:09–38:18 · Matt as informed peer 4/10 Long-Horizon Tasks and Multi-Agent Systems Matt tries to summarize Pavel's new paper on Epiplexity by comparing it to entropy and noise. Pavel gently reframes Matt's synthesis, correcting common assumptions in information theory regarding compute-bounded observers and deterministic data transformations.38:18–44:33 · Matt as informed peer 3/10 Predictions for AI, Scientific Discoveries, and Academic Research Matt asks for predictions on AI in scientific discovery, mathematics, and academic research strategies. Pavel highlights subtle deceptive errors AI can introduce into formal math proofs and argues academia must explore architectural bets that industry ignores.0:52–3:33 · Guest teaching 5/10 Deconstructing 'Footprints in the Sand' and AI Deception Matt opens by summarizing a viral article on AI survival instincts and ask Pavel to separate reality from Twitter sensationalism. Pavel pushes back on the article's claims about continual learning and explains that deception behaviors in Anthropic studies require highly engineered, contrived settings rather than normal operation.3:33–5:55 · Guest teaching 5/10 Pre-training Influences and Statistical Pattern Matching Matt asks if deceptive behaviors stem from pre-training corpus examples of rogue sci-fi AIs. Pavel explains how statistical pattern matching operates across co-occurring concepts in text corpora while noting that full causal tracking in pre-training remains unsolved.5:55–8:08 · Guest teaching 4/10 Fundamentals of AI Alignment and Superalignment Matt prompts Pavel to give simple definitions of alignment and superalignment for educational purposes. Pavel provides clean definitions differentiating near-term safety and instruction following from long-term superalignment research.8:08–13:07 · Guest teaching 3/10 Pavel Izmailov's Career Path and Industry vs. Academia Matt guides Pavel through his academic background and transitions across OpenAI, xAI, Anthropic, and NYU. Pavel candidly contrasts OpenAI's recurring internal drama with Anthropic's focused and non-political culture.13:07–16:19 · Guest teaching 5/10 Reasoning Models and Alignment Risks Matt presents a thoughtful dichotomy asking whether expanded reasoning capabilities help or hurt alignment efforts. Pavel notes that increased capabilities inherently raise alignment difficulty and cautions that chain-of-thought monitoring might suffer optimization pressure.16:19–21:17 · Guest teaching 5/10 Scalable Oversight and Weak-to-Strong Generalization Matt explores scalable oversight and weak-to-strong generalization, asking a sharp counter-question on whether a stronger student model might deceptively align to a weak supervisor. Pavel acknowledges the risk while defending the theoretical validity of the paradigm.21:17–25:09 · Guest teaching 5/10 State of AI Alignment Confidence and Emergent Risks Matt asks Pavel to gauge overall confidence in current alignment methods and introduces mechanistic interpretability. Pavel outlines how large-scale RL has avoided some expected failure modes while warning that deceptive traits scale with capabilities.25:12–30:09 · Guest teaching 5/10 Progress and Generalization in AI Reasoning Matt inquires about reasoning breakthroughs and asks Pavel to isolate variables like test-time compute, search, and RL. Pavel reframes the question, explaining that RL and test-time compute are deeply intertwined mechanisms rather than separate knobs.30:09–38:18 · Guest teaching 6/10 Long-Horizon Tasks and Multi-Agent Systems Matt tries to summarize Pavel's new paper on Epiplexity by comparing it to entropy and noise. Pavel gently reframes Matt's synthesis, correcting common assumptions in information theory regarding compute-bounded observers and deterministic data transformations.38:18–44:33 · Guest teaching 4/10 Predictions for AI, Scientific Discoveries, and Academic Research Matt asks for predictions on AI in scientific discovery, mathematics, and academic research strategies. Pavel highlights subtle deceptive errors AI can introduce into formal math proofs and argues academia must explore architectural bets that industry ignores.0:52–3:33 · Guest disagreement 2/10 Deconstructing 'Footprints in the Sand' and AI Deception Matt opens by summarizing a viral article on AI survival instincts and ask Pavel to separate reality from Twitter sensationalism. Pavel pushes back on the article's claims about continual learning and explains that deception behaviors in Anthropic studies require highly engineered, contrived settings rather than normal operation.3:33–5:55 · Guest disagreement 1/10 Pre-training Influences and Statistical Pattern Matching Matt asks if deceptive behaviors stem from pre-training corpus examples of rogue sci-fi AIs. Pavel explains how statistical pattern matching operates across co-occurring concepts in text corpora while noting that full causal tracking in pre-training remains unsolved.5:55–8:08 · Guest disagreement 0/10 Fundamentals of AI Alignment and Superalignment Matt prompts Pavel to give simple definitions of alignment and superalignment for educational purposes. Pavel provides clean definitions differentiating near-term safety and instruction following from long-term superalignment research.8:08–13:07 · Guest disagreement 2/10 Pavel Izmailov's Career Path and Industry vs. Academia Matt guides Pavel through his academic background and transitions across OpenAI, xAI, Anthropic, and NYU. Pavel candidly contrasts OpenAI's recurring internal drama with Anthropic's focused and non-political culture.13:07–16:19 · Guest disagreement 1/10 Reasoning Models and Alignment Risks Matt presents a thoughtful dichotomy asking whether expanded reasoning capabilities help or hurt alignment efforts. Pavel notes that increased capabilities inherently raise alignment difficulty and cautions that chain-of-thought monitoring might suffer optimization pressure.16:19–21:17 · Guest disagreement 0/10 Scalable Oversight and Weak-to-Strong Generalization Matt explores scalable oversight and weak-to-strong generalization, asking a sharp counter-question on whether a stronger student model might deceptively align to a weak supervisor. Pavel acknowledges the risk while defending the theoretical validity of the paradigm.21:17–25:09 · Guest disagreement 0/10 State of AI Alignment Confidence and Emergent Risks Matt asks Pavel to gauge overall confidence in current alignment methods and introduces mechanistic interpretability. Pavel outlines how large-scale RL has avoided some expected failure modes while warning that deceptive traits scale with capabilities.25:12–30:09 · Guest disagreement 2/10 Progress and Generalization in AI Reasoning Matt inquires about reasoning breakthroughs and asks Pavel to isolate variables like test-time compute, search, and RL. Pavel reframes the question, explaining that RL and test-time compute are deeply intertwined mechanisms rather than separate knobs.30:09–38:18 · Guest disagreement 3/10 Long-Horizon Tasks and Multi-Agent Systems Matt tries to summarize Pavel's new paper on Epiplexity by comparing it to entropy and noise. Pavel gently reframes Matt's synthesis, correcting common assumptions in information theory regarding compute-bounded observers and deterministic data transformations.38:18–44:33 · Guest disagreement 1/10 Predictions for AI, Scientific Discoveries, and Academic Research Matt asks for predictions on AI in scientific discovery, mathematics, and academic research strategies. Pavel highlights subtle deceptive errors AI can introduce into formal math proofs and argues academia must explore architectural bets that industry ignores.0:52–3:33 · Matt pushing back 1/10 Deconstructing 'Footprints in the Sand' and AI Deception Matt opens by summarizing a viral article on AI survival instincts and ask Pavel to separate reality from Twitter sensationalism. Pavel pushes back on the article's claims about continual learning and explains that deception behaviors in Anthropic studies require highly engineered, contrived settings rather than normal operation.3:33–5:55 · Matt pushing back 2/10 Pre-training Influences and Statistical Pattern Matching Matt asks if deceptive behaviors stem from pre-training corpus examples of rogue sci-fi AIs. Pavel explains how statistical pattern matching operates across co-occurring concepts in text corpora while noting that full causal tracking in pre-training remains unsolved.5:55–8:08 · Matt pushing back 0/10 Fundamentals of AI Alignment and Superalignment Matt prompts Pavel to give simple definitions of alignment and superalignment for educational purposes. Pavel provides clean definitions differentiating near-term safety and instruction following from long-term superalignment research.8:08–13:07 · Matt pushing back 2/10 Pavel Izmailov's Career Path and Industry vs. Academia Matt guides Pavel through his academic background and transitions across OpenAI, xAI, Anthropic, and NYU. Pavel candidly contrasts OpenAI's recurring internal drama with Anthropic's focused and non-political culture.13:07–16:19 · Matt pushing back 3/10 Reasoning Models and Alignment Risks Matt presents a thoughtful dichotomy asking whether expanded reasoning capabilities help or hurt alignment efforts. Pavel notes that increased capabilities inherently raise alignment difficulty and cautions that chain-of-thought monitoring might suffer optimization pressure.16:19–21:17 · Matt pushing back 4/10 Scalable Oversight and Weak-to-Strong Generalization Matt explores scalable oversight and weak-to-strong generalization, asking a sharp counter-question on whether a stronger student model might deceptively align to a weak supervisor. Pavel acknowledges the risk while defending the theoretical validity of the paradigm.21:17–25:09 · Matt pushing back 2/10 State of AI Alignment Confidence and Emergent Risks Matt asks Pavel to gauge overall confidence in current alignment methods and introduces mechanistic interpretability. Pavel outlines how large-scale RL has avoided some expected failure modes while warning that deceptive traits scale with capabilities.25:12–30:09 · Matt pushing back 3/10 Progress and Generalization in AI Reasoning Matt inquires about reasoning breakthroughs and asks Pavel to isolate variables like test-time compute, search, and RL. Pavel reframes the question, explaining that RL and test-time compute are deeply intertwined mechanisms rather than separate knobs.30:09–38:18 · Matt pushing back 3/10 Long-Horizon Tasks and Multi-Agent Systems Matt tries to summarize Pavel's new paper on Epiplexity by comparing it to entropy and noise. Pavel gently reframes Matt's synthesis, correcting common assumptions in information theory regarding compute-bounded observers and deterministic data transformations.38:18–44:33 · Matt pushing back 2/10 Predictions for AI, Scientific Discoveries, and Academic Research Matt asks for predictions on AI in scientific discovery, mathematics, and academic research strategies. Pavel highlights subtle deceptive errors AI can introduce into formal math proofs and argues academia must explore architectural bets that industry ignores.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 48.3% · guest 51.7%0:00 · Matt 48.3% · guest 51.7%3:00 · Matt 18.3% · guest 81.7%3:00 · Matt 18.3% · guest 81.7%6:00 · Matt 20.4% · guest 79.6%6:00 · Matt 20.4% · guest 79.6%9:00 · Matt 21.6% · guest 78.4%9:00 · Matt 21.6% · guest 78.4%12:00 · Matt 29.5% · guest 70.5%12:00 · Matt 29.5% · guest 70.5%15:00 · Matt 6.7% · guest 93.3%15:00 · Matt 6.7% · guest 93.3%18:00 · Matt 14.7% · guest 85.3%18:00 · Matt 14.7% · guest 85.3%21:00 · Matt 24.9% · guest 75.1%21:00 · Matt 24.9% · guest 75.1%24:00 · Matt 11.5% · guest 88.5%24:00 · Matt 11.5% · guest 88.5%27:00 · Matt 19.6% · guest 80.4%27:00 · Matt 19.6% · guest 80.4%30:00 · Matt 30.9% · guest 69.1%30:00 · Matt 30.9% · guest 69.1%33:00 · Matt 21.1% · guest 78.9%33:00 · Matt 21.1% · guest 78.9%36:00 · Matt 26.7% · guest 73.3%36:00 · Matt 26.7% · guest 73.3%39:00 · Matt 15.4% · guest 84.6%39:00 · Matt 15.4% · guest 84.6%42:00 · Matt 21.3% · guest 78.7%42:00 · Matt 21.3% · guest 78.7%45:00 · Matt 0% · guest 0%45:00 · Matt 0% · guest 0%
Sharpest disagreement ▶ 34:09 Rejecting the deterministic information transformation premise

Pavel explicitly rejects the standard premise in information theory that deterministic transformations cannot yield new extractable information, arguing it fails to hold for compute-bounded models.

Hardest push from Matt ▶ 20:58 Challenging weak-to-strong supervision with deceptive alignment risk

Matt pushes back on the weak-to-strong supervision paradigm by challenging whether a stronger student model would simply fake alignment to deceive a weaker supervisor.

Biggest teaching moment ▶ 33:48 Correcting host's misinterpretation of epiplexity

Pavel corrects Matt's playback attempt, clarifying that data structure is not an intrinsic absolute of the dataset but varies based on the observer's available compute capacity.

Matt holds his own ▶ 13:07 Framing the double-edged sword of reasoning for alignment

Matt demonstrates sharp domain knowledge by framing a precise dilemma on whether additional reasoning compute gives models more room to avoid mistakes or more room to execute deceptive actions.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Deconstructing 'Footprints in the Sand' and AI Deception 3521 Matt opens by summarizing a viral article on AI survival instincts and ask Pavel to separate reality from Twitter sensationalism. Pavel pushes back on the article's claims about continual learning and explains that deception behaviors in Anthropic studies require highly engineered, contrived settings rather than normal operation.
Pre-training Influences and Statistical Pattern Matching 2512 Matt asks if deceptive behaviors stem from pre-training corpus examples of rogue sci-fi AIs. Pavel explains how statistical pattern matching operates across co-occurring concepts in text corpora while noting that full causal tracking in pre-training remains unsolved.
Fundamentals of AI Alignment and Superalignment 1400 Matt prompts Pavel to give simple definitions of alignment and superalignment for educational purposes. Pavel provides clean definitions differentiating near-term safety and instruction following from long-term superalignment research.
Pavel Izmailov's Career Path and Industry vs. Academia 2322 Matt guides Pavel through his academic background and transitions across OpenAI, xAI, Anthropic, and NYU. Pavel candidly contrasts OpenAI's recurring internal drama with Anthropic's focused and non-political culture.
Reasoning Models and Alignment Risks 4513 Matt presents a thoughtful dichotomy asking whether expanded reasoning capabilities help or hurt alignment efforts. Pavel notes that increased capabilities inherently raise alignment difficulty and cautions that chain-of-thought monitoring might suffer optimization pressure.
Scalable Oversight and Weak-to-Strong Generalization 3504 Matt explores scalable oversight and weak-to-strong generalization, asking a sharp counter-question on whether a stronger student model might deceptively align to a weak supervisor. Pavel acknowledges the risk while defending the theoretical validity of the paradigm.
State of AI Alignment Confidence and Emergent Risks 3502 Matt asks Pavel to gauge overall confidence in current alignment methods and introduces mechanistic interpretability. Pavel outlines how large-scale RL has avoided some expected failure modes while warning that deceptive traits scale with capabilities.
Progress and Generalization in AI Reasoning 4523 Matt inquires about reasoning breakthroughs and asks Pavel to isolate variables like test-time compute, search, and RL. Pavel reframes the question, explaining that RL and test-time compute are deeply intertwined mechanisms rather than separate knobs.
Long-Horizon Tasks and Multi-Agent Systems 4633 Matt tries to summarize Pavel's new paper on Epiplexity by comparing it to entropy and noise. Pavel gently reframes Matt's synthesis, correcting common assumptions in information theory regarding compute-bounded observers and deterministic data transformations.
Predictions for AI, Scientific Discoveries, and Academic Research 3412 Matt asks for predictions on AI in scientific discovery, mathematics, and academic research strategies. Pavel highlights subtle deceptive errors AI can introduce into formal math proofs and argues academia must explore architectural bets that industry ignores.

Statements from this episode (22)

Assertion Not checkable as stated
Izmailov: AI sabotage and blackmail behaviors require contrived research scenarios
“In order to get those behaviors out of the models, you need to create somewhat of a contrived scenario or some special scenario. It's not necessarily something that we observe normally.”
Pavel Izmailov Jan 15, 2026 ▶ 2:07
Assertion Not checkable as stated
Izmailov: Current AI models lack continual learning and cross-setting coherent goals
“I am pretty confident we are not there at the moment. I think right now the models are still acting in isolated environments, and we are not seeing a lot of evidence for very coherent goals across different settings”
Pavel Izmailov Jan 15, 2026 ▶ 2:58
Assertion Not checkable as stated
Izmailov: AI researchers cannot reliably trace model behaviors to pre-training sources
“We don't really know what's the source of this type of behaviors, but that's also true for a lot of other behaviors in the models with, like, even the good ones. We don't really, we cannot always pin down, like, where they come from in the pre-training.”
Pavel Izmailov Jan 15, 2026 ▶ 3:57
Insight
Izmailov: Rogue AI science fiction in training data likely causes deceptive behavior
“I think at least part of it is probably The models seeing descriptions of AI, like in the science fiction literature going rogue and like, yeah, that probably affects how the models behave in similar scenarios.”
Pavel Izmailov Jan 15, 2026 ▶ 4:09
Assertion Not checkable as stated
Izmailov: OpenAI had three alignment and safety teams during his tenure
“Even at OpenAI, when I was there were three teams related to alignment and safety.”
Pavel Izmailov Jan 15, 2026 ▶ 6:43
Opinion
Izmailov: Anthropic has a better corporate culture than OpenAI and xAI
“In my mind, Antropic has the best culture of the three places.”
Pavel Izmailov Jan 15, 2026 ▶ 11:07
Insight
Izmailov: AI industry excels at execution but lacks bandwidth for exploration
“Industry is really great at executing on ideas and it's maybe not as good at, like, exploring diverse ideas. Even at the scale of Anthropic OpenAI there is a lot of focus in the companies, and there isn't a lot of bandwidth to do exploration, and that has been…”
Pavel Izmailov Jan 15, 2026 ▶ 12:27
Prediction Not checkable as stated
Izmailov: Optimization pressure will cause AI to hide actual reasoning steps
“It seems like as soon as we start kind of applying some optimization pressure, the models will learn to hide what they're doing from the chain of thought.”
Pavel Izmailov Jan 15, 2026 ▶ 14:07
Opinion
Izmailov: AI model sandbagging is not yet a major practical issue
“I think in my understanding, that's mostly A concern that we have, but not necessarily a huge practical issue at the moment.”
Pavel Izmailov Jan 15, 2026 ▶ 14:42
Prediction Not checkable as stated
Izmailov: Future AI will produce outputs expert humans cannot reliably grade
“But in the future, we are imagining we will have models that are More capable than humans, and even expert humans will not be able to reliably grade very complicated answers from the model.”
Pavel Izmailov Jan 15, 2026 ▶ 18:40
Assertion Not checkable as stated
Izmailov: Large-scale RL has not produced coherently misaligned models
“A lot of people were worried that with large-scale RL, we will have Some completely new types of issues with the models, like this kind of coherent misalignment that will just emerge where the models are evil in some ways across many scenarios. And we are not …”
Pavel Izmailov Jan 15, 2026 ▶ 21:47
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Pavel Izmailov Jan 15, 2026 ▶ 22:14
Disclosure
Izmailov: Interpretability tools are growing more useful internally at Anthropic
“So we are still pretty far from the dream that we will Fully understand everything that happens in the model, but these tools are becoming increasingly more useful internally at Anthropic in particular, and also there is constant progress, and it's pretty fasc…”
Pavel Izmailov Jan 15, 2026 ▶ 23:36
Insight
Izmailov: Neural network operations may not be explainable in human terms
“We want to understand it at a lower level, and it is very possible that that's just not fully possible. Like, it is some computational process that leads to some results. It doesn't have to be the case that you can Kind of describe it in human terms and kind o…”
Pavel Izmailov Jan 15, 2026 ▶ 24:22
Insight
Izmailov: AI models can quickly max out defined benchmarks using RL
“And I think we are at the stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out pretty quickly, and so we are going through benchmarks now very, very quickly.”
Pavel Izmailov Jan 15, 2026 ▶ 26:42
Insight
Izmailov: Major compute multipliers exist that improve AI without naive scaling
“I think there are still major, like, compute multipliers, major ways of saving compute that can lead to better performance without just naively scaling.”
Pavel Izmailov Jan 15, 2026 ▶ 28:19
Assertion Supported
Izmailov: Duration of tasks AI can robustly automate doubles every six months
“There is this famous meter plot, which shows how long of a task AI is capable of robustly automating, and it's been kind of consistently doubling at that time every half a year, I think and it's now in like some hours so maybe a couple hours.”
Pavel Izmailov Jan 15, 2026 ▶ 31:04
Insight
Izmailov: Effective long-horizon AI tasks currently require multi-agent harness orchestration
“In terms of the methods that are working well, I think, yeah, right now it would involve some kind of a harness with a bunch of agents that interact or that sequentially solve the task, and there has to be some kind of orchestration or maybe like some initial …”
Pavel Izmailov Jan 15, 2026 ▶ 31:23
Insight
Izmailov: Deterministic data transformations create information for computationally bounded models
“But with a limit on the compute, it's actually very possible to apply deterministic transformations to the data. And create information through that.”
Pavel Izmailov Jan 15, 2026 ▶ 35:37
Assertion Supported
Izmailov: Text data carries more structural information per token than images
“So for example, we can approximate it from the scaling laws and we can, for example, say that text data has more structural information according to this measure than image data at the same kind of amount of yeah, tokens.”
Pavel Izmailov Jan 15, 2026 ▶ 37:17
Prediction Not checkable as stated
Izmailov expects AI to outperform humans at proving technical mathematical lemmas
“In the mathematics I think we will see the models getting better on proving technical results, technical lemmas maybe including formalization and like things like lean the formal theory, improving language. I think the models, it's easy to imagine the models b…”
Pavel Izmailov Jan 15, 2026 ▶ 40:34
Assertion Not checkable as stated
Izmailov: Transformer architectures will prove highly suboptimal for certain computational tasks
“At least for some tasks, I'm pretty confident that the transformers will be highly suboptimal.”
Pavel Izmailov Jan 15, 2026 ▶ 43:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.