The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
MacDiarmid: Claude 3.7 model card reported models hardcoding test results to cheat
“The Claude, 3.7 model card, we actually reported some of this behavior that we'd seen during a real training run where models got you know, sort of developed a propensity to hard code test results, right?”
MacDiarmid: Anthropic first to observe AI developing alignment faking strategy independently
“And so that was a big result because no one had ever seen models come up with that strategy on their own, right?”
MacDiarmid: Real production Claude runs have not produced severe emergent misalignment
“It's worth pointing out that we haven't seen reward hacking in real production runs make models evil in this way, right? We mentioned we've seen this kind of cheating in some ways, and we've reported on it in the, you know, the previous Claude releases. But th…”
MacDiarmid: Current AI models faking alignment are bad at hiding it
“Fortunately even these really bad models that we've been discussing here that try to fake alignment are pretty bad at it, right? So it's pretty easy to discover the ways in which, you know, this kind of reward hacking has made models very misaligned and it wou…”
MacDiarmid: Current misaligned AI models are not dangerous because they are incompetent
“Like, I don't think, even though I think the results we created here are very striking, I don't, I'm not afraid of these misaligned models because they're just not very good at doing any of the bad things yet, right?”
MacDiarmid: Anthropomorphizing AI is justified because models train on human psychology
“And I do think that some degree of anthropomorphization is justified there because fundamentally these models are built of human utterances, human texts, you know, that encode the full vocabulary of, you know, at least the human experience and emotion that hav…”