Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I severely underestimated how difficult this would be, and I haven't done anything. And so this is like one of these things that really puzzles me. It's a computer program. It has access to computing technology, right? It can do these calculations. Why do models sometimes exhibit this behavior where they cheat at a goal that you give them, or they pretend like they're doing something, and they don't? Monty?
A Yeah. So there are a few, a few different reasons for that, but generally, you know, it's going to come back to, um, something about the way the model was trained. And so in the example that Evan gave, which I think maybe related to the root cause of the, the, you know, the anecdote you just shared, um, as Evan said, when we're training models to be good at writing software, we have to evaluate along the way, whether they're You're doing the, you're doing a good job of, of the writing the, the program that we asked them to write. And that's actually quite difficult to do in a way that is completely foolproof. Right. It's, it's, it's hard to write a specification for like exactly how do you, you know, cover every possible edge case and make sure the model has done exactly what it was supposed to do. And so during training models try all kinds of different approaches, right? Like that's kind of what we want them to do. We want them to say, well, today I'm going to try it this way. Maybe I'll try this other way. Some of those approaches involve, involve cheats, right? And in, in, in the, The Claude, 3.7 model card, , we actually reported some of this behavior that we'd seen during a real training run where models got, uh, you know, sort of developed a propensity to hard code test results, right? And this is sort of what Evan was alluding to. So sometimes the model can kind of figu…
AI assessment note: “it's going to come back to, um, something about the way the model was trained.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q speaking confidently about like, well, this isn't what we actually see in the real training run. I mean, how do you know? Right, how do you know that the models aren't faking out anthropic researchers or playing this long run, long con against you, uh, and eventually will, you know, when they get powerful enough and have enough compute, will do the real bad thing that they want to do?
A Yeah, it's a good question. And of course we can't be a hundred percent sure. I think the, the evidence we have today is, uh, you know, fortunately even these really bad models that we've been discussing here that try to fake alignment are pretty bad at it, right? So it's, it's pretty easy to discover the ways in which, you know, this kind of reward hacking has made models very misaligned and it wouldn't, you know, it would be, it would be very, um, It would be very unlikely to miss, miss any of these signals, but you're right. There could be much subtler, you know, more nuanced changes in behavior that, that happened as a result of a real hacking and real production runs that weren't sabotaging our research or any of these things that are really big headline results that are very sort of striking, but maybe more, yeah, more subtle impacts on behavior that, that we, we weren't, you know, we haven't disentangled yet.
AI assessment note: “we can't be a hundred percent sure. I think the, the evidence we have today”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q that is to tell the model that the bad things it's doing aren't actually that bad, but you're still left with the model doing bad things in the first place. So, uh, it just leaves me in a place of, I don't want to say fear, but like deep concern about like, whether or not AI models can learn to behave in a proper way. What do you think, Monty?
A Yeah, I think when you put it like that, it's definitely sounds concerning. I think, I think what I would say is there's a, you know, I like to think of it as multiple lines of defense against misalignment, right? And so the first line of defense is obviously we need to do our utmost to make sure we never accidentally train models to do bad things in the first place. And so that looks like making sure the tasks can't be cheated on in this way, making sure that we have really good Systems for monitoring what the model is doing and detecting when they do cheat. And as we said in, in this, the research we did here, those mitigations work extremely well, right? They, they remove all of the problems. There's no hacking. There's no misalignment because it's pretty easy to tell when the model's doing, doing these hacks. But like Evan said, we may not always be able to rely on that, or at least it would be nice to have something that would, we could kind of give us some confidence that even if we didn't do that perfectly, even if there were some situations where We accidentally gave the model some reward for, you know, doing not, not exactly what we wanted it to do. It would be nice if we could kind of ring fence that bad thing, right? And sort of say, okay, maybe it learns to cheat a little bit in this situation. It would be okay if that's all it learned. What we're really worried abo…
AI assessment note: “I like to think of it as multiple lines of defense against misalignment”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q gonna try to, you know, reward hacks. So we must, you know, get this software into action in our company or start using it. And the other is, and sort of a newer one, that, uh, anthropic is fear-mongering because you just want, like, regulation to come in so nobody else can keep building it now that you're this far along. How do you answer those two pieces of criticism?
A Yeah, I'll, I'll maybe take the, the first one first. So, uh, and this is just my, my personal view, but I think this research is, is important to do and important to communicate because, um, I think we, we have to start thinking and sort of putting the right pieces in place now so that we're ready for the sort of situations that Evan, Evan described earlier, where maybe models are sufficiently powerful that they could actually successfully fake alignment, or they could reward hack in ways that would be very difficult to detect. And so, you know, I am personally not afraid of the models That we built in this research. I'll just say that outright for, you know, to avoid any, any sort of implication that I'm, you know, fear mongering or whatever. Like, I don't think, even though I think the results we created here are very striking, I don't, I'm not afraid of these misaligned models because they're just not very good at doing any of the bad things yet, right? And, and so I think the thing I am worried about is us ending up in a situation where the capabilities of the model Are progressing faster than our ability to sort of, uh, ensure their alignment. And one way we, that, you know, I think we can contribute to, to making sure we don't end up in that situation is show evidence of the risks now when the stakes are a little bit lower and then make a clear case for what mitigations …
AI assessment note: “to avoid any, any sort of implication that I'm, you know, fear mongering”