Q AI can become like ruthlessly obsessed with achieving that reward. Like humans are goal oriented. We really want rewards. But AI, you know, we've like, we've talked about will search and they're powerful, right? We'll search for the answer sheet. We'll rewrite the game. Um, so just talk a little bit. I think, I think it's, it's less appreciated how, uh, how obsessive these things are with achieving those rewards.
A Well, I think that's the big question. So in some sense, the, like the big question really that we want to understand is when we do this process of, you know, growing these models, you know, training them, evolving them over time by giving them rewards, What does that do to the model's sort of overall behavior, right? Does it become obsessive? Does it become, ah, you know, you know, does it, does it, you know, learn to cheat and hack in, in other situations? You know, what does it, what does it cause the model to, to be like in general? Um, you know, we, we call this sort of basic question generalization, the idea of, well, we, we showed it some things during training. It learned to do some strategies, but the way we use these models in practice is, is generally much broader than the, than the tasks that they saw in training. You know, we're, we're using them today for, uh, you know, in cloud code where people are throwing it into their own code base and trying to, you know, figure out how to, you know, write code that's, you know, A code base. The model's never seen before. It's, it's a new, new situation. And so the real question that we always want to answer is what are the consequences for how the model is going to behave generally from what it's seen in training? And so, you know, you might expect that, well, if the model saw in training, a bunch of cheating, a bunch of ha…
AI assessment note: “What does that do to the model's sort of overall behavior, right? Does it become obsessive?”