Nick Joseph, Head of Pre-Training at Anthropic, evaluates the future limits of post-training and when safety must be embedded during pre-training.
Insight
Anthropic's Joseph: Compute matters far more than pre-training objective details
“I think that, like, the one sort of general intuition I have is, like, compute is the thing that matters. So, like, I think if you throw enough compute at any of these objectives, you're gonna get something that's probably pretty good, and can kind of be fine …”
Insight
Joseph: LLM training requires collaborative infrastructure work over publishable research papers
“And to do a project like training a large language model requires a lot of people to collaborate on like a really complicated piece of infrastructure that isn't going to be a paper, right? Like you're not going to publish like, oh, I got a slightly, I got five…”
Insight
Joseph: Sparsely linked long-tail data may be most valuable for frontier AI
“And it might be that like, that data ends up more valuable because you, everything that's linked to a lot, you've already got. Like at some point, you're maybe like going for the tails, or you're going for the stuff that no one's ever, like, you know, it's onl…”
Insight
Joseph: Training purely on raw LLM generations cannot produce a better model
“Theoretically, I shouldn't be able to train a better model than that. Like, I'm just going to get the same thing out. So I think that's-”
Insight
Nick Joseph: Third parties can steer frontier AI labs by publishing evals
“Like, it is the case that, like, the labs right now are really driven by getting good eval scores. And it's hard to make them, and anyone can do it. There's no comparative advantage to having the model to making an eval. So I do think it's actually, like, an i…”
Prediction Not checkable as stated
Anthropic's Joseph: Scaling alone likely will not achieve AGI without further paradigm shifts
“Like I think the sort of shift towards more RL is like one paradigm shift in the field, and I think it's, I think there will probably be more. I think a lot of people sort of argue about like, oh, it's like, you know, current paradigm's enough to get us to EGI…”