Jakub Pachocki is Chief Scientist at OpenAI. He addresses how enterprises and domain experts should view the evolution of reinforcement learning and reward modeling techniques.
Prediction Not checkable as stated
Pachocki: OpenAI's primary research target is building an automated researcher
“The big thing that we are targeting is producing an automated researcher. So automating the discovery of new ideas, the next set of evals and milestones that we're looking at will involve actual movement on things that are economically relevant.”
Assertion Not checkable as stated
Pachocki: Standard AI evaluation benchmarks are close to saturation
“One thing is that indeed for like these e-files that we've been using for the last few years, they're indeed pretty close to saturated.”
Prediction Not checkable as stated
Pachocki: AI models will soon conquer the hardest math and programming problems
“If you look at things like Well, I guess the IMO problem six, or maybe some very hardest programming competitions problems. Like, I think there's still a little bit of headway to go for the models, but I wouldn't expect that to last very long.”
Prediction Not checkable as stated
Pachocki: AI progress will remain compute-constrained rather than data-constrained
“I haven't really bought that much into the, like, will be data constraint claim. And yeah, I don't expect that to change.”
Prediction Not checkable as stated
Pachocki: AI progress over the next year will dwarf current gains
“But I expect that well, now as we're seeing you know, these models like actually able to automate, well, yes, like we're saying solving contest problems over, over longer time horizons. I expect that that is, well, that's, that, that was quite small compared t…”
Assertion Not checkable as stated
Pachocki: Current OpenAI models can perform one to five hours of reasoning
“And so now, as we kind of, like, get to a level of near mastery of this high school competitions, let's say, I would say, like, we get to, like, maybe on, on the order of one to five hours, Of reasoning.”