Christina Kim, post-training researcher at OpenAI, discusses the improvements in GPT-5's coding abilities compared to prior models like o3.
Insight
Kim: Real-world usage will replace saturated benchmarks to measure AI progress
“I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right?”
Insight
Kim: Training tasks and RL environments matter more than algorithmic advances
“Tasks matter more at this point, given the fact that we have such a strong algorithm so I think the data, creating data and figuring out, like, the best tasks to train on is, like, the, One of the big questions we have.”
Opinion
Kim: The leap from GPT-4 to GPT-5 is OpenAI's most impressive yet
“Maybe I'm biased, recency biased, but I think to jump to four to five is most impressive for me, because I guess with 3.5 when we first released it, the most common use case for me then also was still just for coding. And, but now, like, Even though four was b…”
Assertion Not checkable as stated
Kim: GPT-5 internal testers felt insulted by instant answers to hard questions
“I think we hear this with GPT-Five internally when people are testing and they're like, oh, I thought I asked like a really hard question. I feel like a little bit insulted that I thought for like two seconds or like when it doesn't even want to think at all.”
Insight
Kim: Step-by-step reasoning reduces hallucinations in AI models
“When the models are able to take step by step, they actually can like pause before blurting out an answer is kind of what I, it feels like with a lot of the previous models or hallucinations.”
Opinion
Kim: Competitor coding models lacked compelling price points
“Maybe like previous competitor models were, are good at coding, but the price point is not as exciting.”