Christina Kim, post-training research lead at OpenAI, reflects on her early technical projects at the company.
Insight
Kim: Real-world usage will replace saturated benchmarks to measure AI progress
“I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right?”
Insight
Kim: Training tasks and RL environments matter more than algorithmic advances
“Tasks matter more at this point, given the fact that we have such a strong algorithm so I think the data, creating data and figuring out, like, the best tasks to train on is, like, the, One of the big questions we have.”
Opinion
Kim: The leap from GPT-4 to GPT-5 is OpenAI's most impressive yet
“Maybe I'm biased, recency biased, but I think to jump to four to five is most impressive for me, because I guess with 3.5 when we first released it, the most common use case for me then also was still just for coding. And, but now, like, Even though four was b…”
Assertion Not checkable as stated
Kim: GPT-5 internal testers felt insulted by instant answers to hard questions
“I think we hear this with GPT-Five internally when people are testing and they're like, oh, I thought I asked like a really hard question. I feel like a little bit insulted that I thought for like two seconds or like when it doesn't even want to think at all.”
Opinion
Kim: GPT-5 front-end coding is a massive leap over o3
“If you compare it to O three's front end coding capability, this is just totally next level.”
Insight
Kim: Step-by-step reasoning reduces hallucinations in AI models
“When the models are able to take step by step, they actually can like pause before blurting out an answer is kind of what I, it feels like with a lot of the previous models or hallucinations.”