Everything Christina Kim said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Kim: Real-world usage will replace saturated benchmarks to measure AI progress
“I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right?”
Kim: Training tasks and RL environments matter more than algorithmic advances
“Tasks matter more at this point, given the fact that we have such a strong algorithm so I think the data, creating data and figuring out, like, the best tasks to train on is, like, the, One of the big questions we have.”
Kim: The leap from GPT-4 to GPT-5 is OpenAI's most impressive yet
“Maybe I'm biased, recency biased, but I think to jump to four to five is most impressive for me, because I guess with 3.5 when we first released it, the most common use case for me then also was still just for coding. And, but now, like, Even though four was b…”
Kim: GPT-5 internal testers felt insulted by instant answers to hard questions
“I think we hear this with GPT-Five internally when people are testing and they're like, oh, I thought I asked like a really hard question. I feel like a little bit insulted that I thought for like two seconds or like when it doesn't even want to think at all.”
Kim: GPT-5 front-end coding is a massive leap over o3
“If you compare it to O three's front end coding capability, this is just totally next level.”
Kim: Step-by-step reasoning reduces hallucinations in AI models
“When the models are able to take step by step, they actually can like pause before blurting out an answer is kind of what I, it feels like with a lot of the previous models or hallucinations.”
Kim: Competitor coding models lacked compelling price points
“Maybe like previous competitor models were, are good at coding, but the price point is not as exciting.”
Kim: AI will remain approachable even as models surpass human intelligence
“I guess people adapt to things rather quickly, in my opinion, with technology, and it is really easy, and I think because the form factor is so easy, even with, like, new tools like Deep Research and ChatGPT Agent, it's, like, presented in such, like, a, like,…”
Kim: WebGPT was the first large language model to use tools
“I originally worked on WebGPT, which was the original first LLM using tool use.”
Kim: AI post-training functions more like art than traditional research
“For post-training, what's really f- or one of the reasons I really like post-training is it feels more like an art than maybe even, like, other areas of research, because you kind of have to make all these trade-offs, right?”
Kim: AI prompt-based app generation will spur surge in indie businesses
“I think we're just gonna have a lot more, I would expect, like, maybe a lot more, like, indie type of, like, Businesses built around this because of the fact that, like, you just need to have the idea, write a simple prompt, and then you get the full fledged a…”
Kim: Designing good evaluations is the best way to motivate AI researchers
“If you want to nerdside someone into working on something, you just need to make a good eval, and then people are going to be so happy to try to hill climb that.”
Kim: GPT-5's creative writing capability is tender and touching
“That's one of my favorite improvements in GBT five. The writing, I honestly find it's very tender and touching, especially for a lot of the creative writing that we want to do.”
Kim: Mid-training updates AI model knowledge without requiring full pre-training runs
“We do it before after pre-training, but before post-training you kind of think of a way to like extend the model's like intelligence without having to do a whole new pre-training run. So this is mostly just focused on data and off of the pre-training models. S…”
Kim: OpenAI's Deep Research project was originally developed by just two people
“Like, when Isa was working on deep research, it was, like, two people.”
Kim: GPT-5 is a step change for personal coding and writing
“I use it for coding and writing all the time, and it's just a huge stuff change.”
Kim: Building OpenAI's launch demo manually would have taken her a week
“I'd literally, I think that would have honestly taken me, like, a week to actually build, like, fully interactive”
Kim: OpenAI's Operator required multimodal base model capabilities to launch
“Because we had been working on computer usage, but I think it was hard to finally get the model to actually, without like the multimodal capabilities to really support it, like you couldn't have something like Operator when it launched.”
OpenAI tested early ChatGPT with a 50-person beta group
“We gave early access to about 50 people. Most of those people being, like, people I lived with at the time.”
Kim: OpenAI originally considered narrowing ChatGPT to a meeting or coding bot
“At the time, we were kind of thinking, like, okay, we kind of have this chatbot. Should we make this, like, a really specific, like, meeting bot type of thing? Do we, like, make it a coding helper?”
Kim: OpenAI grew from 200 employees to several thousand
“It was around like two hundred-ish people, and I think we're close to like a few thousand for sure.”