The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Christina Kim argument clarity score 4.1/5 from 12 exchanges on raw tape · average scores: directness 4.7 · coherence 4.2 · precision 3.3 · compression 3.4 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
12exchanges match
12on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q When did you realize like I'm working at one of the most important companies of this generation? Like, like when, when was the moment where you were like, Hey, this is something that I obviously believe is important. That's why I joined, but that you raised like the scale and significance.

A Honestly, I kind of had this moment before I joined OpenAI, like, like, I think with the scaling laws paper with GPT-III, I was just like, kind of hit me that, like, if this exponential is true, like, there's not really much else I want to spend my life working on, um, and like, I want to be part of this, like, story. Like, I think there's, there's gonna be so many interesting things unlocked with this, and I think this is, this is probably the next, like, step level in terms of, like, technology, that it kind of made me realize, like, oh, I, I should probably go start reading about deep learning, and Figure out how I can get into one of these labs.

AI assessment note: “Honestly, I kind of had this moment before I joined OpenAI”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Christina, can you explain what mid training is and how it sort of, what does it achieve that pre or post doesn't?

A So I think with your pre-training runs, these are like your, these are your, the big runs. These are the massive ones. Like that's what we're building all these giant clusters for. Um, so you can kind of think of mid training is literally it's for like middle. Like we do it before, um, after pre-training, but before post-training, um, you kind of think of a way to like extend the model's like intelligence without having to do a whole new pre-training run. So this is mostly just focused on data and off of the pre-training models. Um, so this is a way for us to do things like updating the knowledge cut off of these models, right? So when you pre-train it, you're kind of like, okay, shoot, now we're kind of stuck in this date and we can never update it again. And does it quite make sense to put all that data into post-training? Um, and so mid-training is just a smaller pre-training run to help expand like the model's intelligence and like up-to-dateness.

AI assessment note: “a way to like extend the model's like intelligence without having to do a whole”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q When did you realize like I'm working at one of the most important companies of this generation? Like, like when, when was the moment where you were like, Hey, this is something that I obviously believe is important. That's why I joined, but that you raised like the scale and significance.

A Honestly, I kind of had this moment before I joined OpenAI, like, like, I think with the scaling laws paper with GPT-III, I was just like, kind of hit me that, like, if this exponential is true, like, there's not really much else I want to spend my life working on, um, and like, I want to be part of this, like, story. Like, I think there's, there's gonna be so many interesting things unlocked with this, and I think this is, this is probably the next, like, step level in terms of, like, technology, that it kind of made me realize, like, oh, I, I should probably go start reading about deep learning, and Figure out how I can get into one of these labs.

AI assessment note: “Honestly, I kind of had this moment before I joined OpenAI, like, with the scaling laws paper”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Christina, can you explain what mid training is and how it sort of, what does it achieve that pre or post doesn't?

A So I think with your pre-training runs, these are like your, these are your, the big runs. These are the massive ones. Like that's what we're building all these giant clusters for. Um, so you can kind of think of mid training is literally it's for like middle. Like we do it before, um, after pre-training, but before post-training, um, you kind of think of a way to like extend the model's like intelligence without having to do a whole new pre-training run. So this is mostly just focused on data and off of the pre-training models. Um, so this is a way for us to do things like updating the knowledge cut off of these models, right? So when you pre-train it, you're kind of like, okay, shoot, now we're kind of stuck in this date and we can never update it again. And does it quite make sense to put all that data into post-training? Um, and so mid-training is just a smaller pre-training run to help expand like the model's intelligence and like up-to-dateness.

AI assessment note: “mid-training is just a smaller pre-training run to help expand like the model's intelligence”

Answered raw tape D 5 · C 4 · P 3 · Cm 4 4.05

Q Say more about how you received such, uh, this kind of reduction in hallucinations, but also, also deception. What's the relationship between those?

A I guess, like, for me, I find hallucinations, deceptions, like, pretty related. So the model, um, and we kind of saw this a lot with the reasoning models. Like, they, the reasoning model would understand that it didn't have some ability, but then it still really wanted to respond. I think if we really baked it into the models that they want to be helpful, and so they're like, whatever I can say to be helpful in that moment. Um, and that's kind of what we consider for, like, deception. Versus hallucinations, sometimes the model, like, literally, uh, it seems that they will just say something quickly. Um, and we kind of see a lot of this reduction with the thinking, with, When the models are able to take step by step, they actually can like pause before blurting out an answer is kind of what I, it feels like with a lot of the previous models or hallucinations.

AI assessment note: “for me, I find hallucinations, deceptions, like, pretty related.”

Answered raw tape D 5 · C 4 · P 3 · Cm 3 3.90

Q Let's talk about, uh, creative writing. Maybe you talk about the, the improvements there. How do you think about it?

A That's one of my favorite improvements in GBT five. Um, the writing, I honestly find it's very tender and touching, especially for a lot of the creative writing that we want to do. Um, we were thinking through like a bunch of different samples for the live stream. And like, every time I was like, oh, that's like, actually like that, like hits like, and it's like spooky. And I'm just like, oh, this feels like, Someone, like, someone should have written this. Um, but I think it's really cool, because you can actually really use it for, um, like, helping you with things. Like, like, I, like, my example I did in the livestream was, like, writing, helping me write the eulogy. Something that, like, that's, like, kind of hard to write. Especially since writing isn't really something a lot of people are good at. Like, I'm personally a very, very bad writer.

AI assessment note: “the writing, I honestly find it's very tender and touching, especially for a lot”

Answered raw tape D 5 · C 4 · P 3 · Cm 3 3.90

Q Over the next few weeks as, as you're evaluating usage, what are the biggest questions that you're having or that you're sort of anticipating, uh, being potentially answered?

A I'm just really curious to see how all of these things, um, reflect in usage, right? Like I think coding is way, way better. Like what does this actually unlock for people? And I think we're really excited to be offering these models at the price points that we have. Cause I think this actually like unlocks like a lot more use cases that really weren't there before. Maybe like previous competitor models were, are good at coding, but the price point is not as exciting. And so I think with this number of capabilities that we have in this model and the price point, I'm kind of excited to see like all the new startups and like developers, like doing things on top of it.

AI assessment note: “I'm just really curious to see how all of these things, um, reflect in usage”

Answered raw tape D 5 · C 4 · P 3 · Cm 3 3.90

Q Let's talk about, uh, creative writing. Maybe you talk about the, the improvements there. How do you think about it?

A That's one of my favorite improvements in GBT five. Um, the writing, I honestly find it's very tender and touching, especially for a lot of the creative writing that we want to do. Um, we were thinking through like a bunch of different samples for the live stream. And like, every time I was like, oh, that's like, actually like that, like hits like, and it's like spooky. And I'm just like, oh, this feels like, Someone, like, someone should have written this. Um, but I think it's really cool, because you can actually really use it for, um, like, helping you with things. Like, like, I, like, my example I did in the livestream was, like, writing, helping me write the eulogy. Something that, like, that's, like, kind of hard to write. Especially since writing isn't really something a lot of people are good at. Like, I'm personally a very, very bad writer.

AI assessment note: “That's one of my favorite improvements in GBT five. Um, the writing”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q Say more about how you received such, uh, this kind of reduction in hallucinations, but also, also deception. What's the relationship between those?

A I guess, like, for me, I find hallucinations, deceptions, like, pretty related. So the model, um, and we kind of saw this a lot with the reasoning models. Like, they, the reasoning model would understand that it didn't have some ability, but then it still really wanted to respond. I think if we really baked it into the models that they want to be helpful, and so they're like, whatever I can say to be helpful in that moment. Um, and that's kind of what we consider for, like, deception. Versus hallucinations, sometimes the model, like, literally, uh, it seems that they will just say something quickly. Um, and we kind of see a lot of this reduction with the thinking, with, When the models are able to take step by step, they actually can like pause before blurting out an answer is kind of what I, it feels like with a lot of the previous models or hallucinations.

AI assessment note: “When the models are able to take step by step, they actually can like pause”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q Over the next few weeks as, as you're evaluating usage, what are the biggest questions that you're having or that you're sort of anticipating, uh, being potentially answered?

A I'm just really curious to see how all of these things, um, reflect in usage, right? Like I think coding is way, way better. Like what does this actually unlock for people? And I think we're really excited to be offering these models at the price points that we have. Cause I think this actually like unlocks like a lot more use cases that really weren't there before. Maybe like previous competitor models were, are good at coding, but the price point is not as exciting. And so I think with this number of capabilities that we have in this model and the price point, I'm kind of excited to see like all the new startups and like developers, like doing things on top of it.

AI assessment note: “curious to see how all of these things, um, reflect in usage”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q Yeah. Um, how about in the, in the broader sort of, uh, AGI discourse, like, what is this, um, what, what does this mean or accelerator or not, or like, how do we think about sort of the broader, um, AI discourse in terms of what does GPT-V mean here or change the conversation in any sort of way?

A I think with GPT-V, um, Because it's, like, a new, it's obviously state of the art in, like, all the things we talked about, um, but I think if you're showing that, like, you know, we can continue pushing the frontier here, and I feel like there's always people like, oh, we're hurting a wall, like, things aren't actually improving, um, and I think the interesting thing is, I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right? Like, who, what are the new use cases that are being unlocked, and, like, what, how, like, how many more people are using this in their daily lives to help them, like, across multiple So I feel like that's actually, like, the ultimate usage in terms, like, that I'm excited about for, in terms of, like, are we getting to AGI?

AI assessment note: “showing that, like, you know, we can continue pushing the frontier here”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q Yeah. Um, how about in the, in the broader sort of, uh, AGI discourse, like, what is this, um, what, what does this mean or accelerator or not, or like, how do we think about sort of the broader, um, AI discourse in terms of what does GPT-V mean here or change the conversation in any sort of way?

A I think with GPT-V, um, Because it's, like, a new, it's obviously state of the art in, like, all the things we talked about, um, but I think if you're showing that, like, you know, we can continue pushing the frontier here, and I feel like there's always people like, oh, we're hurting a wall, like, things aren't actually improving, um, and I think the interesting thing is, I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right? Like, who, what are the new use cases that are being unlocked, and, like, what, how, like, how many more people are using this in their daily lives to help them, like, across multiple So I feel like that's actually, like, the ultimate usage in terms, like, that I'm excited about for, in terms of, like, are we getting to AGI?

AI assessment note: “we can continue pushing the frontier here... in terms of, like, are we getting to AGI?”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.