The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Dylan Patel no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay. That's fascinating. Uh, we've touched on post-training, uh, for a bit, but just to recap, post-training is, so you have a model that's good at predicting the next word. And in post-training, you sort of give it a personality by inputting sample conversations to make The model want to emulate the certain values that you want it to take on?

A Yeah, so post training can be a number of different things. The most simple way of doing it is, is yeah, uh, pay for humans to label a bunch of data, um, take a bunch of example conversations, um, et cetera, and input that data and train on that at the end, right? Um, and so that, that example data is, is useful, uh, but this is not scalable, right? Like using humans to train models is just so expensive, right? So then there's the magic of sort of reinforcement learning and, And other synthetic data technologies, right? Where the model is helping teach the model, right? So you have many models in, in, in a sort of, in a post training where, yes, you have some example human data, but human data does not scale that fast, right? Cause, cause the internet is trillions and trillions of words out there. Whereas, you know, even if you had, you know, Alex and I write words all day long for our whole lives, it would, we would have Millions or, you know, hundreds of millions of words written, right? It's nothing. It's like orders of magnitude off in terms of the number of words required. Um, so then you have the model, you know, take some of this example data, um, and you have various models that are surrounding the main model that you're training, right? And these can be policy models, right? Teaching it. Hey, is this, is this what you want or that what you want? Uh, reward models, righ…

AI assessment note: “post training can be a number of different things. The most simple way”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Right, so how does the model then learn how to, you know, figure this out, and in what context is it accurate to say blue and red? Right.

A So, so I mean, um, first of all, the model doesn't just output one token, right? It outputs a distribution. Um, it turns out the most, the way most people take it is they take the top, uh, top K, i.e. the most high probability. So yes, blue is obviously the right answer if you give it to anyone on this planet. Um, but there are situations and contexts where the sky is red is the appropriate sentence, but that's not just an isolation, right? It's like if the prior passage is all about Mars and all this, and then all of a sudden it's like, And that's like a quote from a Martian settler, and it's like the sky is, and then the correct token is actually red, right? The correct word. Um, and so it has to know this through the attention mechanism, right? Um, if it was just the sky is blue, always you're going to output blue because blue is, let's say, 80%, 90%, 99% likely to be the right option. But as you, as you start to add context about Mars or any other planet, right? Other planets have different colored, colored atmospheres, I presume. Um, you end up with this, um, distribution starts to shift. Right? If, if I add, we're on Mars, the sky is, you know, then, then all of a sudden blue goes from 99%, you know, in the prior context window, right? The, uh, the text that you sent to the model, the attention of it, all of a sudden it realizes the sky is blue, uh, preceded by that, the,…

AI assessment note: “and so it has to know this through the attention mechanism, right?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. And so why is it called pre-training?

A So, so pre-training is, is, is sort of called that because it is what happens, you know, before the actual, uh, training of the model, right? Uh, the objective function in pre-training is to just predict the next token. Uh, but predicting the next token is not what humans want to use AIs for, right? I want it to ask a question and answer it. Uh, but in, in most cases, asking a question does not necessarily mean that the next most likely token is, is the answer, right? Oftentimes it is another question, right? Uh, for example, if I ingested the entire SAT, um, you know, and I asked a question, the next, like, five answers, then all the next tokens would be like, A is this, B is this, C is this, D is this, like, no, I just want the answer. Right? Um, and so pre-training is, the reason it's called pre-training is because you're ingesting humongous volumes of text no matter the use case. Right? Um, and you're learning the general patterns across all of language. Right? I don't actually know that king and queen relate to each other in this way. And I don't know that king and queen are opposites in these ways. Right? Um, and so this is why it's called pre-training is because you must get a broad general understanding of the entire sort of world of text. Before you're able to then do post training or fine tuning, which is let me train it on more specific data that is specifically usef…

AI assessment note: “called that because it is what happens, you know, before the actual, uh, training”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q then coming out with its, with a new prediction. I think Carpathie does a very interesting job, uh, in the YouTube video talking about how models think with tokens. The more tokens there are, the more computes they, compute they use, because they're running these predictions through the transformer model, which we discussed, and therefore they can come to better answers. Is that the right way to think about reasoning?

A So, so I think that, um, humans are also fantastic at fat pattern matching, right? Um, we're really good at like recognizing things, um, but a lot of tasks, it's not like an immediate response, right? We are thinking, um, whether that's thinking through words out loud, thinking through words in an inner monologue in our head, or it's just like processing somehow, and then we know the answer, right? Um, and this is the same for models, right? Models are horrendous at math, right? Historically have been, Right. Um, you could ask it, you know, what is nine dot one, one bigger than nine dot nine. Um, and it would say, yes, it's bigger, even though like everyone knows that nine dot one, one is, is way smaller than nine dot nine. Right. Um, and that's just like a thing that happened in models because they didn't think or reason. Right. And then it's the same for you, Alex, right? Like, you know, or myself, right? Like if someone asked me, you know, uh, 17 times 34, I'd be like, I don't know, like right off the top of my head, but You know, give me, give me a little bit of time. I can do some long form multiplication and I can get the answer. Right. And that's because I'm thinking about it. Um, and this is the same thing with reasoning for models, um, is, you know, when, when you look at a transformer, every word is this, every token output, it has the same amount of compute behind it…

AI assessment note: “and this is the same thing with reasoning for models”

Answered raw tape D 4 · C 4 · P 5 · Cm 4 4.25

Q but we also have massive, massive data center build outs. Um, I think you, it would be great to hear you kind of recap the size of these data center build outs. And then answer this question. If we are getting more efficient, why are these data centers getting so much bigger? And what might that added scale get in the world of generative AI for the companies building them?

A Yeah. So when we look across the ecosystem at data center buildouts, um, you know, we track all the buildouts and, and server purchases and supply chains here, and, and the pace of construction is incredible, right? You can just, you can pick a state and you can see new data centers going up, um, all across the U S and, and around the world. Right. Um, and so you see things like, um, capacity in, you know, for example, of, of the largest scale training supercomputers goes from, Hey, it's a few hundred million dollars. It's, it's, It's not even a few hundred million dollars years ago, but like, you know, hey, for GPT-IV, it was a few hundred million dollars, um, and it's, it's one building full of GPUs, too. Uh, GPT-IV, uh, 4.5, um, and, uh, the reasoning models, like, oh, one, oh, three were done in a, in three buildings on the same site, and, you know, billions of dollars to, hey, these next generation things that people are making are tens of billions of dollars, Um, like OpenAI's data center in Texas called Stargate, right? Uh, with Crusoe and Oracle and et cetera, right? Um, and likewise applies to Elon Musk, who's building these data centers in old fa- in an old factory where he's got, like, a bunch of, like, gas generation, uh, you know, outside, and he's doing all these crazy things to get the data center up as fast as possible, right? Um, and, and, and you can go to jus…

AI assessment note: “and because of the scaling laws, right? You know, 10 X more, uh, compute”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q to be 11. And so what it's doing in with tokens is taking words, breaking them down to their component parts, assigning them a numerical value, and then basically in its own word, in its own language, learning to predict what number comes next because computers are better at numbers and then converting that number back to text. And that's what we see come out. Is that, is that accurate?

A Yeah. And each individual token, uh, is actually, it's not just like one number, right? It's multiple vectors. You could think of like, Well, the tokenizer needs to learn king and queen are actually extremely similar on, on most in terms of like the English language, extremely similar, right? Except there is like one vector in which they're super different, right? Because a king is a male and a female, a queen is a female, right? And then from there, like, you know, in language, oftentimes kings are considered conquerors and, And you know, all these are the things and like, these are just like historical things, right? So a lot of the texts around them, while they're both like royal regal, right? Like, you know, uh, monarchy, et cetera, there are many vectors in which they differ. So like, it's not just like converting a word into one number, right? It's like converting it into multiple vectors and each of these vectors, the model learns what it means, right? Um, you don't initialize the model with like, Hey, you know, King means male, uh, monarch. And it's associated with, like, war and conquering, because that's what all the writing about kings is on, you know, in history and all that, right? Like, people don't talk about the daily lives of kings that much, uh, or they mostly talk about, like, their wars and conquests and stuff, um, and so, like, there will be each of these n…

AI assessment note: “Yeah. And each individual token, uh, is actually, it's not just like one number”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.