The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Jack Morris no published score: only 4 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 4 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
4exchanges match
4on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q is cool. And we can talk about that as well. But the other thing I think is that around about seven, eight B, maybe four B is when you start switching from like a single GPU setup to like a distributed setup. And I'm wondering, like, do grad students get HPC training? How much do they teach you of like, just how to work with like large clusters of stuff?

A Oh, to be clear, they don't teach you anything, like anything, like if you see a paper coming out from even, you know, Stanford, they're probably the best school in AI if you had to choose. And it's not like they're learning how to do like multi-node distributed FSTP training, like with whatever deep speed, you have to learn that from the internet and from other people. And like, there's no classes that really do that. I mean, it's, that's hard to facilitate, like, As one person, I would say most grad students are doing stuff on single GPU. Some people are doing multi-GPU training. There's probably basically no grad students doing multi-node training. I mean, there's probably a few, especially if they have like company affiliations, but that's really unusual, I think.

AI assessment note: “Oh, to be clear, they don't teach you anything, like anything”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q this all ends up, that's great. But like people aren't, are not doing that. Instead we're building, You know, five hundred billion dollar data centers in the middle of Texas and like, you know, all hail the, the, the God cluster, uh, that just will, you know, eventually wrap around the sun and consume solar energy because that's, that's what we need. Do we finish out the universal geometry thing?

A Let me finish the kind of, uh, methodological description. So, so we had this goal. So, so yeah, back to the embedding universality, we started with going from embeddings to text. We know about this platonic representation hypothesis, and maybe I'll skip over the details, but basically we had total inspiration from computer vision and this model from 20 17 called CycleGAN, which is among other things, uh, it's a way to map between two different distributions without any underlying notion of like which thing should be mapped where. It's just based on some kind of Idea of closeness. So like the cool thing about this, if you look at the top left, so I guess the, the top left is Monet. So impressionist paintings and this picture on the right is a photograph. So like it's learning this kind of like semantic notion of what content goes, where just by mapping a distribution of Monet pictures to a distribution of photographs without actually telling it which Monet picture should map to which photograph. It's kind of a subtle point I'm making. It takes a little bit of time to wrap your head around, or maybe like go to the middle one, if you don't mind the zebras and the horses. So like, it's clearly learning like what an animal is and what legs are and sort of like more abstract stuff, like what, uh, the camera position should be and, and what grass is and stuff like that. And it's lear…

AI assessment note: “Let me finish the kind of, uh, methodological description. So, so we had this goal.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. Do you think this is a hard limit? Do you think someone can come up with a better algorithm, but better architecture, and then sort of just change the slope?

A There are two axes here. One is the ability of the model to store data, and I think we can definitely improve that. I think, like, maybe even if we tested this with LALAMA architecture, like, there's sort of like a GPT++ architecture, like, I would guess that can store better data just because the kind of numerical flow is a little bit better, the nonlinearities are maybe, like, A little bit more suitable to training. Like that will probably raise the bound a little bit. And then the second axis is that our measurement tools are just not that good. Like this is, you know, me, I'm a grad student. I'm running all these hyperparameter sweeps and sort of like where we draw conclusions from them. But even that being said, like there are probably ways to measure this better. And, but all that would do is push the number up. So it's possible. Like there is a way to store five bits per parameter. If you have like a better optimization technique or If you were a super genius and you could just perfectly set the weights to store the data, then maybe you can do better. And this is just sort of like what we can reach through optimization is this 3.6 bits per parameter. But I would be happy if someone came along with a much better measurement tool. Like, uh, this is just sort of like the first measurement. I mean, I would, I would guess in the future, like, you know, people will look back a…

AI assessment note: “There are two axes here. One is the ability of the model to store data”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q off of, um, just a lot of pre-trained data that is, that is potentially collided. Okay. There was one, there's two more papers that we wanted to cover and then we can, we can sort of wrap it. You had an approximating You had a language model training data. I think this is a little bit, uh, also newer. How does this rank in terms of your, your overall work?

A Let's return to the kind of information theory question. So yeah, maybe we'll, we'll skip over the contextual embeddings in the case of time, but we'll group those papers. Great paper. Hopefully people start training with that technique. It's kind of a free lunch. Those questions are all about information and model activations. Like how much Can we recover from this given vector? Or like, what data does this vector represent? Or what computation does this vector represent? And there's really two types of, like, if you want to taxonomize, there's, there's two types of, whatever you call it, dense information storage mechanisms. One of them is activations or embeddings, which we were discussing already. And then the other is weights, which are the things that Are used to perform the computation, but not the computation itself. And so we have now two papers in this direction of what is stored in the weights. The first one is about language model capacity, which is called how much can language models memorize or how much do language models memorize? I never remember which one we settled on. And then the other one is called approximating language model training data from weights. The first one is like, I think has a lot of deep Messages about how language models store information and how they work in general. The second thing is like a proof of concept of maybe like a longer term re…

AI assessment note: “The second thing is like a proof of concept of maybe like a longer term”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.