The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Bryan Catanzaro argument clarity score 4.1/5 from 14 exchanges on raw tape · average scores: directness 3.9 · coherence 4.5 · precision 4 · compression 3.4 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
14exchanges match
14on raw tape
5redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Is what you just described, uh, called latent MOE, or is that a different concept?

A Latent MOE is a specific, uh, innovation that we have in NemoTron three family. And, uh, what it does is actually, um, reduces the amount of communication that has to be sent through NVLink during MOE computations by basically down projecting it. So, you know, every token produces a vector and the idea is like, we're going to take that vector and learn a way to compress it. And then send that compressed thing through the network, and then we're gonna uncompress it at the other end. And as a result, we save on network bandwidth, and we also get four times the number of experts for the same inference cost. So you could think about it as like, you know, our library of books got four times bigger, um, and we get to, you know, read four times more books, uh, at the same inference cost because of, because of this particular innovation.

AI assessment note: “Latent MOE is a specific, uh, innovation that we have in NemoTron three family.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q It's been dead for years? Like, why, why is that?

A Well, you just look at the, the progress, uh, in semiconductor manufacturing, you know, the, the original statement of Moore's law was economic, right? It was about, we can afford to put twice as many transistors on the same chip in every, whatever, 24 months, whatever the, the time period is. And, um, these days that is, Absolutely not the case. It hasn't been for probably five or 10 years, right? Now, we are still scaling our systems, right, um, through a number of ways. One is just applying a lot more silicon to it, right? Uh, we are also getting, transistors are continuing to get smaller and, and more efficient, although at a slower pace, but they're also getting quite a bit more expensive at the same time. Um, uh, so the, uh, you know, in an era where, where Moore's law was alive, the best way to make the system of the future was to take the system of the present and then just shrink it and, and maybe double it at the same time, right? But in an era where, where we've been living for a while now, where you don't get economic benefits from taking your existing design and shrinking it, uh, you really have to be more clever about how you use every part of the system. Uh, that, that's, uh, you know, an era where accelerated computing is, is much more valuable than ever because the, the work of thinking through the prop problem from first principles and co-designing absolutely …

AI assessment note: “you don't get economic benefits from taking your existing design and shrinking it”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q To double click on this, uh, at a high level, NemoTron family is focused on urgent reasoning with a particular focus on making it efficient. Is that, is that the right headline?

A That's right. Yeah. Um, NemoTron has always been, um, uh, speed first approach to building models because NVIDIA is an accelerated computing company. As I was saying, we're trying to think through what is the problem here computationally from first principles. And, um, you know, NemoTron, ah, three family has a lot of things in it that are, ah, we're really proud of. For example, ah, NemoTron Ultra and Super, ah, were pre-trained using four-bit arithmetic. We pre-trained those in MVFP four, um, which, you know, ah, is, ah, a not trivial thing to do, to invent the algorithm so that your model can converge to an excellent result using such coarse arithmetic, ah, required a lot of invention. Really proud of that.

AI assessment note: “That's right. Yeah. Um, NemoTron has always been, um, uh, speed first approach”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Great. Can you talk about the multi-token prediction, uh, which is also very interesting?

A If you're running at a low batch size, um, which is when you are trying to get the most interactivity if you're in a data center, so you want, you want the model to respond as quickly as possible, and it's okay for it to be more expensive. Your token, your cost per token might be higher, but you want the result as quickly as possible. Or if you're running, um, locally, um, so you might be running At batch size one just because you're the only person using it. It turns out that the GPU has extra execution capabilities that are just lying there unused. The bulk of the work when you're running in these scenarios is actually fetching the weights from memory, and then you push the token past those weights, and then you fetch more, more, uh, weights from memory. But it turns out if you, if you push two tokens or even five tokens, Through those same, um, weights, it would cost basically the same amount of time because the, the expensive thing is not doing the math to push the token through the weights. The expensive thing is just reading all of those weights from memory, all those parameters they have to come in. And so the idea with multi-token prediction is to take advantage of this by having the model predict multiple tokens at once. Let's say that the model predicts five tokens. We know the first token is correct. The next four tokens may or may not be correct. So then what we do …

AI assessment note: “having the model predict multiple tokens at once.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Great. What's the current state of the Neumontron family? You got Nano, you got Super, you got Ultra. What do those models do and what are the use cases for them?

A So Nano is a thirty billion, um, uh, total three billion active parameter model. Super is one 20 and 12, and Ultra is five 50 and 55. Um, they're designed, um, really to fit, you know, it's kind of small, medium, and large, um, uh, deployment scenarios. Um, uh, you know, nano can be really capable for things that, um, you know, don't require nearly as much knowledge or reasoning, but obviously for the, for the most, um, capable model, you go for ultra. Um, super in a lot of ways is our most popular model because it represents kind of a great balance between, Um, cost and, and intelligence. So we, we kind of like, um, having this small, medium, and large, um, approach to building a family just because our customers, um, seem to respond to that, um, pretty well. But, um, you know, uh, the most important thing from NVIDIA's point of view that people are doing with, uh, uh, with LLMs is agents, right, is, um, building agentic workflows Having it, having an agent working on your behalf, solving problems for you night and day, um, is such an exciting way of approaching the problems that we have to solve. Um, and, um, it's our dream to make NemoTron amazing for that purpose. That's, that's our goal.

AI assessment note: “Nano is a thirty billion, um, uh, total three billion active parameter model”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Another important characteristic of Numatron, free ultra is a one million token context, the, the long context window. How important is that in the overall mix and what does it enable the model to do?

A The longer the context length, the more challenging problems we can solve with a language model, um, that allows us to do things like append all sorts of information to a query, which could be a code base. It could be instructions. Um, you know, uh, in, in the longterm, I'm hoping that I have my own personal LLM that's able to read all of my emails, you know, and help me answer questions about that. You know, the more information that we can attach to a particular query, Um, the, the more useful the model can be. Um, now, uh, it can get more and more expensive, right, to reason over large amounts of, of input data. And, um, and so that's one of the, the reasons why there's usually a limit on how big the context length can be. But with Nemotron III, we, we tried to push it as far as we could go. Uh, we think a million tokens is a lot of tokens, um, and you can do a lot of things with that.

AI assessment note: “allows us to do things like append all sorts of information to a query”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q big question is, can they become great at, uh, Law and consulting, and then, you know, all sorts of different domains, and part of the black box of closed models is like how people go about doing all of this, where do they get the data from? To the extent that you can talk about all of this, I'd be very curious about how you guys have gone about it.

A It's not an easy question to answer because it is quite complex, but I would say, um, we rely on a number of things. One is that we do purchase data from, uh, companies that, um, uh, that, you know, are, are building data sets that you can purchase. Um, and to the extent that, you know, we have the rights to redistribute, uh, or to, to, to open up that data, we do as part of, um, our, our, uh, Mnemotron data effort. Um, you know, with, with Nemo Tron, we are trying to be maximally open with the data that we release because our goal is to support the ecosystem, right? Our goal is, is not to be the only model out there, and we love it when we hear of other models around the industry that are using our data sets, um, to, to make their AI stronger, because that means we're succeeding in our job to keep the ecosystem thriving and growing. Um, now, uh, We also are big believers in synthetic data generation. Um, we use an enormous amount of compute, um, uh, running language models on our own systems to create synthetic data that then helps our models be better at, uh, solving problems in specific domains, and we release a lot of that data as well. Now it's, of course, not very straightforward to do this. Like, you know, AI is always garbage in, garbage out. So you have to work really hard to make sure that any synthetic data that you create is actually adding value. That's actually he…

AI assessment note: “we rely on a number of things. One is that we do purchase data”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q All right. So you mentioned, uh, you know, making 500 people work together, and I said that we would get back to it because it's so interesting. So just taking a step back, like, tell us about the research organization at NVIDIA. Like, how is it structured? How does it all work?

A Well, NVIDIA is not structured according to an org chart. Um, we have one, but it's not actually the best way of understanding how we work. Um, Uh, my team, for example, is not part of the official NVIDIA research team. My team is actually part of the organization that builds the GPU. And my team is not the only team building NemoTron. There's probably 10 teams around the company that have significant involvement in building NemoTron, um, in different parts of the company, in, in enterprise software, um, in, uh, uh, our AI software, um, uh, division. The part of, uh, NVIDIA that actually designs the GPU also significantly is involved in, in building NemoTron. Um, so there's, there's so many different teams that, that have to work together. Um, we, um, always like to say that the mission is the boss, um, rather than, uh, the organization, but, um, what that implies is That people have to figure out how to work together, which is challenging in the sense that humans are naturally tribal creatures, and it's not natural for us to be friendly with people we don't know very well or trust co-workers that we don't have, you know, success working with in the past. And, you know, Actually, the name NemoTron reflects that. We had the Nemo team, which was building software for AI, and the Megatron team, which was building primarily focused on systems research for, um, for building large la…

AI assessment note: “NVIDIA is not structured according to an org chart.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q substantial open source frontier AI research effort that's been happening. So it's really interesting to hear that, you know, there's been this Progression. And now there's this family of models that we're going to talk about, uh, in a second. Another important moment seems to be, uh, the creation, uh, just in March, uh, three months ago of the Mnemotron coalition. Do you want to explain briefly what that is?

A So Mnemotron exists to help support the ecosystem. And we were thinking, well, this is a different kind of AI project than other Projects around the industry, right? Because, um, we're not actually trying to dominate in any way. We're just trying to support. We don't, we're not trying to control, uh, the way that AI is, um, uh, being integrated into all these companies. We're just trying to make sure there's good AI, but we thought, well, maybe if we worked with people while we develop it, then it's going to be more useful for them. It'll be easier to integrate. Because we will consider what they need from the beginning. And, um, you know, Nematron has always been collaborative. I was telling you that, you know, long, long time ago, our first big model that we trained, we, we did with Microsoft, right? It was a joint effort where NVIDIA and Microsoft researchers worked side by side to build that. And that, um, that ended up, I think, helping both NVIDIA and Microsoft. I think we both learned a lot from that, um, experience. So because Nemo Tron is not trying to compete with other companies, but rather support, because we're going to be putting it out there openly anyway, why not collaborate before the thing is built rather than Nemo Tron being a project that Nvidia does all on its own and then posts on the internet and says, hey, why don't you try this? We think it might be goo…

AI assessment note: “Mnemotron exists to help support the ecosystem.”

Redirected raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q Do you guys feel the AI backlash that seems to be forming internally? Is that, is that something that you all perceive think about? And if so, do you think it's a communication problem that our industry may have, you know, in particular, given what you just said about, uh, all the obvious Potential of AI?

A You know, I'm always worried about the way that the public thinks about technology and interacts with it. It matters a lot. Um, and it is definitely the case that, um, societies that want technological advancement have more technological advancement than societies that, that don't want change. Um, so I think it is, um, actually important to think about it. One thing that's interesting about AI is that, um, I believe it tends to be Uh, much more accepted when it is, uh, part of everyday life. And then at that point, people stop thinking about it as AI. It's just, oh, this is the tool that I use. Like, do you care whether it's AI that's helping you, uh, route your car when you ask the map application to help you drive somewhere? Like, I mean, it is, there is actually sophisticated AI that's going into that. Um, uh, uh, but you're not really thinking about that, right? You're just using a tool. And, um, so I feel like people's acceptance of AI, um, you know, comes with experience, right? Um, the more experience we have working with it, the more we learn how to work with it productively, um, I think the more comfortable we, we become with it.

AI assessment note: “people's acceptance of AI, um, you know, comes with experience, right?”

Not addressed raw tape D 2 · C 4 · P 4 · Cm 3 3.25

Q How do you balance useful research with great exploratory research?

A My belief is that research needs to be bootstrapped. Um, research is a chicken and egg problem. So, um, it is always the case that every researcher believes if I just had a lot more resources, my idea would change the world. Actually, it's important that researchers feel that way, because if you didn't feel that way, you wouldn't have the conviction that's required to go do something crazy and new. Right? So you have to believe, um, and, and so of course, um, uh, you start with that belief, uh, but then how do you translate that belief into something that other people can understand, right? That other people are willing to invest in. Um, this is what I call the, the chicken and egg problem, right? Because like, once your research idea is, is obviously good and impactful, it's easy to get resources, but how do you get it to be obviously good and impactful without those resources, right? So, so the way you solve chicken and egg problems is by bootstrapping. This is an iterative, uh, problem solving approach where you do something small. You get some sort of signal about this is a good idea, and you tell people about that, and then you ask for just a little bit more. And, um, if people saw like, oh yeah, that, you know, that experiment turned out pretty well. That's pretty intriguing. We should probably do a little bit more there. Um, then you're on track, right? And, uh, that, th…

AI assessment note: “My belief is that research needs to be bootstrapped.”

Redirected raw tape D 2 · C 4 · P 4 · Cm 2 3.10

Q Isn't a GPU a thing for gaming as well?

A Oh, right. Yeah. There's, there's also that, right? Which we, we continue to, um, run into that, uh, that idea. Actually a GPU is whatever NVIDIA says it is. You know, we make that. So a GPU is, is a, is a thing that we make in order to accelerate the world's most important computations, um, which in 1995 was graphics, and, you know, for a long time now, it's been AI. So, um, anyway, I started at NVIDIA. Um, I was in the research group, um, doing, uh, strange things about trying to make, um, uh, compilers, libraries, uh, for AI on the GPU. Um, that led to, um, the creation of Uh, first Copperhead, which was a, it was a, a Python, uh, embedded language that, um, compiled to the GPU, um, which I think foreshadowed a lot of things in TensorFlow and, and PyTorch. Um, and then, um, uh, then that led to the creation of QDNN, which was NVIDIA's first product for, um, uh, for deep learning on the GPU. Um, and, uh, I, I really enjoyed working on that, um, but I was always wanting to see more, uh, First hand about the applications of AI, and at NVIDIA, I was mostly working on, you know, libraries and compilers for AI, so I thought, um, well, you know, when Andrew Ng asked me to go build the Silicon Valley AI Lab with him at Baidu, I thought, oh, this is a great opportunity because even back then, Baidu, um, was very advanced in its application of AI to its core business. And so, um, uh, …

AI assessment note: “Oh, right. Yeah. There's, there's also that, right?”

Redirected raw tape D 2 · C 4 · P 2 · Cm 2 2.60

Q community that's wondering whether open source as an ecosystem, not, not Nvidia, but in general has been progressing in part based on the ability to distill closed source models and in a world where we seeing the anthropics and fable Fives of the world starting to discourage distillation. Do you think there is a chance that open source AI progress may slow down in that context or as a result?

A You know, in my mind, there's no question that when the, um, technology community decides to make huge investments in the most transformational technology of our time, that there's going to be rapid progress. Um, and also that that technology is not going to be controlled by a small group of people. Um, because that's just not the way that, um, the industry works. You know, we, we, um, do our best work. Um, uh, we have the most impact with our work when we're able to, uh, each think about it in our own way and apply it in our own way. So, um, you know, uh, I love, uh, the, uh, closed AI APIs, uh, whether from Anthropic or other people, I think they're amazing. You know, I'm really, really impressed with the work that those labs are doing, but they're not the only labs in the world. There's lots of labs around the world, and lots of people have a good idea. Um, it's not the case that there's only a few labs that have the monopoly on all good ideas. That's just not true. That's not how humanity operates. There's a, there's a lot of bright people on this planet. And, um, you know, the community, Uh, of course, cares deeply about this technology. It's obviously so transformational. It has such profound impacts on so many things, um, that, that, of course, uh, many people, uh, wanna be involved in that. And, um, so I think over time, we're gonna see that, um, community-oriented appr…

AI assessment note: “community-oriented approaches to developing and deploying AI are gonna continue to strengthen”

Redirected raw tape D 1 · C 4 · P 2 · Cm 2 2.30

Q weights model in the U S and that was just a few days ago. And then, uh, even more recently, GLM 5.2 came out and that was another moment. So it seems that things are accelerating in open source AI. It feels like a great Uh, place to start. What's your assessment, uh, about where we are and how wide the gap between closed source and open source currently is?

A Well, it's really exciting to see all of the energy going into open technologies for AI because we know that, um, open technologies make it possible for people to innovate. You know, the internet is such a great example of that. Um, we actually did have closed internets. I don't know if you remember things like America Online and Prodigy back in the day. Um, and they were great. Um, and open internet has also, uh, been amazing, right? Like so many different companies have been able to figure out how to transform their work, um, thanks to, uh, an open technology. The application of the internet to retail is very different from the application of the internet to healthcare or manufacturing, but all of them have been totally transformed, um, by the internet. Um, AI, uh, I believe is, uh, Um, also a very transformational technology and also a technology that needs to be applied in very diverse ways. And because of that, I believe that open technologies for AI are really fundamental. Um, and it's very exciting to see continued, um, investment and development of open technologies from, uh, for AI from so many different organizations around the world. Um, uh, and, uh, you know, I, I hope that that continues.

AI assessment note: “it's really exciting to see all of the energy going into open technologies”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.