Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q Before, before we move to inference, I, I do just want us to stay on chips and just say, okay, so we have TPUs, we have NVIDIA, we have, um, AMD. Is this in terms of distribution of gains? Is this a winner take all market? Is this cloud where you have several providers who are dominant? What does the distribution of gains look like in the chip market?
A So I would divide it in two categories. Well, three categories. Um, The GPUs you can buy or rent, the TPUs you can rent, and the TPUs you can buy. This is how the market is structured today, right? Right now, if you are, you want to go dedicated, there's, at least in the cloud, there's two options, TPUs and Tranium. TPUs on Google, Tranium on Amazon. So these are, you know, available chips, you can rent them today. If you want to buy, um, GPUs or RAN GPUs, you know, they're GPUs. We, we, we know it all the time. And there's this new wave of computing, which are dedicated, you know, chips you can actually buy. The Tense Torrent, the Etched, the Vsora. So I think it will be a mix of, you know, whoever, you know, whatever you get. For instance, let's say you are in Google Cloud. Of course, you don't want to do Nvidia. You get ripped off. Cause here's the, here's the dirty secret is that Nvidia, like a TSMC sells you at 60% margin. Nvidia sells you at, you know, 90% margin. And on top of that, there's Amazon that takes, let's say a 30% margin. So you are a very thin crust on a very big cake. And so that's why to me, it's, it's, it's not really, it's a big, it's a bit of a losing game. If you, you know, go all in on one provider, you want, you know, optionality.
AI assessment note: “I think it will be a mix of, you know, whoever, you know, whatever you get.”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q Which is why you don't, or are you saying that to get into one of these buy processes, you have to buy so much that it prohibits you?
A It's actually, it's actually both. So the buy-in is very high. So to make it worth it, you have to buy a lot. And if you buy a lot, this is, you know, what every, we talk to all of them. They, they always have the same questions and it's completely understandable. They say, this is great, but who's the customer? Because on the other side, let's take Amazon, for instance, with Tranium. Apple just came and said, Hey, we're going to buy a 100,000 of them. Oh, so you want to buy, you know, 10,000, you feel like the big shot, right? Yeah, but yeah, take, you know, go back to the queue because there's Apple before you, right? So they have to have very high commitments to make it, you know, you cannot be incrementally better. It's very hard and also very hard. I can give you a, I can give you one metric if you want. Um, I know for a fact that being seven times better and whatever, take whatever metric you want. Uh, whether it's spend, whether it's whatever. It's not enough to get people to switch. People will choose nothing over something. Right. I've like, I have stories. So this is a very hard market to enter into because you cannot also compete of incremental gains. Right. It's very hard. Right. So you have to convince a lot of people. Um, maybe you can go the, um, Middle East route in which, you know, they sprinkle everything and they, you know, evaluate everything, but you know, …
AI assessment note: “It's actually, it's actually both. So the buy-in is very high.”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q Sorry. Let me just. Is distillation. Wrong. And if we're all progressively moving towards a better future for humanity, more efficient models, is distillation not effectively open source in, in another wrapper?
A I think it's fair game to be honest. Honestly, I, I, I do not, I will not shed a tear. Uh, it's fair game. If you, there were like some people who tried to ask, I think it was, I don't know if it was an open, I don't, I don't remember if it was an open AI model. But they asked it to, um, so a diffusion model image, right? They asked it to generate, uh, an image from, uh, a Star Wars movie at whatever timestamp. And it came out with the Star Wars movie, you know, screenshot, right? So obviously it was trained with it. I think it's fair game because the, I mean, there's no free lunch, right? Uh, it was trained with data. You had a good ride. Somebody was, was sneaky and, and took it, but you took it from the beginning too. So let's just accept that, you know, it's, it's fair game.
AI assessment note: “I think it's fair game to be honest.”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q Just so I understand, why does it work for coding and not for other things?
A Because you don't use the mod, the AI model to generate, you know, output. You use the machine, you use the, you, you, you just run the code, right? And you see what it makes and you run all these code. And you create data out of it. Whereas if you run an LLM and you say to an LLM, all right, generate me two trillion tokens of text. It will do it with its, you know, so you may inject and stuff. So there's a lot of tricks, but ultimately my guts tell me that it's, it feels wrong, right? Because you re-inject data that was there. And so it will deteriorate. There's, you know, there's loss. So yeah, I'm, I'm a bit bullish. I'm not sure exactly on what vertical. Code is one. Um, but, um, we'll see. Distillation is in some sense a bit like that, right? You distill, you know, you create synthetic data from a bigger model into a smaller one. Probably the most, I would say mind blowing thing about distillation is that sometimes the smaller models become better than the, than the, than the bigger model through distillation. So, um, so, so we'll see, uh, but I love this.
AI assessment note: “You use the machine... you just run the code, right? And you see what it makes”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q Why are GPUs not built for AI? And if not, what is better?
A So, The, the way it worked is that a screen is, you can think of a screen as a matrix, right? And if you have to render, you know, pixels on a screen, there's a lot of pixels and everything has to happen in parallel, right? So that you don't waste time. Uh, turns out, you know, matrices are very, are a very important thing in AI. So there was this cool trick in which we essentially tricked the GPU into back that was like Probably 20 years ago, we would trick the GPU into believing it was doing graphics rendering where actually we would making it do parallel work, right? It was called GPGPU at the time, right? So it was always a cool trick, right? And very cool and very successful at that, mind you. Um, but it was not dedicated for this. Uh, the pioneers probably were, uh, of course, Google with GPU. Uh, which are very much more advanced, uh, on the architectural level, but essentially the way they work, um, it kind of works for, for AI, but for LLMs that starts to, you know, to crack because they're so big and there's a lot of memory transfers and so on. Actually, that's why Grok achieves, ah, not Grok, but Grok, Cerebras, and all these folks, they achieve very high performance single stream is because the data is right in the chip that doesn't have to get it from memory, which is slow, which GPU has to do. So there's a lot of these things that ultimately make it a good trick. …
AI assessment note: “we would trick the GPU into believing it was doing graphics rendering”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q How do you do that then? Do you have agreements with all the different providers?
A Oh yeah, yeah, yeah. Not agreements, but like we, we work with them, uh, to support their, their chips. But the thing is my, at least as, you know, a, I would say a user myself of You know, of our tech is that if I can, you know, if it's free for me to switch or to choose whichever provider I want in terms of compute, right? AMD, Nvidia, whatever. Uh, then I can take whatever is best today and I can take whatever is best tomorrow and I can run both. I can run, you know, three different platforms at the same time. I don't care. I only run, you know, what is good at the moment. And that unlocks to me a very cool thing, which is incremental, you know, improvement. If you are 30% better, I'll switch to you.
AI assessment note: “Not agreements, but like we, we work with them, uh, to support their, their chips.”
Answered raw tape
D 5 · C 3 · P 3 · Cm 3 3.60
Q Can I just ask, what's latent space reasoning?
A So the way models reason today is the reason it was in, in, in tokens. So it's as if, if you think to yourself, you would, you know, say out loud what you're thinking. So yes, it works, but it is a bit, you know, inefficient, right? And you lose, uh, information doing this. Latent space reasoning is this without going, I would say to English or whatever, right? So staying, you know, in what's called the latent space, which is where all the information of an LLM, let's take an LLM, an LLM lives. So this is very much how we, uh, how we, you know, work as humans. Uh, and we move toward what, what Yann Lequin calls, uh, energy based model in which We have different types of, um, uh, longer or shorter, I would say, thinking times, if you will. So that fundamentally, GPUs cannot deliver, deliver this, plain and simple at scale. So these chips can deliver the chips.
AI assessment note: “Latent space reasoning is this without going, I would say to English”
Answered raw tape
D 5 · C 3 · P 3 · Cm 3 3.60
Q Sorry. Let me just. Is distillation. Wrong. And if we're all progressively moving towards a better future for humanity, more efficient models, is distillation not effectively open source in, in another wrapper?
A I think it's fair game to be honest. Honestly, I, I, I do not, I will not shed a tear. Uh, it's fair game. If you, there were like some people who tried to ask, I think it was, I don't know if it was an open, I don't, I don't remember if it was an open AI model. But they asked it to, um, so a diffusion model image, right? They asked it to generate, uh, an image from, uh, a Star Wars movie at whatever timestamp. And it came out with the Star Wars movie, you know, screenshot, right? So obviously it was trained with it. I think it's fair game because the, I mean, there's no free lunch, right? Uh, it was trained with data. You had a good ride. Somebody was, was sneaky and, and took it, but you took it from the beginning too. So let's just accept that, you know, it's, it's fair game.
AI assessment note: “I think it's fair game to be honest. Honestly, I, I, I do not”
Answered raw tape
D 5 · C 3 · P 3 · Cm 3 3.60
Q Just so I understand, why does it work for coding and not for other things?
A Because you don't use the mod, the AI model to generate, you know, output. You use the machine, you use the, you, you, you just run the code, right? And you see what it makes and you run all these code. And you create data out of it. Whereas if you run an LLM and you say to an LLM, all right, generate me two trillion tokens of text. It will do it with its, you know, so you may inject and stuff. So there's a lot of tricks, but ultimately my guts tell me that it's, it feels wrong, right? Because you re-inject data that was there. And so it will deteriorate. There's, you know, there's loss. So yeah, I'm, I'm a bit bullish. I'm not sure exactly on what vertical. Code is one. Um, but, um, we'll see. Distillation is in some sense a bit like that, right? You distill, you know, you create synthetic data from a bigger model into a smaller one. Probably the most, I would say mind blowing thing about distillation is that sometimes the smaller models become better than the, than the, than the bigger model through distillation. So, um, so, so we'll see, uh, but I love this.
AI assessment note: “You just run the code, right? And you see what it makes”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q So can you just help me understand? Sorry, why is training more is more and that's great? And in inference less is more? Why do we have that?
A Yeah, think of it think of it like, you know, doing a Painting and doing a million paintings, right? The tools you will use, the process you will do. If you do one painting, what you favor is the speed at which you can do a stroke and do some iteration. If you do a million, what you want is a process, a process that is reliable, that can deliver you efficiently a million paintings, right? Um, so that is the same for, for, for training versus inference. Um, if you run around, you know, Millions of instances of a model. You cannot, you know, hack your way to do that. By the way, people do hide their way, uh, today. Um, but this is probably the fundamental difference.
AI assessment note: “doing a Painting and doing a million paintings... that is the same for, for, for training versus inference”
Partly raw tape
D 3 · C 4 · P 3 · Cm 3 3.30
Q How do you do that then? Do you have agreements with all the different providers?
A Oh yeah, yeah, yeah. Not agreements, but like we, we work with them, uh, to support their, their chips. But the thing is my, at least as, you know, a, I would say a user myself of You know, of our tech is that if I can, you know, if it's free for me to switch or to choose whichever provider I want in terms of compute, right? AMD, Nvidia, whatever. Uh, then I can take whatever is best today and I can take whatever is best tomorrow and I can run both. I can run, you know, three different platforms at the same time. I don't care. I only run, you know, what is good at the moment. And that unlocks to me a very cool thing, which is incremental, you know, improvement. If you are 30% better, I'll switch to you.
AI assessment note: “Not agreements, but like we, we work with them”
Answered raw tape
D 4 · C 3 · P 3 · Cm 3 3.30
Q What does everyone think they know about inference that they actually don't, or what does everyone get wrong about inference?
A Probably Not a lot of people are accustomed to what it entails to, to run production. So that inference is production and production is hard. Somebody has to wake up at night. And I used to be that guy, right? I don't want to do it again. Um, so production is hard. Thankfully, we have a lot of, uh, uh, software nowadays to do that a lot better, but there's not a lot of reuse because the AI field at least Is not really accustomed to that yet. It's changing. Uh, but you know, the discussions I had, you know, a year ago, and the discussions I had, you know, today are not the same. They're going to the right direction, but they're not there exactly yet. So probably that would be the number one thing. That is only, you know, um, uh, training code running only forward pass, right? This is not what it is.
AI assessment note: “inference is production and production is hard”
Answered raw tape
D 3 · C 3 · P 4 · Cm 3 3.25
Q Before, before we move to inference, I, I do just want us to stay on chips and just say, okay, so we have TPUs, we have NVIDIA, we have, um, AMD. Is this in terms of distribution of gains? Is this a winner take all market? Is this cloud where you have several providers who are dominant? What does the distribution of gains look like in the chip market?
A So I would divide it in two categories. Well, three categories. Um, The GPUs you can buy or rent, the TPUs you can rent, and the TPUs you can buy. This is how the market is structured today, right? Right now, if you are, you want to go dedicated, there's, at least in the cloud, there's two options, TPUs and Tranium. TPUs on Google, Tranium on Amazon. So these are, you know, available chips, you can rent them today. If you want to buy, um, GPUs or RAN GPUs, you know, they're GPUs. We, we, we know it all the time. And there's this new wave of computing, which are dedicated, you know, chips you can actually buy. The Tense Torrent, the Etched, the Vsora. So I think it will be a mix of, you know, whoever, you know, whatever you get. For instance, let's say you are in Google Cloud. Of course, you don't want to do Nvidia. You get ripped off. Cause here's the, here's the dirty secret is that Nvidia, like a TSMC sells you at 60% margin. Nvidia sells you at, you know, 90% margin. And on top of that, there's Amazon that takes, let's say a 30% margin. So you are a very thin crust on a very big cake. And so that's why to me, it's, it's, it's not really, it's a big, it's a bit of a losing game. If you, you know, go all in on one provider, you want, you know, optionality.
AI assessment note: “Nvidia sells you at, you know, 90% margin. And on top of that, there's Amazon”
Answered raw tape
D 4 · C 3 · P 3 · Cm 2 3.15
Q So what, what chips are great and why aren't they on the market?
A I mean, if you look at, you know, let's say for instance, Cerebrus, right? Incredible technology. Incredibly expensive. So will, you know, how will the market value the premium of having single stream very high tokens per second? Uh, there is a value into that, right? As we saw with Mistral and perplexity, but I'm not sure, you know, that was done. I think that was done at the loss. I don't know. I don't have the details, but I think it was done at the loss, uh, that Cerebris, you know, put, put it out. Um, so today there's three actors on the market that can, you know, deliver this. I think this will be, I would, I would say the, uh, the, The pushing force for change in the inference, uh, landscape, uh, agents and reasoning. So that is, you know, very high tokens per second only for you and not to a, you know, an aggregate of people.
AI assessment note: “Cerebrus, right? Incredible technology. Incredibly expensive.”
Answered raw tape
D 3 · C 3 · P 2 · Cm 3 2.75
Q So can you just help me understand? Sorry, why is training more is more and that's great? And in inference less is more? Why do we have that?
A Yeah, think of it think of it like, you know, doing a Painting and doing a million paintings, right? The tools you will use, the process you will do. If you do one painting, what you favor is the speed at which you can do a stroke and do some iteration. If you do a million, what you want is a process, a process that is reliable, that can deliver you efficiently a million paintings, right? Um, so that is the same for, for, for training versus inference. Um, if you run around, you know, Millions of instances of a model. You cannot, you know, hack your way to do that. By the way, people do hide their way, uh, today. Um, but this is probably the fundamental difference.
AI assessment note: “that is the same for, for, for training versus inference.”