Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q of your first goals is like build the framework grind time and driver for, for AMD. Uh, and then on, On June third on Twitch, uh, you weren't as excited about AMD anymore. Maybe let's talk a bit about that. And like, uh, you compared the quality of like commit messages from like the AMD kernel to like the Intel work that people are doing there. What's important to know?
A So when I said I wanted to, I want to write a framework, I didn't never intended on writing a kernel driver. I mean, like I flirted with that idea briefly, but like realistically, I like there's three parts to it, right? There's like the ML framework, there's the driver, and then there's the user space runtime. I was even down to rewrite the user space runtime. I have, I have a GitHub repo called CUDA IO control sniffer. It's terribly called, but you can actually launch a CUDA kernel without CUDA. So you don't need CUDA installed, just the Nvidia open source driver. And this open source repo can launch a CUDA kernel. So rewriting the user space runtime is doable. Rewriting the kernel driver? I don't even have docs. I don't have any docs for the GPU. Like, it would just be a massive reverse engineering project. Um, so that is, when I saw that there, like, it wasn't, like, I wasn't complaining about it being slow. I wasn't complaining about PyTorch not compiling. I was complaining about the thing crashing my entire computer. It panics my kernel, and I have to wait five minutes while it reboots because it's a server motherboard, and they take Like five minutes to reboot. Um, so I was like, look, if you guys do not care enough to get me a decent kernel driver, there's no way I'm wasting my time on this, especially when I can use Intel GPUs. Intel GPUs have a stable kernel driver an…
AI assessment note: “I was complaining about the thing crashing my entire computer. It panics my kernel”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q Awesome. Um, yeah. Any other thoughts? Tiny corp bounties?
A Um, yeah, so we have, you know, I've been thinking a lot about, like, what it means to hire in today's world. What actually is the, like, core? Okay, look, I'm a believer that machines are gonna replace everything in about 20 years. Uh, so, okay. What is that? What is that thing that people can still do that computers can't, right? Um, and this is a narrowing list, but like, you know, back in the day, like, imagine I was starting your company in 1960, right? Oh, and we're gonna have to hire a whole bunch of calculators in the basement to do all the, you know, math to support the calculus. Dude, have you heard about computers? Why don't we just buy a few of those? Oh. Oh, wow, man. You're right. Um, so like, I feel like that's kind of happening again. And I'm thinking about, I will post in my discord. I'll be like, okay, who wants to like, okay, I just changed my unary ops used to be log and exp in like E. Um, I changed them to be log two and exp two because hardware has log two and exp two accelerators. Yeah. And of course you can just change a base. It's one multiply to, to get it back to E, but like I made the primitives log two and exp two. Right. And this is the kind of I just posted in the Discord. I'm like, did someone put this pull request up? And someone eventually did and I merged it. But I'm like, this is almost to the level where models can do it. We're almost to the…
AI assessment note: “I just posted in the Discord. I'm like, did someone put this pull request up?”
Redirected raw tape
D 2 · C 4 · P 4 · Cm 3 3.25
Q Um, John Carmack mentioned there's about six insights we have left. Do you have an intuition for what some of the paths people should be taking? Obviously, you're working on one. Um, what are some of the other branches of the tree that people should go under?
A I don't think I'm working on one of the six insights. I don't think TinyGrid's any one of the six insights. Um, something I, I really like that Elon does, and I try to take it from, uh, try to be inspired by it is, um, Look at the boring tunnel machine and ask how you can build a 10 x cheaper one, right? Look at the rocket. How can I build a 10 x cheaper one, right? Look at the electric car and say, how can I build a 10 x cheaper, like, cheaper or, you know, can go further or whatever, whatever, whatever, right? And you just do the straight up physics math, right? Like, I'm trying to do the same thing with, with, uh, ML frameworks, right? And in, in, in doing so, making sure that this stuff remains accessible, right? You could imagine a world where it, Google TPUs were actually the ultimate. Google TPUs were actually the best training things. I mean, actually, you know, I'm kind of grateful for NVIDIA, right? Like, because if Google TPUs were the ultimate, now you have this huge closed source compiler in between XLA and the hardware, and yeah, that's just a really bad thing. So, I mean, something that is somewhat upsetting about the TinyCorp is that it is trying to prevent downside. But, uh, it's not all trying to prevent downside. Like, we're also building computers, and we're gonna build some awesome, powerful, cheap computers, uh, along the way. Uh, so no, I'm not really wor…
AI assessment note: “I don't think I'm working on one of the six insights.”
Partly raw tape
D 2 · C 3 · P 4 · Cm 3 2.95
Q Yeah. Um, so that's kind of like the lowest level of the stack. And then at a slightly higher level, obviously there's tiny grad, there's module, uh, there's ggml. How are you thinking about breadth versus like depth and like where you decided to focus early on?
A Um, so ggml is very much like a, okay, everyone has M ones, right? Actually I was thinking in the beginning, I was thinking of something more like ggml, focus on the M ones, but ggml showed up and was just like, we're actually just focusing on the M ones. Um, So, and actually M one PyTorch is considerably better than AMD PyTorch. M one PyTorch works. It only gives wrong answers sometimes, and it only crashes sometimes, but like some models kind of run. Um, when I was writing the, uh, metal backend, I was comparing to MPS PyTorch, and I had like a, I had a discrepancy. A tiny grid checks all its outputs compared to Torch, and I had one where it didn't match. I'm like, I really, I, I checked the matrix by hand. It matches TinyGrad. I don't understand. And then I switched PyTorch back to CPU and it matched. I'm like, oh. Yeah. Well, there's like bugs. Like if you like transpose the matrix, because like, I think it's like has to do with like multi views and PyTorch and like weird under the hood stuff that's not exposed to you. Like there, there's bugs and maybe they fix them, but like, you know, it seems like there was a lot of momentum again, because you're getting a huge variety. You're getting how many engineers care about making PyTorch work on M one, right? 1010 of thousands.
AI assessment note: “in the beginning, I was thinking of something more like ggml, focus on the M ones”
Redirected raw tape
D 1 · C 4 · P 4 · Cm 3 2.95
Q No, I think like the, you're planning to release the first tiny box ship them in like two to six, eight months, something like that. Uh, what's up on mind for you in terms of building a team? Who should, who are you calling for?
A Yeah. Uh, well, to, to, to stay on the tiny box for, for, for, for one minute. Um, so I have the GPUs picked out and you're like, well, I could make that computer with the GPUs. And my answer is, can you, do you know how to put, do you know how hard it is to put six GPUs in a computer? People think it's really easy, and it's really easy to put one GPU in a computer. It's really easy to put two GPUs in a computer, but now you want to put in eight. Ok, so I'll tell you a few things about these GPUs. They take up four slots. What kind of computer? You can buy the nicest super micro. You can't put eight of those in there. You need two slot blowers. If you want to use one of those, those four year super micros, you need two slot blowers or water cooling, right? If you're trying to get the four slot cards in there, you're going to need some form of water cooling. Uh, or you're going to need, there are some like Chinese 40 nineties that are blowers, right? You have any blowers or water cooling if you're trying to get it in those things, right?
AI assessment note: “to stay on the tiny box for, for, for, for one minute.”
Partly raw tape
D 3 · C 2 · P 3 · Cm 2 2.55
Q Oh, I mean, like, what, what makes, what makes, what's, what about Qualcomm architecture?
A Oh, what, what makes it doable? Well, because the world has spent how many millions of man hours to make the video fast? And Qualcomm has a team of 10 Qualcomm engineers? Okay, well, who can I beat here? Let's, like, like, what I propose, what I propose with TinyGrad is that developer efficiency is much higher. But even if I have 10 x higher developer efficiency, I still lose on a video, right? You know, ok, I didn't put a 100,000 man hours into it, right? If they put a million, like, like, that's what I'm saying, but that's what I'm saying, we can get. And we are going to close this speed gap a lot. Like, I don't support Tensor course yet. That's, that's, that's a big one that's just gonna, ok, massively close the gap. And then AMD. Uh, I can't even get, I don't even have a benchmark for AMD because I couldn't get it compiled. Oh, and I tried. Oh, I tried. I spent a day, like, I spent actually a day trying to get PyTorch, and I got it built. I got it kind of working. Then when I tried to run a model, like, there's all kinds of weird errors, and the rabbit holes are so deep on this. I'm like, um, so we, you know, you can compare the speed. Right now, you can run Llama. You can run anything you want on AMD. It already all works. Any OpenCL backend works, and it's not terribly slow. I mean, it's a lot faster than crashing. So it's an infinitely times faster than PyTorch on AMD. U…
AI assessment note: “Qualcomm has a team of 10 Qualcomm engineers? Okay, well, who can I beat here?”