The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Hotz: GPT-4 is an 8-way mixture model with 220B parameters per head
“GPT-IV is two hundred twenty billion in each head, and then it's an eight-way mixture model.”
Hotz: Meta attracts researchers who want to publish while OpenAI keeps ideologues
“OpenAI can keep ideologues who, you know, believe ideological stuff, and Facebook can keep every researcher who's like, dude, I just want to build AI and publish it.”
Hotz: AI application wrappers like Cursor and Windsurf will fail
“And then on top you have, like, completely worthless things like cursor and windsurf you know, these character AI, all these people who think, oh, we're the app, we're gonna get the ARR. No, that worked in the web, it won't work for AI”
Hotz: AI scaling laws deliver linear returns for exponential capital
“AI scaling laws. You can put in exponentially more money to get linear returns.”
Hotz: Waymo robotaxis require approximately 1.2 remote human operators per car
“There's probably about, I would say there's 1.2 operators per car. But it's not, they don't have a steering wheel and pedals. It is an autonomous system that they're probably doing some higher level inputs on.”
Hotz: Analog computing for AI won't work and clockless chips aren't practical
“Analog computing just won't work. And clockless computing sure, it might work in theory, but your ETA tools are, maybe AIs will be able to design clockless chips, but not humans.”
Hotz: Systolic arrays are the wrong architectural choice for AI chips
“I think systolic arrays are the wrong choice. Systolic array, I think they have systolic arrays because that was the guy's PhD. And of course Amazon makes... They are very power efficient, but it becomes hard to schedule a lot of stuff. On them, if you're not …”
Hotz: Tinybox will be 5x faster than H100 systems per dollar
“For 90% of most companies model training use cases, the tiny box will be five X faster for the same price.”
Hotz: Machines will replace all human labor in about 20 years
“I'm a believer that machines are gonna replace everything in about 20 years.”
Hotz: Maximally compressed human brain represents only a couple gigabytes
“Quantization is a poor man's compression. I think we're only talking really here about, like, maybe a couple gigabytes, right? And then if you have, like, a couple gigabytes of true information of yourself up there, cool man. Like, what does it mean for me to …”
Hotz: Real AI alignment problem is corporate and government misalignment
“I think it's actually not a question of whether the computer is aligned with the company who owns the computer. It's a question of whether that company's aligned with you or that government's aligned with you. And the answer is no. And that's how you end up de…”
Hotz: AI is a far better military technology than nuclear weapons
“So, like, as a military technology, nukes are not that good. AI is way better.”
Hotz: Robotaxis will be a scooter-style capital race to the bottom
“Self-driving cars look like scooters. The only thing that it's going to take to roll out big fleets of self-driving cars is capital, right? It's just strictly a capital market... So self-driving cars are going to be this awesome race to the bottom, right? It's…”
Hotz: Robotic arm hardware is good enough, robotics is strictly software
“The arms are already good enough. It's the off the shelf arms are fine. It's all software. Again, it's always all software. Autonomous vehicles are all software. Robotics is all software”
Hotz: Shift to diffusion AI models will erode centralized cloud moats
“With AI, I think there's going to be a much less of a moat. Especially when you look at the move from autoregression to diffusion. So autoregression can run in large batch sizes. When you run ChatGPT, you're running with a whole bunch of other people on that s…”
Hotz: Tinygrad runs all ML models with only 25 primitive operations
“Tiny grad is, we are going to make a risk offset for all ML models. And yeah, it can run all ML models with basically 25 instead of the two 50 of XLA or PrimTorch. So about 10 X less complex.”
Hotz: Google TPUs are the only successful non-Nvidia training chips
“The only company, there's one other company aside from Nvidia who's succeeded at all at making training chips... Mid journey is trained on TPU, right? Like a lot of startups do actually train on TPUs, and they're the only other successful training chip aside f…”
Hotz: AMD CEO Lisa Su sent pre-release ROCm 5.6 to fix panics
“Lisa Sue reached out, connected with a whole bunch of different people. They sent me a pre-release version of Rock M 5.6. They told me you can't release it, which I'm like, okay, Why do you care? But they say they're going to release it by the end of the month…”
Hotz: Best chatbots will be smaller models with 1,000 training runs
“I don't think that the best chatbot models are going to be the big ones. I think the best chatbot models are going to be the ones where you had a thousand training runs instead of one. And I don't think that the interconnect bandwidth is going to matter that m…”
Hotz: Third company will build an AI girlfriend product
“The third company's the first one that's gonna build a real product, and that product is A girlfriend? No, like I'm dead serious, right? Like this is the dream product, right? This is the absolute dream product.”
Hotz: Merging with machines requires an AI companion, not Neuralink electrodes
“So I don't need to put, you know, electrodes in my brain to merge with a machine. I need an AI girlfriend, right?”
Hotz: Comma.ai rejects stealth mode because competitors cannot out-execute him
“Nah, we don't believe in stealth. I'm a really open guy. You are pretty open. I tell you everything I'm doing. Come on. Here's what I say. Here's what I say. I'm going to tell you what I'm doing. And you can try to compete, but I'll still crush you.”
Hotz: Either Nvidia is overvalued or AMD is significantly undervalued
“Either Nvidia's really overvalued or AMD's really undervalued. It has to be one or the other.”
Hotz: Automated flight booking will arrive via APIs, not agentic web browsing
“We're gonna pretty soon have computer use models that are actually capable of going to Delta.com and booking a flight. But then what's actually gonna happen is Delta's gonna partner with whatever company does that, and they're gonna put it behind the stupid th…”
Hotz: tinycorp will eventually partner or manufacture its own chips
“So I'd like to start another organization that eventually in the limit either works with people to make chips or makes chips itself and makes them available to anybody.”
Hotz: Tinygrad runs OpenPilot in production 2x faster than Qualcomm's library
“TinyGrad is used to run the model in OpenPilot. Like right now, it's been live in production now for six months. And TinyGrad is about two X faster on the GPU than Qualcomm's library.”
Hotz: Tinygrad could replicate PyTorch's API in two engineer-months
“Replicating the PyTorch API. Is something I can do with a couple, you know, like an engineer month or two.”
Hotz: Nobody has successfully trained models in INT8
“No one's gotten training to work with Indate yet. There's a few papers that vaguely show it, but if you're training, you're going to need BF-sixteen or float-sixteen.”
Hotz: None of 20 tested PCIe extenders work at PCIe 4.0
“No PCI extender I've tested and I've bought 20 of them works at PCIe four point out. So you're going to need PCIe redrivers now.”
Hotz: Except for Apple, secretive tech companies hide underwhelming tech
“Whenever a company is secretive, with the exception of Apple, Apple's the only exception, whenever a company is secretive, it's because they're hiding something that's not that cool.”
Hotz: Training on internet cross-entropy loss yields mediocre AI responses
“The problem is, if your loss function is categorical across entropy on the internet, your responses will always be mid.”
Hotz: RLHF models adopt customer support personalities
“I don't like the RLHF models. I don't like the tuned versions of them. I think that they become, you take on the personality of a customer support agent, right?”
Hotz: Current machine learning models are 1,000x less data efficient than humans
“Like current machine learning algorithms, like a thousand X less data efficient than humans. So yeah, you need a thousand X more data, right? If a human can learn something in one example, or 10 examples, the computer is going to need a thousand or 10,000.”
Hotz: Google's TPU compiler is a closed-source 32MB binary blob
“Not only is the chip closed source, But all of XLA is open source, but the XLA to TPU compiler is a 32 megabyte binary blob called lib TPU on Google's cloud instances. It's all closed source.”
Hotz: Tinygrad beats CoreML on ONNX tests and will soon pass ONNX Runtime
“We're below Onyx runtime, but we're beyond CoreML. So, like, that's, like, where we are in Onyx support now, but we will pass Onyx runtime soon, because it becomes very easy to add ops, because of how, like, you don't need to do anything at the lower levels.”
Hotz: Tinygrad is about 5x slower than PyTorch on Nvidia GPUs
“The correctness for both forwards and backwards passes is there, but on Nvidia, it's about five X slower than PyTorch right now.”
Hotz: Intel GPUs feature stable kernel drivers and public register docs
“Intel GPUs have a stable kernel driver and they have all their hardware documented. You can go and you can find all the register docs on Intel GPUs.”
Hotz: Consumer AMD GPUs lack peer-to-peer support
“If you have a consumer AMD GPU, they don't support peer to peer.”
Hotz: Halving GPU power yields 80% of peak performance
“Now, you can limit power on GPUs and still get, you can use like half the power and get 80% of the performance. This is a known fact about GPUs”
Hotz: Qualcomm SNPE cannot run transformers due to missing outer product ops
“Qualcomm's S and PE can't run transformers for this reason. So most matrix multiplies in neural networks are weights times values, right? Whereas you know, when you get to the outer product in in transformers, well, it's waste times weight. It's a, it's values…”