Everything Matt Garman said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Garman: Expecting one dominant AI model is a fundamentally flawed premise
“And this is where a lot of people started, and I just think it's fundamentally the wrong way of thinking about it. Which is a lot of times people are thinking about there's just going to be this one model and I want to have the one model that's going to be the…”
Garman: AWS is open to hosting OpenAI and Google Gemini models
“We're always open to having OpenAI be available in AWS someday, or having Gemini models be available in AWS someday, and maybe someday we will spend more time focused on our own models, for sure.”
Garman: Trainium 2 Delivers 30% to 40% Better Price-Performance Than GPUs
“We see 30 to 40% price performance games versus our instances that are GPU powered today.”
Garman: AWS automated reasoning mathematically eliminates AI hallucinations in bounded domains
“So by that, by this kind of mechanism, you're, like, systematically able to actually mathematically prove that you got the right answer coming out of this and completely eliminate hallucinations for that area, right? It doesn't mean that we've eliminated hallu…”
Garman: Small modular nuclear reactors will be major energy sources starting 2030
“We do think that you know, over the next probably starting somewhere in, in 20, 30 and beyond these small modular reactors, which is what X-energy builds, are gonna be a huge component of this.”
Garman calls Anthropic Claude models the best performing in the world
“I think people love using Anthropix Claude models. Those are fantastic, and right now, those are the best performing models in the world which is fantastic.”
Nvidia runs its own AI training clusters on AWS infrastructure
“NVIDIA actually runs their AI training clusters in AWS, because we actually have the most stable infrastructure of anyone else, and so they actually get the best performance from us”
Garman: AWS Project Rainier will deploy hundreds of thousands of Trainium2 chips in 2025
“We're building together with them what we call Project Rainier and so it's using our next generation Tranium II chips, and this cluster that we're building for them in twenty-twenty-five will be five times the size, the number of exaflops that they use to trai…”
Garman: Past AI Accelerators Failed Due to Weak Software Versus NVIDIA
“I think that's one of the things where people who've tried to build Accelerated platforms before have fallen down is the software support has not been as good as Nvidia software support is fantastic.”
Garman says Google and Microsoft forced new app paradigms in early cloud
“I think if you look at some of those others like Google and later Microsoft they kind of went at this space for first, we were the first ones out there that had anything like this, but even soon after that, I think they went after the space, like that was goin…”
Garman estimates 85% of enterprise workloads still run on-premise
“Now that we're at a hundred billion run rate you look at, you still go out there, and I think, 85% of workloads are still running on-prem today, by most estimation, somewhere in that range, you know, pick your number, whether it's 80 to 90, whatever it is, lik…”
Secret intelligence contract win against IBM marked AWS inflection point
“One of the big Inflection points we saw is we went after the intelligence agencies for the U S government, and we won that contract and it was secret. And it, you know, we pushed really hard to go in that. It was against all the incumbents, HPs and IBMs and Or…”
Garman predicts foundation models will command less attention over time
“Today the models is the front and center thing that everybody pays attention to, but I think increasingly it'll become a smaller percentage of the thing that people pay attention to”
Multi-year fab lead times will keep AI hardware supply chains constrained
“Look, I think we're probably going to be in a constrained world for the next little bit of time. Just, you know, that some of these things are, they take time. Like, look how long it takes to build a semiconductor fab. Like, it is, it's not a short lead time, …”
Inference workloads must dominate training for AI capital investments to pay off
“Inference is, is one of those workloads that today it's, you know, fifty-fifty maybe of training in Inference, but in order for the math to work out, Inference workloads have to dominate, otherwise all this investment in, in these big models isn't really gonna…”
Garman: NVIDIA treats AWS very fairly in allocating GPU supplies
“They're very fair in dealing with us and we give long-term forecasts and they tell us what they can supply.”
Garman: Generative AI workloads are 50/50 today but shifting toward inference
“You know, we're still probably seeing about that ratio of fifty-fifty. I think more and more it's more inference than training, and increasingly we'll see more and more of the workload shift that way.”
Garman: Nvidia Blackwell delivers about 2.5x compute performance of H100
“I think we were expecting about two and a half times the compute performance out of the Blackwell chip that you get out of an H-one hundred”
Garman: Developers only spend about one hour a day actively coding
“It turns out, also, developers, on average, code about one hour a day. The rest of their day is spent doing documentation. It's spent writing unit tests. It's spent writing, doing code reviews. It's spent doing, You know, going to meetings, it's spent doing up…”
Garman: Cloud Optimization Is Largely Finished, Funding AI and AWS Growth
“Customers, I think number one, a lot of them have been optimized, right? And there's only so much you can kind of squeeze into an Optimize place and customers are still looking for optimizations, but a lot of that work has been done and they're using some of t…”
Garman: Andy Jassy's bureaucracy critique applied inside AWS too
“Yeah, I think it's across across Amazon. So it wasn't specific to the rest of Amazon. It was definitely inside of AWS too.”
Garman: Enterprise cloud adoption will flip from 20% today to over 80%
“But I do think, and I actually think that at a minimum I think that, that percentage could flip, and it could be eighty-twenty versus 28 where it is today, or even less.”
Garman predicts liquid cooling will make on-prem AI clusters too difficult
“Increasingly, I think that's going to get harder and harder as you move to liquid cooling and larger clusters”
Amazon Titan is by far the most popular embeddings model on Bedrock
“In fact, the Titan embeddings model, Is by far the most popular embeddings model that we have inside of Bedrock today for people that are building search indices and thinking about things like that.”