The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Madra: The three fastest AI inference companies are not NVIDIA
“The three fastest companies in inference right now are not Nvidia.”
Madra: NVIDIA's CUDA moat does not exist for inference workloads
“There is no
Tie into CUDA that's required to go faster.
That's required to get the models running, right?
Obviously none of the three companies run CUDA.
And so that moat doesn't exist around inference.”
Madra: Alibaba's 30-billion parameter Qwen model matches GPT-4o performance
“There was a release today of a Quen You know, thirty billion parameter model, which is performing as good as GPT-IV-O.”
Madra: Chinese open-source AI offers 90% of top quality at 90% discount
“One of the things that you see with the leading open source models now, which are the Chinese is 90% of the quality in terms of intelligence, but at a 90% price discount.”
Madra: US open-source AI will hit the global top three by 2026
“I think if we look at Q four this year or Q one next year, I'd be willing to say, you know, that a top three worldwide model will be a U S based open source model.”
Madra: Enterprise usage will shift back to US models if OpenAI open-sources
“With OpenAI's release and, you know, Meta charges back, or even if some of these startups emerge that, you know, we can point at, I think we'll see a huge shift back towards those models versus, versus the Chinese ones.”
Madra predicts multi-agent systems will enable transactional AI within one year
“You can have a thousand agents working together.
You can have one that's making sure that the credit card charge is not too big.
You can have another one to make sure that the address is right.
You can have another one checking against your calendar.
And so al…”
Madra: NVIDIA's internal AI productivity gains for chip design exceed 100%
“I really, you know, been thinking a lot about Jensen's point in the pod about, you know, how much AI they're using internally for design, design verification for all those pieces. Right. And I think, you know, it's not 30%. I actually think sort of that's an u…”
Madra: Fewer developers will touch CUDA long-term, weakening NVIDIA's software moat
“I think there's going to be fewer people touching that. And I do think that's a point where they're the moat is not as strong as a longer term, as you say, and think about like, you know, the way the analogy that I would go with is like, think about the number…”
Madra: Meta's open-source stance is its primary recruiting weapon for AI talent
“I think part of the pitch to get everyone there is that it's open. Because everyone that they're pulling over are coming from closed places, and so if you're really passionate about the work you're doing, and you're passionate about where this is going to go, …”
Madra: Google's monthly AI token volume exploded 200x to 1 quadrillion
“In the Mary Meeker bond deck, you know, they have a slide there that shows Google went from five trillion tokens a month to 480 trillion tokens a month, and they had just put some press out that they crossed, like, you know, 800 trillion, and I saw something t…”
Madra: Inference clusters will be smaller and more distributed than training
“You'll see inference clusters be large, but not as large as a training clusters and be a lot more distributed because you don't need it to be all in the same place.”
Madra: Real-time AI inference cannot run across geographically distributed data centers
“You can train a model across a distributed site and it may just take you a, you know, a month longer because you have to move traffic around. And so instead of taking three months, it takes you four months, but you can't really run a model across a distributed…”