The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Sohmers: Hardening silicon for specific AI models is obsolete in months
“Doing any of that, like, hardening for specific Model things. I don't think lasts more than, you know, two or three months at the rate that the industry moves at.”
Sohmers: AI hardware over-indexes on raw FLOPS instead of memory bandwidth
“Everyone else was focusing on the wrong things. They were just trying to have more and more flops when memory bandwidth, memory capacity were the real, real bottlenecks.”
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Sohmers: Requiring workload recompilation creates fatal friction for AI chip adoption
“If you are requiring a user or having yourself as the company needing to actually recompile a workload, that's already one step too far, even if you assume it works perfectly.”
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Sohmers: Google Veo and Imagen 3 are pure autoregressive transformers, not diffusion
“A lot of things have actually been moving away from diffusion to being pure autoregressive transformers for image and video generation. So like the latest, yeah, there's a VO three and since image and three on, on Google side have been pure autoregressive movi…”
Sohmers: Positron AI hardware achieves 93% of theoretical memory bandwidth
“And so our fundamental architecture is enabling us, you know, today with hardware that we're shipping right now to be achieving, you know, 93% of the theoretical memory bandwidth of our device consistently across all use cases.”
Sohmers: Transformer inference is memory-bound with a 1:1 FLOP-to-byte ratio
“And on the other side of this chart, you have the case of a transformer where when you're actually, you know, doing attention, Or really just any case where you're doing, you're fundamentally doing matrix vector multiplication rather than matrix matrix multipl…”
Sohmers: Cray-2 was the last major system with balanced memory-to-compute ratio
“And if you look at sort of traditional big iron compute systems, the last like major compute, compute platform that had that balance of memory to compute ratio was the Cray two supercomputer.”
Sohmers: Matrix-vector multiplication in transformer inference is fundamentally uncacheable
“So the second level of this is that matrix vector multiplication is basically uncacheable. When you're doing transformer inference, matrix A is the weights of your model. And so if you're talking about model weights that are tens of gigabytes, hundreds of giga…”
Sohmers: AMD's PyTorch fork lagged official releases by 6-9 months
“And AMD had their own separate you know, non-mainline PyTorch distribution for years. That was always six to nine months behind any new PyTorch releases.”
Sohmers: People will continue buying NVIDIA for AI training
“We are betting that people are going to continue to train on NVIDIA for at least the foreseeable future, where, since we're able to, you know, and I'll say, I really hope others are able to be successful in, in providing competition against NVIDIA, but Given t…”
Sohmers: NVIDIA's TF32 is actually a 19-bit precision format
“NVIDIA's TF-thirty-two number format is a nineteen-bit number format. They just call it thirty-two-bit.”
Sohmers: NVIDIA will have a very good decade ahead despite startup challengers
“The reality is NVIDIA is going to have a very, very good decade ahead for them. And The market is growing so fast that all of us in the space trying to take them on can be very happy with, you know, very, very small wins in the space.”
Sohmers: Reasoning models shift inference workloads to 100 output tokens per input
“If you go back a year from today in July of last year, the ratios of like input to output for LLMs were very, very heavily on, on inputs where you could be doing, you know, 10, 10, 15 to one ratio of input to output. But that has completely flipped and it's ob…”