Aug 18, 2025 · 47m · latent-space

⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI

Thomas Sohmers · 28m spoken Mitesh Agrawal · 12m spoken Shawn Wang · 1m spoken Alessio Fanelli · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Positron AI founders Thomas Sohmers and Mitesh Agrawal discuss how their memory-bandwidth-optimized hardware architecture accelerates transformer inference, featuring zero-step Hugging Face compatibility and a clear silicon roadmap from commercial FPGAs to custom ASICs backed by a $51M Series A.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 7.5% of the talking time here. How this is scored →

The hosts as informed peer 5.7 Guest teaching 4.3 Guest disagreement 1.6 The hosts pushing back 1.2
05100:0015:0030:0045:000:03–4:33 · The hosts as informed peer 4/10 Introductions and Founding Story of Positron AI Swyx kicks off the interview with relevant context regarding the founders' backgrounds at Lambda Labs and Groq. The guests share their founding story and career history in a very collaborative and conversational manner.4:34–7:38 · The hosts as informed peer 6/10 Hardware Bottlenecks and the Failure of Software-Only Optimization Alessio asks an astute technical question on whether software-only optimization is hitting a local maximum. Thomas and Mitesh explain their hardware thesis, citing historical DSP insights and Rich Sutton's Bitter Lesson.7:39–12:03 · The hosts as informed peer 5/10 Roofline Analysis: Memory-Bound Transformers Versus Compute-Bound CNNs Thomas walks through roofline analysis slides, educating the hosts on arithmetic intensity and how transformer decoding is strictly memory-bandwidth bound compared to compute-bound CNNs or training workloads.12:03–17:34 · The hosts as informed peer 5/10 Matrix-Vector Computation and Saturated Memory Bandwidth Utilization Thomas explains the mathematical difference between matrix-matrix and matrix-vector operations, pointing out that transformer weights are inherently uncacheable and detailing why modern GPUs only hit low fractions of theoretical memory bandwidth.17:34–22:51 · The hosts as informed peer 6/10 Zero-Step CUDA Compatibility and Compiler-Free Weight Ingestion Swyx asks about the operational secret behind rapid shipping. Thomas and Mitesh explain their zero-step workflow that directly ingests raw PyTorch/SafeTensors binaries to bypass compilers entirely, critiquing past startup compiler traps and AMD ROCm friction.22:52–29:02 · The hosts as informed peer 7/10 FPGA Efficiency, Precision Formats, and the Dedicated ASIC Roadmap Swyx pushes on the apparent contradiction of using power-hungry FPGAs while claiming top perf-per-watt metrics. Thomas and Mitesh break down why customized balance beats raw GPU flops and call out Nvidia's TF32 naming convention.29:02–36:20 · The hosts as informed peer 6/10 General Linear Algebra Acceleration Versus Hardened Model ASICs Alessio and Swyx ask about model-specific ASICs like Etched and George Hotz's criticism. Thomas and Mitesh strongly reject the premise of hardening specific transformer architectures into silicon, arguing that rapid algorithmic shifts like DeepSeek's MLA make hardwired chips obsolete in months.36:20–40:56 · The hosts as informed peer 5/10 Capital Efficiency, Return on Invested Capital, and Series A Announcement Thomas and Mitesh announce their $51M Series A and explain why capital efficiency and ROIC matter more than massive fundraising rounds, emphasizing their goal of selling systems rather than taking cloud infrastructure onto their balance sheet.40:56–45:33 · The hosts as informed peer 7/10 The Dominance of Autoregressive Decoding in Reasoning and Multimodal AI Swyx synthesizes the architecture thesis around autoregressive decoding speed. Thomas and Mitesh validate and extend this by showing how test-time reasoning and multimodal autoregressive generation flip token input/output ratios dramatically.0:03–4:33 · Guest teaching 1/10 Introductions and Founding Story of Positron AI Swyx kicks off the interview with relevant context regarding the founders' backgrounds at Lambda Labs and Groq. The guests share their founding story and career history in a very collaborative and conversational manner.4:34–7:38 · Guest teaching 3/10 Hardware Bottlenecks and the Failure of Software-Only Optimization Alessio asks an astute technical question on whether software-only optimization is hitting a local maximum. Thomas and Mitesh explain their hardware thesis, citing historical DSP insights and Rich Sutton's Bitter Lesson.7:39–12:03 · Guest teaching 7/10 Roofline Analysis: Memory-Bound Transformers Versus Compute-Bound CNNs Thomas walks through roofline analysis slides, educating the hosts on arithmetic intensity and how transformer decoding is strictly memory-bandwidth bound compared to compute-bound CNNs or training workloads.12:03–17:34 · Guest teaching 6/10 Matrix-Vector Computation and Saturated Memory Bandwidth Utilization Thomas explains the mathematical difference between matrix-matrix and matrix-vector operations, pointing out that transformer weights are inherently uncacheable and detailing why modern GPUs only hit low fractions of theoretical memory bandwidth.17:34–22:51 · Guest teaching 5/10 Zero-Step CUDA Compatibility and Compiler-Free Weight Ingestion Swyx asks about the operational secret behind rapid shipping. Thomas and Mitesh explain their zero-step workflow that directly ingests raw PyTorch/SafeTensors binaries to bypass compilers entirely, critiquing past startup compiler traps and AMD ROCm friction.22:52–29:02 · Guest teaching 4/10 FPGA Efficiency, Precision Formats, and the Dedicated ASIC Roadmap Swyx pushes on the apparent contradiction of using power-hungry FPGAs while claiming top perf-per-watt metrics. Thomas and Mitesh break down why customized balance beats raw GPU flops and call out Nvidia's TF32 naming convention.29:02–36:20 · Guest teaching 4/10 General Linear Algebra Acceleration Versus Hardened Model ASICs Alessio and Swyx ask about model-specific ASICs like Etched and George Hotz's criticism. Thomas and Mitesh strongly reject the premise of hardening specific transformer architectures into silicon, arguing that rapid algorithmic shifts like DeepSeek's MLA make hardwired chips obsolete in months.36:20–40:56 · Guest teaching 3/10 Capital Efficiency, Return on Invested Capital, and Series A Announcement Thomas and Mitesh announce their $51M Series A and explain why capital efficiency and ROIC matter more than massive fundraising rounds, emphasizing their goal of selling systems rather than taking cloud infrastructure onto their balance sheet.40:56–45:33 · Guest teaching 6/10 The Dominance of Autoregressive Decoding in Reasoning and Multimodal AI Swyx synthesizes the architecture thesis around autoregressive decoding speed. Thomas and Mitesh validate and extend this by showing how test-time reasoning and multimodal autoregressive generation flip token input/output ratios dramatically.0:03–4:33 · Guest disagreement 0/10 Introductions and Founding Story of Positron AI Swyx kicks off the interview with relevant context regarding the founders' backgrounds at Lambda Labs and Groq. The guests share their founding story and career history in a very collaborative and conversational manner.4:34–7:38 · Guest disagreement 1/10 Hardware Bottlenecks and the Failure of Software-Only Optimization Alessio asks an astute technical question on whether software-only optimization is hitting a local maximum. Thomas and Mitesh explain their hardware thesis, citing historical DSP insights and Rich Sutton's Bitter Lesson.7:39–12:03 · Guest disagreement 1/10 Roofline Analysis: Memory-Bound Transformers Versus Compute-Bound CNNs Thomas walks through roofline analysis slides, educating the hosts on arithmetic intensity and how transformer decoding is strictly memory-bandwidth bound compared to compute-bound CNNs or training workloads.12:03–17:34 · Guest disagreement 2/10 Matrix-Vector Computation and Saturated Memory Bandwidth Utilization Thomas explains the mathematical difference between matrix-matrix and matrix-vector operations, pointing out that transformer weights are inherently uncacheable and detailing why modern GPUs only hit low fractions of theoretical memory bandwidth.17:34–22:51 · Guest disagreement 3/10 Zero-Step CUDA Compatibility and Compiler-Free Weight Ingestion Swyx asks about the operational secret behind rapid shipping. Thomas and Mitesh explain their zero-step workflow that directly ingests raw PyTorch/SafeTensors binaries to bypass compilers entirely, critiquing past startup compiler traps and AMD ROCm friction.22:52–29:02 · Guest disagreement 2/10 FPGA Efficiency, Precision Formats, and the Dedicated ASIC Roadmap Swyx pushes on the apparent contradiction of using power-hungry FPGAs while claiming top perf-per-watt metrics. Thomas and Mitesh break down why customized balance beats raw GPU flops and call out Nvidia's TF32 naming convention.29:02–36:20 · Guest disagreement 3/10 General Linear Algebra Acceleration Versus Hardened Model ASICs Alessio and Swyx ask about model-specific ASICs like Etched and George Hotz's criticism. Thomas and Mitesh strongly reject the premise of hardening specific transformer architectures into silicon, arguing that rapid algorithmic shifts like DeepSeek's MLA make hardwired chips obsolete in months.36:20–40:56 · Guest disagreement 1/10 Capital Efficiency, Return on Invested Capital, and Series A Announcement Thomas and Mitesh announce their $51M Series A and explain why capital efficiency and ROIC matter more than massive fundraising rounds, emphasizing their goal of selling systems rather than taking cloud infrastructure onto their balance sheet.40:56–45:33 · Guest disagreement 1/10 The Dominance of Autoregressive Decoding in Reasoning and Multimodal AI Swyx synthesizes the architecture thesis around autoregressive decoding speed. Thomas and Mitesh validate and extend this by showing how test-time reasoning and multimodal autoregressive generation flip token input/output ratios dramatically.0:03–4:33 · The hosts pushing back 0/10 Introductions and Founding Story of Positron AI Swyx kicks off the interview with relevant context regarding the founders' backgrounds at Lambda Labs and Groq. The guests share their founding story and career history in a very collaborative and conversational manner.4:34–7:38 · The hosts pushing back 2/10 Hardware Bottlenecks and the Failure of Software-Only Optimization Alessio asks an astute technical question on whether software-only optimization is hitting a local maximum. Thomas and Mitesh explain their hardware thesis, citing historical DSP insights and Rich Sutton's Bitter Lesson.7:39–12:03 · The hosts pushing back 1/10 Roofline Analysis: Memory-Bound Transformers Versus Compute-Bound CNNs Thomas walks through roofline analysis slides, educating the hosts on arithmetic intensity and how transformer decoding is strictly memory-bandwidth bound compared to compute-bound CNNs or training workloads.12:03–17:34 · The hosts pushing back 1/10 Matrix-Vector Computation and Saturated Memory Bandwidth Utilization Thomas explains the mathematical difference between matrix-matrix and matrix-vector operations, pointing out that transformer weights are inherently uncacheable and detailing why modern GPUs only hit low fractions of theoretical memory bandwidth.17:34–22:51 · The hosts pushing back 1/10 Zero-Step CUDA Compatibility and Compiler-Free Weight Ingestion Swyx asks about the operational secret behind rapid shipping. Thomas and Mitesh explain their zero-step workflow that directly ingests raw PyTorch/SafeTensors binaries to bypass compilers entirely, critiquing past startup compiler traps and AMD ROCm friction.22:52–29:02 · The hosts pushing back 3/10 FPGA Efficiency, Precision Formats, and the Dedicated ASIC Roadmap Swyx pushes on the apparent contradiction of using power-hungry FPGAs while claiming top perf-per-watt metrics. Thomas and Mitesh break down why customized balance beats raw GPU flops and call out Nvidia's TF32 naming convention.29:02–36:20 · The hosts pushing back 2/10 General Linear Algebra Acceleration Versus Hardened Model ASICs Alessio and Swyx ask about model-specific ASICs like Etched and George Hotz's criticism. Thomas and Mitesh strongly reject the premise of hardening specific transformer architectures into silicon, arguing that rapid algorithmic shifts like DeepSeek's MLA make hardwired chips obsolete in months.36:20–40:56 · The hosts pushing back 0/10 Capital Efficiency, Return on Invested Capital, and Series A Announcement Thomas and Mitesh announce their $51M Series A and explain why capital efficiency and ROIC matter more than massive fundraising rounds, emphasizing their goal of selling systems rather than taking cloud infrastructure onto their balance sheet.40:56–45:33 · The hosts pushing back 1/10 The Dominance of Autoregressive Decoding in Reasoning and Multimodal AI Swyx synthesizes the architecture thesis around autoregressive decoding speed. Thomas and Mitesh validate and extend this by showing how test-time reasoning and multimodal autoregressive generation flip token input/output ratios dramatically.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 19.6% · guest 80.4%0:00 · the hosts 19.6% · guest 80.4%3:00 · the hosts 17.1% · guest 82.9%3:00 · the hosts 17.1% · guest 82.9%6:00 · the hosts 14.5% · guest 85.5%6:00 · the hosts 14.5% · guest 85.5%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 7% · guest 93%12:00 · the hosts 7% · guest 93%15:00 · the hosts 3.8% · guest 96.2%15:00 · the hosts 3.8% · guest 96.2%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 16.6% · guest 83.4%21:00 · the hosts 16.6% · guest 83.4%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 11.6% · guest 88.4%27:00 · the hosts 11.6% · guest 88.4%30:00 · the hosts 10.6% · guest 89.4%30:00 · the hosts 10.6% · guest 89.4%33:00 · the hosts 2% · guest 98%33:00 · the hosts 2% · guest 98%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 10.6% · guest 89.4%39:00 · the hosts 10.6% · guest 89.4%42:00 · the hosts 0.3% · guest 99.7%42:00 · the hosts 0.3% · guest 99.7%45:00 · the hosts 5.7% · guest 94.3%45:00 · the hosts 5.7% · guest 94.3%
Sharpest disagreement ▶ 29:23 Dismissing silicon hardening for transient model architectures

Thomas forcefully rejects the philosophy of etching specific transformer mechanisms into ASICs, declaring that model hardening becomes obsolete within two or three months.

Hardest push from the hosts ▶ 22:51 Swyx challenges FPGA power efficiency narrative

Swyx directly challenges the guests on how they claim 3x power efficiency while running on FPGAs, which are conventionally known for being power-hungry.

Biggest teaching moment ▶ 8:50 Thomas delivers masterclass on roofline intensity of transformers

Thomas uses a visual roofline model to educate the hosts on why transformer inference collapses into a 1:1 flop-to-byte memory-bound regime compared to compute-dense training and CNNs.

The host holds their own ▶ 40:55 Swyx summarizes the architectural thesis to pre-fill versus decode

Swyx cleanly encapsulates Positron's core technical differentiation into a sharp pre-fill versus autoregressive decode performance framing.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions and Founding Story of Positron AI 4100 Swyx kicks off the interview with relevant context regarding the founders' backgrounds at Lambda Labs and Groq. The guests share their founding story and career history in a very collaborative and conversational manner.
Hardware Bottlenecks and the Failure of Software-Only Optimization 6312 Alessio asks an astute technical question on whether software-only optimization is hitting a local maximum. Thomas and Mitesh explain their hardware thesis, citing historical DSP insights and Rich Sutton's Bitter Lesson.
Roofline Analysis: Memory-Bound Transformers Versus Compute-Bound CNNs 5711 Thomas walks through roofline analysis slides, educating the hosts on arithmetic intensity and how transformer decoding is strictly memory-bandwidth bound compared to compute-bound CNNs or training workloads.
Matrix-Vector Computation and Saturated Memory Bandwidth Utilization 5621 Thomas explains the mathematical difference between matrix-matrix and matrix-vector operations, pointing out that transformer weights are inherently uncacheable and detailing why modern GPUs only hit low fractions of theoretical memory bandwidth.
Zero-Step CUDA Compatibility and Compiler-Free Weight Ingestion 6531 Swyx asks about the operational secret behind rapid shipping. Thomas and Mitesh explain their zero-step workflow that directly ingests raw PyTorch/SafeTensors binaries to bypass compilers entirely, critiquing past startup compiler traps and AMD ROCm friction.
FPGA Efficiency, Precision Formats, and the Dedicated ASIC Roadmap 7423 Swyx pushes on the apparent contradiction of using power-hungry FPGAs while claiming top perf-per-watt metrics. Thomas and Mitesh break down why customized balance beats raw GPU flops and call out Nvidia's TF32 naming convention.
General Linear Algebra Acceleration Versus Hardened Model ASICs 6432 Alessio and Swyx ask about model-specific ASICs like Etched and George Hotz's criticism. Thomas and Mitesh strongly reject the premise of hardening specific transformer architectures into silicon, arguing that rapid algorithmic shifts like DeepSeek's MLA make hardwired chips obsolete in months.
Capital Efficiency, Return on Invested Capital, and Series A Announcement 5310 Thomas and Mitesh announce their $51M Series A and explain why capital efficiency and ROIC matter more than massive fundraising rounds, emphasizing their goal of selling systems rather than taking cloud infrastructure onto their balance sheet.
The Dominance of Autoregressive Decoding in Reasoning and Multimodal AI 7611 Swyx synthesizes the architecture thesis around autoregressive decoding speed. Thomas and Mitesh validate and extend this by showing how test-time reasoning and multimodal autoregressive generation flip token input/output ratios dramatically.

Statements from this episode (24)

Assertion Supported
Agrawal: Lambda Labs generates well over $500M in ARR
“Really grew the company from zero dollars in revenue all the way to where it is now well over half a billion in ARR, right?”
Mitesh Agrawal Aug 18, 2025 ▶ 2:45
Insight
Sohmers: AI hardware over-indexes on raw FLOPS instead of memory bandwidth
“Everyone else was focusing on the wrong things. They were just trying to have more and more flops when memory bandwidth, memory capacity were the real, real bottlenecks.”
Thomas Sohmers Aug 18, 2025 ▶ 6:26
Insight
Sohmers: Transformer inference is memory-bound with a 1:1 FLOP-to-byte ratio
“And on the other side of this chart, you have the case of a transformer where when you're actually, you know, doing attention, Or really just any case where you're doing, you're fundamentally doing matrix vector multiplication rather than matrix matrix multipl…”
Thomas Sohmers Aug 18, 2025 ▶ 9:27
Assertion Contradicted
Sohmers: Cray-2 was the last major system with balanced memory-to-compute ratio
“And if you look at sort of traditional big iron compute systems, the last like major compute, compute platform that had that balance of memory to compute ratio was the Cray two supercomputer.”
Thomas Sohmers Aug 18, 2025 ▶ 10:39
Insight
Sohmers: Matrix-vector multiplication in transformer inference is fundamentally uncacheable
“So the second level of this is that matrix vector multiplication is basically uncacheable. When you're doing transformer inference, matrix A is the weights of your model. And so if you're talking about model weights that are tens of gigabytes, hundreds of giga…”
Thomas Sohmers Aug 18, 2025 ▶ 13:33
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Thomas Sohmers Aug 18, 2025 ▶ 14:53
Assertion Open · timeframe Aug 2026
Sohmers: Positron AI hardware achieves 93% of theoretical memory bandwidth
“And so our fundamental architecture is enabling us, you know, today with hardware that we're shipping right now to be achieving, you know, 93% of the theoretical memory bandwidth of our device consistently across all use cases.”
Thomas Sohmers Aug 18, 2025 ▶ 15:26
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Thomas Sohmers Aug 18, 2025 ▶ 16:28
Disclosure
Agrawal: Positron AI shipped hardware within 15 months using only seed funding
“It's like, we got to market shipping within 15 months of founding of the company with only the seed round raised, and again, are now planning our next generation of Silicon within 18 months of the first generation.”
Mitesh Agrawal Aug 18, 2025 ▶ 16:56
Assertion Supported
Sohmers: AMD's PyTorch fork lagged official releases by 6-9 months
“And AMD had their own separate you know, non-mainline PyTorch distribution for years. That was always six to nine months behind any new PyTorch releases.”
Thomas Sohmers Aug 18, 2025 ▶ 20:06
Insight
Sohmers: Requiring workload recompilation creates fatal friction for AI chip adoption
“If you are requiring a user or having yourself as the company needing to actually recompile a workload, that's already one step too far, even if you assume it works perfectly.”
Thomas Sohmers Aug 18, 2025 ▶ 21:05
Assertion Supported
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Thomas Sohmers Aug 18, 2025 ▶ 21:32
Prediction Not checkable as stated
Sohmers: People will continue buying NVIDIA for AI training
“We are betting that people are going to continue to train on NVIDIA for at least the foreseeable future, where, since we're able to, you know, and I'll say, I really hope others are able to be successful in, in providing competition against NVIDIA, but Given t…”
Thomas Sohmers Aug 18, 2025 ▶ 22:17
Assertion Supported
Sohmers: Positron AI FPGA cards consume only 150 watts
“Like our cards are only using a 150 watts.”
Thomas Sohmers Aug 18, 2025 ▶ 24:08
Prediction Open · timeframe Dec 2027
Agrawal: Positron ASIC will lead all silicon in memory capacity by 2027
“So we are going to be coming out with our ASIC and then like later, it will have more memory capacity than any other silicon in late 2026 or in 27, actually.”
Mitesh Agrawal Aug 18, 2025 ▶ 26:43
Assertion Supported
Sohmers: NVIDIA's TF32 is actually a 19-bit precision format
“NVIDIA's TF-thirty-two number format is a nineteen-bit number format. They just call it thirty-two-bit.”
Thomas Sohmers Aug 18, 2025 ▶ 27:29
Opinion
Sohmers: Hardening silicon for specific AI models is obsolete in months
“Doing any of that, like, hardening for specific Model things. I don't think lasts more than, you know, two or three months at the rate that the industry moves at.”
Thomas Sohmers Aug 18, 2025 ▶ 30:11
Prediction Not checkable as stated
Sohmers: NVIDIA will have a very good decade ahead despite startup challengers
“The reality is NVIDIA is going to have a very, very good decade ahead for them. And The market is growing so fast that all of us in the space trying to take them on can be very happy with, you know, very, very small wins in the space.”
Thomas Sohmers Aug 18, 2025 ▶ 34:26
Disclosure
Agrawal: Positron cannot beat NVIDIA on matrix-matrix performance per dollar today
“Thomas said that, you know, we are accelerating matrix vector. It doesn't mean that we can't do matrix, matrix. We can, and we can do it fairly well. It's just that the point becomes is like, are you really beating NVIDIA on it from a perf per dollar? And if y…”
Mitesh Agrawal Aug 18, 2025 ▶ 35:54
Disclosure
Agrawal: Positron AI counts Cloudflare and Parasale as early hardware customers
“So we have customers in Cloudflare and Parasale, both using us Cloudflare because of performance per watt.”
Mitesh Agrawal Aug 18, 2025 ▶ 37:54
Disclosure
Agrawal: Positron AI raised a $51M Series A for next-gen silicon
“And then we just did a series A close with Valor, Atreides, and DFJ Growth for a roughly fifty one million dollar series A.”
Mitesh Agrawal Aug 18, 2025 ▶ 38:34
Opinion
Agrawal: Cerebras and Groq offer cloud APIs because their software struggles
“You know, if you look at Cerebrus and Grok and others, they've really tried to do this, their cloud kind of portal. And for us, you know, it shows two things. One is the difficulty of software for third party to implement that, that that's why they're kind of …”
Mitesh Agrawal Aug 18, 2025 ▶ 39:13
Assertion Not checkable as stated
Sohmers: Reasoning models shift inference workloads to 100 output tokens per input
“If you go back a year from today in July of last year, the ratios of like input to output for LLMs were very, very heavily on, on inputs where you could be doing, you know, 10, 10, 15 to one ratio of input to output. But that has completely flipped and it's ob…”
Thomas Sohmers Aug 18, 2025 ▶ 42:06
Assertion Contradicted
Sohmers: Google Veo and Imagen 3 are pure autoregressive transformers, not diffusion
“A lot of things have actually been moving away from diffusion to being pure autoregressive transformers for image and video generation. So like the latest, yeah, there's a VO three and since image and three on, on Google side have been pure autoregressive movi…”
Thomas Sohmers Aug 18, 2025 ▶ 43:48
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.