Thomas Sohmers

Co-Founder and CTO, Positron AI · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderengineerexecutive@trsohmers ↗LinkedIn ↗trsohmers.com ↗

A 2013 Thiel Fellow who began computing research at MIT at age 14, Thomas Sohmers previously founded REX Computing and worked at Groq and Lambda Labs. In 2023, he co-founded Positron AI to develop energy-efficient hardware accelerators for AI inference workloads.

17statements → 12claims → 7claims resolved → 57%fully supported → 3.94/5average certainty → 2.41/5average debate potential →

4 supported 0 partly supported 3 contradicted 2 not yet assessed 3 not checkable as stated how the 12 claims stand · each chip opens the sources

3 predictions · 9 assertions · 1 opinion · 4 insights · every statement was checked. The predictions and assertions are the 12 claims: statements the public record can support or contradict. 7 are resolved, 2 are not yet assessed, and 3 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Thomas argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Thomas Sohmers Aug 18, 2025 ▶ 21:32 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI

Their most notable contradicted claim

Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Thomas Sohmers Aug 18, 2025 ▶ 14:53 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
0% certainty 3
60% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Thomas Sohmers on measured tape to publish a rate. This says nothing about how they speak.

Everything Thomas Sohmers said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Opinion
Sohmers: Hardening silicon for specific AI models is obsolete in months
“Doing any of that, like, hardening for specific Model things. I don't think lasts more than, you know, two or three months at the rate that the industry moves at.”
Thomas Sohmers Aug 18, 2025 ▶ 30:11 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Insight
Sohmers: AI hardware over-indexes on raw FLOPS instead of memory bandwidth
“Everyone else was focusing on the wrong things. They were just trying to have more and more flops when memory bandwidth, memory capacity were the real, real bottlenecks.”
Thomas Sohmers Aug 18, 2025 ▶ 6:26 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Thomas Sohmers Aug 18, 2025 ▶ 14:53 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Thomas Sohmers Aug 18, 2025 ▶ 16:28 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Insight
Sohmers: Requiring workload recompilation creates fatal friction for AI chip adoption
“If you are requiring a user or having yourself as the company needing to actually recompile a workload, that's already one step too far, even if you assume it works perfectly.”
Thomas Sohmers Aug 18, 2025 ▶ 21:05 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Supported
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Thomas Sohmers Aug 18, 2025 ▶ 21:32 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Contradicted
Sohmers: Google Veo and Imagen 3 are pure autoregressive transformers, not diffusion
“A lot of things have actually been moving away from diffusion to being pure autoregressive transformers for image and video generation. So like the latest, yeah, there's a VO three and since image and three on, on Google side have been pure autoregressive movi…”
Thomas Sohmers Aug 18, 2025 ▶ 43:48 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Open · timeframe Aug 2026
Sohmers: Positron AI hardware achieves 93% of theoretical memory bandwidth
“And so our fundamental architecture is enabling us, you know, today with hardware that we're shipping right now to be achieving, you know, 93% of the theoretical memory bandwidth of our device consistently across all use cases.”
Thomas Sohmers Aug 18, 2025 ▶ 15:26 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Insight
Sohmers: Transformer inference is memory-bound with a 1:1 FLOP-to-byte ratio
“And on the other side of this chart, you have the case of a transformer where when you're actually, you know, doing attention, Or really just any case where you're doing, you're fundamentally doing matrix vector multiplication rather than matrix matrix multipl…”
Thomas Sohmers Aug 18, 2025 ▶ 9:27 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Contradicted
Sohmers: Cray-2 was the last major system with balanced memory-to-compute ratio
“And if you look at sort of traditional big iron compute systems, the last like major compute, compute platform that had that balance of memory to compute ratio was the Cray two supercomputer.”
Thomas Sohmers Aug 18, 2025 ▶ 10:39 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Insight
Sohmers: Matrix-vector multiplication in transformer inference is fundamentally uncacheable
“So the second level of this is that matrix vector multiplication is basically uncacheable. When you're doing transformer inference, matrix A is the weights of your model. And so if you're talking about model weights that are tens of gigabytes, hundreds of giga…”
Thomas Sohmers Aug 18, 2025 ▶ 13:33 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Supported
Sohmers: AMD's PyTorch fork lagged official releases by 6-9 months
“And AMD had their own separate you know, non-mainline PyTorch distribution for years. That was always six to nine months behind any new PyTorch releases.”
Thomas Sohmers Aug 18, 2025 ▶ 20:06 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Prediction Not checkable as stated
Sohmers: People will continue buying NVIDIA for AI training
“We are betting that people are going to continue to train on NVIDIA for at least the foreseeable future, where, since we're able to, you know, and I'll say, I really hope others are able to be successful in, in providing competition against NVIDIA, but Given t…”
Thomas Sohmers Aug 18, 2025 ▶ 22:17 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Supported
Sohmers: NVIDIA's TF32 is actually a 19-bit precision format
“NVIDIA's TF-thirty-two number format is a nineteen-bit number format. They just call it thirty-two-bit.”
Thomas Sohmers Aug 18, 2025 ▶ 27:29 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Prediction Not checkable as stated
Sohmers: NVIDIA will have a very good decade ahead despite startup challengers
“The reality is NVIDIA is going to have a very, very good decade ahead for them. And The market is growing so fast that all of us in the space trying to take them on can be very happy with, you know, very, very small wins in the space.”
Thomas Sohmers Aug 18, 2025 ▶ 34:26 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Not checkable as stated
Sohmers: Reasoning models shift inference workloads to 100 output tokens per input
“If you go back a year from today in July of last year, the ratios of like input to output for LLMs were very, very heavily on, on inputs where you could be doing, you know, 10, 10, 15 to one ratio of input to output. But that has completely flipped and it's ob…”
Thomas Sohmers Aug 18, 2025 ▶ 42:06 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Supported
Sohmers: Positron AI FPGA cards consume only 150 watts
“Like our cards are only using a 150 watts.”
Thomas Sohmers Aug 18, 2025 ▶ 24:08 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI

Appearances (1)

EpisodeDateSpeaking time
⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, P Aug 18, 2025 28m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.