Misha Smulyanski

Co-founder, Marlowe · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderengineerscientistexecutiveLinkedIn ↗

Misha Smelyanskiy is a computer architect and AI systems expert who previously served as Senior Director of AI Infrastructure at NVIDIA and Director at Meta. He has led datacenter-scale AI hardware-software co-design and co-founded Marlowe to develop heterogeneous AI inference infrastructure.

4statements → 1claims → 1claims resolved → 4/5average certainty → 1.5/5average debate potential →

1 supported 0 partly supported 0 contradicted how the 1 claim stands · each chip opens the sources

1 assertion · 3 insights · every statement was checked. The predictions and assertion are the 1 claim: statements the public record can support or contradict. 1 is resolved. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Misha argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Smulyanski: LLM decode workloads remain bandwidth-bound even across large batch sizes
“Pre-fill is generally very compute bound. Right because you basically do, ah, like, attention, ah, you do, ah, ah, a lot of, ah work for, you know, for, you know, all the tokens that you are fetching, right? You can, ah, you fetch the weights while I'm once an…”
Misha Smulyanski Jul 29, 2026 ▶ 51:45 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator

Everything Misha Smulyanski said on the Y Combinator Startup Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Smulyanski: AI inference demands heterogeneous hardware co-designed for different phases
“Inference Is a very heterogeneous workload, right? Different phases of inference exercise, compute, network, storage, memory bandwidths differently, and so when we look at it makes sense to actually co-design the systems that will opt to, you know, use differe…”
Misha Smulyanski Jul 29, 2026 ▶ 47:35 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Insight
Smulyanski: On-die SRAM accelerators excel at LLM decode due to high bandwidth
“The SRA machine basically keeps the entire weight matrix in SRA memory on DAI, so the, you got a lot more bandwidth, right, because it's on chip, right, so you can access, you know bytes over cycles, right, the chip interconnect is also fast, you can go, like …”
Misha Smulyanski Jul 29, 2026 ▶ 55:29 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Assertion Supported
Smulyanski: LLM decode workloads remain bandwidth-bound even across large batch sizes
“Pre-fill is generally very compute bound. Right because you basically do, ah, like, attention, ah, you do, ah, ah, a lot of, ah work for, you know, for, you know, all the tokens that you are fetching, right? You can, ah, you fetch the weights while I'm once an…”
Misha Smulyanski Jul 29, 2026 ▶ 51:45 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Insight
Smulyanski: GPU throughput drops sharply at low concurrency due to kernel overheads
“The moment that you start basically going to lower concurrency because you want better interactivity and better latency, Right? The performance the throughput drops. And it drops very sharply because all of a sudden you have a lot of, like, smaller kernels, yo…”
Misha Smulyanski Jul 29, 2026 ▶ 59:40 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator

Appearances (1)

EpisodeDateSpeaking time
Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Clu Jul 29, 2026 16m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.