People, every show

Misha Smulyanski

Co-founder, Marlowe. On 1 show, 1 appearance. The Shows tab opens the full record on each.

founderengineerscientistexecutiveLinkedIn ↗

Misha Smelyanskiy is a computer architect and AI systems expert who previously served as Senior Director of AI Infrastructure at NVIDIA and Director at Meta. He has led datacenter-scale AI hardware-software co-design and co-founded Marlowe to develop heterogeneous AI inference infrastructure.

1shows
1appearances
4statements
1resolved
1supported
0contradicted
100%fully supported

Everything Misha Smulyanski said on any show that made the record, most notable first. Each card names its show and opens the statement there.

Smulyanski: AI inference demands heterogeneous hardware co-designed for different phases
“Inference Is a very heterogeneous workload, right? Different phases of inference exercise, compute, network, storage, memory bandwidths differently, and so when we look at it makes sense to actually co-design the systems that will opt to, you know, use differe…”
Misha Smulyanski Jul 29, 2026 ▶ 47:35 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Smulyanski: On-die SRAM accelerators excel at LLM decode due to high bandwidth
“The SRA machine basically keeps the entire weight matrix in SRA memory on DAI, so the, you got a lot more bandwidth, right, because it's on chip, right, so you can access, you know bytes over cycles, right, the chip interconnect is also fast, you can go, like …”
Misha Smulyanski Jul 29, 2026 ▶ 55:29 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Y COMBINATOR Assertion Supported
Smulyanski: LLM decode workloads remain bandwidth-bound even across large batch sizes
“Pre-fill is generally very compute bound. Right because you basically do, ah, like, attention, ah, you do, ah, ah, a lot of, ah work for, you know, for, you know, all the tokens that you are fetching, right? You can, ah, you fetch the weights while I'm once an…”
Misha Smulyanski Jul 29, 2026 ▶ 51:45 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
Smulyanski: GPU throughput drops sharply at low concurrency due to kernel overheads
“The moment that you start basically going to lower concurrency because you want better interactivity and better latency, Right? The performance the throughput drops. And it drops very sharply because all of a sudden you have a lot of, like, smaller kernels, yo…”
Misha Smulyanski Jul 29, 2026 ▶ 59:40 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator

One line per show, most statements first. The link opens Misha's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
Y COMBINATORLEDGER Co-founder, Marlowe 1 4 100% 1/1 full record on the Y Combinator Startup Podcast →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.