Jun 13, 2025 · 1h 18m · latent-space

The Shape of Compute (Chris Lattner of Modular)

Chris Lattner · 1h 1m spoken Shawn Wang · 6m spoken Alessio Fanelli · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this in-depth discussion, Modular founder and CEO Chris Lattner joins Alessio Fanelli and Shawn 'swyx' Wang to explore the architecture of Mojo and MAX, the shifting economics of AI inference, and the engineering principles required to build portable, high-performance compute infrastructure.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 13.4% of the talking time here. How this is scored →

The hosts as informed peer 4.5 Guest teaching 5.9 Guest disagreement 2.4 The hosts pushing back 1.6
05100:0020:0040:001:00:000:05–6:55 · The hosts as informed peer 3/10 Modular's Three-Year Evolution from R&D to Product Execution Swyx and Alessio set up the context around Modular's progress. Chris provides a deep dive into Modular's first three stealth R&D years, explaining why they refused to use CUDA and set a high bar of beating NVIDIA on LLaMA 3 serving before publicly launching.6:55–11:15 · The hosts as informed peer 4/10 Engineering Roadmap: From CPU Compilers to GPU LLM Serving Alessio asks about moving from CPU compilers to GPUs. When Swyx questions whether industry naysayers meant 'impossible or very hard', Chris details how conventional wisdom repeatedly declared LLVM, Swift, and MLIR impossible before their adoption.11:15–23:02 · The hosts as informed peer 5/10 Architectural Breakdown: The MAX Engine and Mojo Concentric Circles Alessio prompts Chris to walk through the architecture of the MAX engine. Chris details the concentric circle model starting from Mojo, offering a sharp critique of C++ complexity and explaining binding-free Python acceleration.23:02–32:25 · The hosts as informed peer 5/10 Inference Engines and the Power of Modular Software Architecture Swyx asks about community rivalries between vLLM and SGLang and whether modular design entails architectural trade-offs. Chris explains how monolithic frameworks become disposable while composable systems survive hardware shifts.32:25–37:08 · The hosts as informed peer 4/10 Democratizing Inference and Fostering Ecosystem Contributions Swyx notes the democratization theme in Modular's mission. Chris reframes the history, noting that PyTorch successfully democratized training in 2017 but inference remained an opaque black art reserved for big tech labs.37:08–44:29 · The hosts as informed peer 4/10 Modular's Business Strategy and Cloud GPU Economics Swyx inquires about Modular's commercial monetization model versus cloud inference endpoints. Chris breaks down the economic divergence between stateless CPU cloud scaling and stateful GPU reserved commitments facing rapid hardware obsolescence.44:29–48:50 · The hosts as informed peer 4/10 Lessons in Language Launching: Comparing Swift and Mojo Swyx asks for a comparison between rolling out Swift at Apple and launching Mojo. Chris politely restates the prompt into a lessons-learned framing, contrasting Swift's secret 1.0 NDA launch with Mojo's transparent 0.1 dogfooding.48:50–58:35 · The hosts as informed peer 6/10 Big Tech AI Dynamics, Google's Legacy, and the DeepSeek Catalyst Alessio and Swyx guide the conversation to Google's AI history and DeepSeek's open source shockwave. Chris explains how DeepSeek reverse engineering PTX tensor core instructions mirrors standard practices of top compiler engineers and reveals why Blackwell kernel rewrites are needed.58:35–1:02:32 · The hosts as informed peer 7/10 Reasoning Models, Reinforcement Learning, and the Centrality of Inference Swyx demonstrates strong domain knowledge by noting that reasoning models and reinforcement learning collapse the distinction between training and inference. Chris affirms the insight, connecting it back to early DeepMind AlphaGo infrastructure challenges.1:02:32–1:09:08 · The hosts as informed peer 3/10 Founder Leadership, Team Scaling, and Daily Productivity Routines Alessio asks Chris about his founder routine and managing startup team dynamics. Chris reflects candidly on personal emotional adjustments when running a startup versus scaling teams in large tech corporations.1:09:08–1:14:37 · The hosts as informed peer 4/10 AI Coding Agents, The Role of Human Code, and Research Tracking Swyx asks how AI coding assistants influence language adoption and compiler writing. Chris argues that programming languages remain fundamental because code exists primarily for human reasoning and constraint management rather than machine execution alone.0:05–6:55 · Guest teaching 6/10 Modular's Three-Year Evolution from R&D to Product Execution Swyx and Alessio set up the context around Modular's progress. Chris provides a deep dive into Modular's first three stealth R&D years, explaining why they refused to use CUDA and set a high bar of beating NVIDIA on LLaMA 3 serving before publicly launching.6:55–11:15 · Guest teaching 7/10 Engineering Roadmap: From CPU Compilers to GPU LLM Serving Alessio asks about moving from CPU compilers to GPUs. When Swyx questions whether industry naysayers meant 'impossible or very hard', Chris details how conventional wisdom repeatedly declared LLVM, Swift, and MLIR impossible before their adoption.11:15–23:02 · Guest teaching 7/10 Architectural Breakdown: The MAX Engine and Mojo Concentric Circles Alessio prompts Chris to walk through the architecture of the MAX engine. Chris details the concentric circle model starting from Mojo, offering a sharp critique of C++ complexity and explaining binding-free Python acceleration.23:02–32:25 · Guest teaching 6/10 Inference Engines and the Power of Modular Software Architecture Swyx asks about community rivalries between vLLM and SGLang and whether modular design entails architectural trade-offs. Chris explains how monolithic frameworks become disposable while composable systems survive hardware shifts.32:25–37:08 · Guest teaching 6/10 Democratizing Inference and Fostering Ecosystem Contributions Swyx notes the democratization theme in Modular's mission. Chris reframes the history, noting that PyTorch successfully democratized training in 2017 but inference remained an opaque black art reserved for big tech labs.37:08–44:29 · Guest teaching 7/10 Modular's Business Strategy and Cloud GPU Economics Swyx inquires about Modular's commercial monetization model versus cloud inference endpoints. Chris breaks down the economic divergence between stateless CPU cloud scaling and stateful GPU reserved commitments facing rapid hardware obsolescence.44:29–48:50 · Guest teaching 6/10 Lessons in Language Launching: Comparing Swift and Mojo Swyx asks for a comparison between rolling out Swift at Apple and launching Mojo. Chris politely restates the prompt into a lessons-learned framing, contrasting Swift's secret 1.0 NDA launch with Mojo's transparent 0.1 dogfooding.48:50–58:35 · Guest teaching 7/10 Big Tech AI Dynamics, Google's Legacy, and the DeepSeek Catalyst Alessio and Swyx guide the conversation to Google's AI history and DeepSeek's open source shockwave. Chris explains how DeepSeek reverse engineering PTX tensor core instructions mirrors standard practices of top compiler engineers and reveals why Blackwell kernel rewrites are needed.58:35–1:02:32 · Guest teaching 4/10 Reasoning Models, Reinforcement Learning, and the Centrality of Inference Swyx demonstrates strong domain knowledge by noting that reasoning models and reinforcement learning collapse the distinction between training and inference. Chris affirms the insight, connecting it back to early DeepMind AlphaGo infrastructure challenges.1:02:32–1:09:08 · Guest teaching 4/10 Founder Leadership, Team Scaling, and Daily Productivity Routines Alessio asks Chris about his founder routine and managing startup team dynamics. Chris reflects candidly on personal emotional adjustments when running a startup versus scaling teams in large tech corporations.1:09:08–1:14:37 · Guest teaching 5/10 AI Coding Agents, The Role of Human Code, and Research Tracking Swyx asks how AI coding assistants influence language adoption and compiler writing. Chris argues that programming languages remain fundamental because code exists primarily for human reasoning and constraint management rather than machine execution alone.0:05–6:55 · Guest disagreement 2/10 Modular's Three-Year Evolution from R&D to Product Execution Swyx and Alessio set up the context around Modular's progress. Chris provides a deep dive into Modular's first three stealth R&D years, explaining why they refused to use CUDA and set a high bar of beating NVIDIA on LLaMA 3 serving before publicly launching.6:55–11:15 · Guest disagreement 4/10 Engineering Roadmap: From CPU Compilers to GPU LLM Serving Alessio asks about moving from CPU compilers to GPUs. When Swyx questions whether industry naysayers meant 'impossible or very hard', Chris details how conventional wisdom repeatedly declared LLVM, Swift, and MLIR impossible before their adoption.11:15–23:02 · Guest disagreement 3/10 Architectural Breakdown: The MAX Engine and Mojo Concentric Circles Alessio prompts Chris to walk through the architecture of the MAX engine. Chris details the concentric circle model starting from Mojo, offering a sharp critique of C++ complexity and explaining binding-free Python acceleration.23:02–32:25 · Guest disagreement 3/10 Inference Engines and the Power of Modular Software Architecture Swyx asks about community rivalries between vLLM and SGLang and whether modular design entails architectural trade-offs. Chris explains how monolithic frameworks become disposable while composable systems survive hardware shifts.32:25–37:08 · Guest disagreement 2/10 Democratizing Inference and Fostering Ecosystem Contributions Swyx notes the democratization theme in Modular's mission. Chris reframes the history, noting that PyTorch successfully democratized training in 2017 but inference remained an opaque black art reserved for big tech labs.37:08–44:29 · Guest disagreement 2/10 Modular's Business Strategy and Cloud GPU Economics Swyx inquires about Modular's commercial monetization model versus cloud inference endpoints. Chris breaks down the economic divergence between stateless CPU cloud scaling and stateful GPU reserved commitments facing rapid hardware obsolescence.44:29–48:50 · Guest disagreement 3/10 Lessons in Language Launching: Comparing Swift and Mojo Swyx asks for a comparison between rolling out Swift at Apple and launching Mojo. Chris politely restates the prompt into a lessons-learned framing, contrasting Swift's secret 1.0 NDA launch with Mojo's transparent 0.1 dogfooding.48:50–58:35 · Guest disagreement 3/10 Big Tech AI Dynamics, Google's Legacy, and the DeepSeek Catalyst Alessio and Swyx guide the conversation to Google's AI history and DeepSeek's open source shockwave. Chris explains how DeepSeek reverse engineering PTX tensor core instructions mirrors standard practices of top compiler engineers and reveals why Blackwell kernel rewrites are needed.58:35–1:02:32 · Guest disagreement 1/10 Reasoning Models, Reinforcement Learning, and the Centrality of Inference Swyx demonstrates strong domain knowledge by noting that reasoning models and reinforcement learning collapse the distinction between training and inference. Chris affirms the insight, connecting it back to early DeepMind AlphaGo infrastructure challenges.1:02:32–1:09:08 · Guest disagreement 1/10 Founder Leadership, Team Scaling, and Daily Productivity Routines Alessio asks Chris about his founder routine and managing startup team dynamics. Chris reflects candidly on personal emotional adjustments when running a startup versus scaling teams in large tech corporations.1:09:08–1:14:37 · Guest disagreement 2/10 AI Coding Agents, The Role of Human Code, and Research Tracking Swyx asks how AI coding assistants influence language adoption and compiler writing. Chris argues that programming languages remain fundamental because code exists primarily for human reasoning and constraint management rather than machine execution alone.0:05–6:55 · The hosts pushing back 1/10 Modular's Three-Year Evolution from R&D to Product Execution Swyx and Alessio set up the context around Modular's progress. Chris provides a deep dive into Modular's first three stealth R&D years, explaining why they refused to use CUDA and set a high bar of beating NVIDIA on LLaMA 3 serving before publicly launching.6:55–11:15 · The hosts pushing back 3/10 Engineering Roadmap: From CPU Compilers to GPU LLM Serving Alessio asks about moving from CPU compilers to GPUs. When Swyx questions whether industry naysayers meant 'impossible or very hard', Chris details how conventional wisdom repeatedly declared LLVM, Swift, and MLIR impossible before their adoption.11:15–23:02 · The hosts pushing back 2/10 Architectural Breakdown: The MAX Engine and Mojo Concentric Circles Alessio prompts Chris to walk through the architecture of the MAX engine. Chris details the concentric circle model starting from Mojo, offering a sharp critique of C++ complexity and explaining binding-free Python acceleration.23:02–32:25 · The hosts pushing back 2/10 Inference Engines and the Power of Modular Software Architecture Swyx asks about community rivalries between vLLM and SGLang and whether modular design entails architectural trade-offs. Chris explains how monolithic frameworks become disposable while composable systems survive hardware shifts.32:25–37:08 · The hosts pushing back 1/10 Democratizing Inference and Fostering Ecosystem Contributions Swyx notes the democratization theme in Modular's mission. Chris reframes the history, noting that PyTorch successfully democratized training in 2017 but inference remained an opaque black art reserved for big tech labs.37:08–44:29 · The hosts pushing back 1/10 Modular's Business Strategy and Cloud GPU Economics Swyx inquires about Modular's commercial monetization model versus cloud inference endpoints. Chris breaks down the economic divergence between stateless CPU cloud scaling and stateful GPU reserved commitments facing rapid hardware obsolescence.44:29–48:50 · The hosts pushing back 2/10 Lessons in Language Launching: Comparing Swift and Mojo Swyx asks for a comparison between rolling out Swift at Apple and launching Mojo. Chris politely restates the prompt into a lessons-learned framing, contrasting Swift's secret 1.0 NDA launch with Mojo's transparent 0.1 dogfooding.48:50–58:35 · The hosts pushing back 2/10 Big Tech AI Dynamics, Google's Legacy, and the DeepSeek Catalyst Alessio and Swyx guide the conversation to Google's AI history and DeepSeek's open source shockwave. Chris explains how DeepSeek reverse engineering PTX tensor core instructions mirrors standard practices of top compiler engineers and reveals why Blackwell kernel rewrites are needed.58:35–1:02:32 · The hosts pushing back 2/10 Reasoning Models, Reinforcement Learning, and the Centrality of Inference Swyx demonstrates strong domain knowledge by noting that reasoning models and reinforcement learning collapse the distinction between training and inference. Chris affirms the insight, connecting it back to early DeepMind AlphaGo infrastructure challenges.1:02:32–1:09:08 · The hosts pushing back 1/10 Founder Leadership, Team Scaling, and Daily Productivity Routines Alessio asks Chris about his founder routine and managing startup team dynamics. Chris reflects candidly on personal emotional adjustments when running a startup versus scaling teams in large tech corporations.1:09:08–1:14:37 · The hosts pushing back 1/10 AI Coding Agents, The Role of Human Code, and Research Tracking Swyx asks how AI coding assistants influence language adoption and compiler writing. Chris argues that programming languages remain fundamental because code exists primarily for human reasoning and constraint management rather than machine execution alone.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 36.4% · guest 63.6%0:00 · the hosts 36.4% · guest 63.6%3:00 · the hosts 1.2% · guest 98.8%3:00 · the hosts 1.2% · guest 98.8%6:00 · the hosts 7% · guest 93%6:00 · the hosts 7% · guest 93%9:00 · the hosts 7% · guest 93%9:00 · the hosts 7% · guest 93%12:00 · the hosts 2.7% · guest 97.3%12:00 · the hosts 2.7% · guest 97.3%15:00 · the hosts 8.9% · guest 91.1%15:00 · the hosts 8.9% · guest 91.1%18:00 · the hosts 4.5% · guest 95.5%18:00 · the hosts 4.5% · guest 95.5%21:00 · the hosts 9.3% · guest 90.7%21:00 · the hosts 9.3% · guest 90.7%24:00 · the hosts 15.4% · guest 84.6%24:00 · the hosts 15.4% · guest 84.6%27:00 · the hosts 6.5% · guest 93.5%27:00 · the hosts 6.5% · guest 93.5%30:00 · the hosts 3.7% · guest 96.3%30:00 · the hosts 3.7% · guest 96.3%33:00 · the hosts 11.2% · guest 88.8%33:00 · the hosts 11.2% · guest 88.8%36:00 · the hosts 25.4% · guest 74.6%36:00 · the hosts 25.4% · guest 74.6%39:00 · the hosts 4.2% · guest 95.8%39:00 · the hosts 4.2% · guest 95.8%42:00 · the hosts 17% · guest 83%42:00 · the hosts 17% · guest 83%45:00 · the hosts 2.9% · guest 97.1%45:00 · the hosts 2.9% · guest 97.1%48:00 · the hosts 16.4% · guest 83.6%48:00 · the hosts 16.4% · guest 83.6%51:00 · the hosts 43.5% · guest 56.5%51:00 · the hosts 43.5% · guest 56.5%54:00 · the hosts 0.6% · guest 99.4%54:00 · the hosts 0.6% · guest 99.4%57:00 · the hosts 27.1% · guest 72.9%57:00 · the hosts 27.1% · guest 72.9%1:00:00 · the hosts 39% · guest 61%1:00:00 · the hosts 39% · guest 61%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 21.3% · guest 78.7%1:06:00 · the hosts 21.3% · guest 78.7%1:09:00 · the hosts 17.6% · guest 82.4%1:09:00 · the hosts 17.6% · guest 82.4%1:12:00 · the hosts 18.4% · guest 81.6%1:12:00 · the hosts 18.4% · guest 81.6%1:15:00 · the hosts 3.8% · guest 96.2%1:15:00 · the hosts 3.8% · guest 96.2%1:18:00 · the hosts 81.3% · guest 18.7%1:18:00 · the hosts 81.3% · guest 18.7%
Sharpest disagreement ▶ 9:00 Chris dismisses the narrative that replacing CUDA is impossible

Chris forcefully rejects the conventional tech wisdom that a startup cannot displace CUDA, citing his career track record overcoming identical skepticism with LLVM, Swift, and MLIR.

Hardest push from the hosts ▶ 8:57 Swyx presses Chris on the definition of impossible

Swyx directly challenges Chris's framing of industry critics, pressing him to clarify whether commentators literally meant impossible or simply very difficult.

Biggest teaching moment ▶ 38:20 Chris breaks down the economics of GPU vs CPU cloud infrastructure

Chris educates the hosts on the structural difference between stateless elastic CPU hosting and stateful GPU capacity planning where hardware depreciates before long-term commits end.

The host holds their own ▶ 1:00:45 Swyx links chain-of-thought inference directly into training loops

Swyx demonstrates technical command by pointing out that modern test-time compute and RL make inference an essential phase of model training, validating Modular's architectural focus.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Modular's Three-Year Evolution from R&D to Product Execution 3621 Swyx and Alessio set up the context around Modular's progress. Chris provides a deep dive into Modular's first three stealth R&D years, explaining why they refused to use CUDA and set a high bar of beating NVIDIA on LLaMA 3 serving before publicly launching.
Engineering Roadmap: From CPU Compilers to GPU LLM Serving 4743 Alessio asks about moving from CPU compilers to GPUs. When Swyx questions whether industry naysayers meant 'impossible or very hard', Chris details how conventional wisdom repeatedly declared LLVM, Swift, and MLIR impossible before their adoption.
Architectural Breakdown: The MAX Engine and Mojo Concentric Circles 5732 Alessio prompts Chris to walk through the architecture of the MAX engine. Chris details the concentric circle model starting from Mojo, offering a sharp critique of C++ complexity and explaining binding-free Python acceleration.
Inference Engines and the Power of Modular Software Architecture 5632 Swyx asks about community rivalries between vLLM and SGLang and whether modular design entails architectural trade-offs. Chris explains how monolithic frameworks become disposable while composable systems survive hardware shifts.
Democratizing Inference and Fostering Ecosystem Contributions 4621 Swyx notes the democratization theme in Modular's mission. Chris reframes the history, noting that PyTorch successfully democratized training in 2017 but inference remained an opaque black art reserved for big tech labs.
Modular's Business Strategy and Cloud GPU Economics 4721 Swyx inquires about Modular's commercial monetization model versus cloud inference endpoints. Chris breaks down the economic divergence between stateless CPU cloud scaling and stateful GPU reserved commitments facing rapid hardware obsolescence.
Lessons in Language Launching: Comparing Swift and Mojo 4632 Swyx asks for a comparison between rolling out Swift at Apple and launching Mojo. Chris politely restates the prompt into a lessons-learned framing, contrasting Swift's secret 1.0 NDA launch with Mojo's transparent 0.1 dogfooding.
Big Tech AI Dynamics, Google's Legacy, and the DeepSeek Catalyst 6732 Alessio and Swyx guide the conversation to Google's AI history and DeepSeek's open source shockwave. Chris explains how DeepSeek reverse engineering PTX tensor core instructions mirrors standard practices of top compiler engineers and reveals why Blackwell kernel rewrites are needed.
Reasoning Models, Reinforcement Learning, and the Centrality of Inference 7412 Swyx demonstrates strong domain knowledge by noting that reasoning models and reinforcement learning collapse the distinction between training and inference. Chris affirms the insight, connecting it back to early DeepMind AlphaGo infrastructure challenges.
Founder Leadership, Team Scaling, and Daily Productivity Routines 3411 Alessio asks Chris about his founder routine and managing startup team dynamics. Chris reflects candidly on personal emotional adjustments when running a startup versus scaling teams in large tech corporations.
AI Coding Agents, The Role of Human Code, and Research Tracking 4521 Swyx asks how AI coding assistants influence language adoption and compiler writing. Chris argues that programming languages remain fundamental because code exists primarily for human reasoning and constraint management rather than machine execution alone.

Statements from this episode (23)

Insight
Alternative AI stacks must match state-of-the-art performance to be useful
“Building An alternate AI stack that is not as good as the existing one isn't very useful because people will always evaluate it against state of the art.”
Chris Lattner Jun 13, 2025 ▶ 2:05
Disclosure
Modular will soon launch support for AMD MI300 and MI325 GPUs
“We're about to launch our AMD and my 303 25 support. That'll be a big deal for the industry.”
Chris Lattner Jun 13, 2025 ▶ 5:21
Assertion Partly supported
First-time Mojo developers built a GPU training system in one day
“The winning team for the hackathon took that four person team. And one day they had not used Mojo before they hadn't programmed GPUs before. And they had built a training system. They wrote an atom optimizer, a bunch of training kernels. They built a simple ba…”
Chris Lattner Jun 13, 2025 ▶ 6:06
Assertion Supported
Modular beat Intel MKL matrix multiplication speed in its first year
“First year was prove compilation philosophy. And so this was writing very abstract compiler stuff. And then prove that we can make a matrix multiplication go faster than Intel MKL's matrix multiplication on Intel Silicon and make it configurable and multiple D…”
Chris Lattner Jun 13, 2025 ▶ 7:06
Opinion
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Chris Lattner Jun 13, 2025 ▶ 12:25
Opinion
Lattner says C++ 'sucks' for modern accelerated compute
“Let me be the first to tell you, and I can say this now, I feel comfortable saying this, that C++ sucks.”
Chris Lattner Jun 13, 2025 ▶ 13:38
Assertion Not checkable as stated
Lattner says Mojo is 10,000x faster than Python and beats Rust
“Mojo is not just a little faster than Python, it's faster than Rust. So it's like tens of thousands of times faster than Python, and it's in the Python family.”
Chris Lattner Jun 13, 2025 ▶ 15:29
Opinion
Lattner calls vLLM a 'hot mess' due to too many stakeholders
“VLM seems much more like a massive community with a lot of stakeholders, a lot of stuff going on, and it's kind of a hot mess.”
Chris Lattner Jun 13, 2025 ▶ 23:52
Assertion Not publicly verifiable
Modular's Mojo FlashAttention beats Tri Dao's reference implementation
“We're beating the tree DAO reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our, GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks, r…”
Chris Lattner Jun 13, 2025 ▶ 30:59
Disclosure
Modular's team has grown to over 100 employees
“So we're a small team. I mean, we're over a hundred people, but we're compared to the size of the problem we're taking on.”
Chris Lattner Jun 13, 2025 ▶ 32:26
Opinion
PyTorch democratized model training, but AI inference remains undemocratized
“And so things like PyTorch came on the scene and I think PyTorch gets all credit for democratizing model training, right? It's taught to pretty much every computer science student that graduates. That's a huge deal, but nobody democratized inference. Inference…”
Chris Lattner Jun 13, 2025 ▶ 35:44
Assertion Not checkable as stated
Intelligent compute routing delivers 5x performance TCO benefit
“You can get like five X performance TCO benefit by doing intelligent routing.”
Chris Lattner Jun 13, 2025 ▶ 39:13
Disclosure
Modular's Mojo and MAX are free on NVIDIA and CPUs
“The max framework and the mojo language, free to use on NVIDIA and CPUs, any scale, go nuts, do whatever you want.”
Chris Lattner Jun 13, 2025 ▶ 41:47
Disclosure
Modular monetizes cluster management and support on a per-GPU basis
“If you want cluster management, and you want enterprise support, and you want things like this, then you can pay on a per GPU basis, and you can contact our sales team, and then we can work out a deal, and we can work with you directly.”
Chris Lattner Jun 13, 2025 ▶ 42:19
Prediction Not checkable as stated
An industry-leading state-of-the-art AI model may launch on MAX soon
“We may have an industry leading state of the art model launching first on Mac soon.”
Chris Lattner Jun 13, 2025 ▶ 43:16
Assertion Not checkable as stated
Only 250 people knew about Apple's Swift language before its 2014 launch
“Started in 2010, it launched publicly by Apple in 2014, ok? And by the time it launched in 2014, only about 250 people in the world knew about it. Most of whom were in my team, about 200 and something of them were in my team, and then it was senior execs, mark…”
Chris Lattner Jun 13, 2025 ▶ 45:18
Assertion Partly supported
Modular has open-sourced 650,000 lines of Mojo code
“Modular is Mojo's first customer. Like we have more Mojo code in our repository than any other language. And it's open source. Like, and we open source like 650,000 lines of Mojo code.”
Chris Lattner Jun 13, 2025 ▶ 47:44
Assertion Supported
NVIDIA Blackwell is not backward-compatible with Hopper kernels
“Well, so it's not widely known, but Blackwell is not compatible with Hopper. Hopper kernels don't always run on Blackwell, for example.”
Chris Lattner Jun 13, 2025 ▶ 53:18
Disclosure
Modular completely eliminates and replaces NVIDIA's CUDA stack
“In the case of Modular, we go literally, like, we only work at that level, because we get rid of all of CUDA, right? And so we've replaced the entire stack, and so we only do that.”
Chris Lattner Jun 13, 2025 ▶ 55:03
Opinion
DeepSeek's open releases pulled global AI progress forward by six months
“But I think that what it did is it pulled forward progress in AI by like six months.”
Chris Lattner Jun 13, 2025 ▶ 57:54
Insight
AI training scales with research teams; inference scales with customer base
“Training scales the size of your research team. Inference scales the size of your customer base.”
Chris Lattner Jun 13, 2025 ▶ 59:12
Assertion Supported
CPU latency in KV cache management bottlenecks GPU utilization
“Like your eviction policy runs on a CPU. Like that radix hashing algorithm and block hashing and all that stuff happens like primarily CPU. That's really important for performance because if you have latency in these steps, like you're not keeping your GPU uti…”
Chris Lattner Jun 13, 2025 ▶ 1:00:09
Insight
AI coding assistants excel at boilerplate, not inventing new algorithms
“People within the company are dabbling with cloud code and some of the stuff, and so I don't have personal experience with that, but for me, I found that it's mostly, it is very useful, but it's really about boilerplate, and so it's not about inventing new alg…”
Chris Lattner Jun 13, 2025 ▶ 1:10:18
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.