Aug 4, 2025 · 28m · latent-space

⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs

Stefano Ermon · 19m spoken Shawn Wang · 4m spoken Alessio Fanelli · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Stanford professor and Inception Labs CEO Stefano Ermon discusses the breakthrough mechanics, ultra-fast inference capabilities, and systems architecture of discrete diffusion language models on the Latent Space Lightning podcast.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.7% of the talking time here. How this is scored →

The hosts as informed peer 5.3 Guest teaching 5.8 Guest disagreement 1.3 The hosts pushing back 1.7
05100:0010:0020:004:04–6:47 · The hosts as informed peer 4/10 Mechanics of Discrete Diffusion versus Autoregressive Text Generation Alessio asks foundational questions about discrete diffusion mechanics and dataset generation compared to standard autoregressive LLMs. Stefano provides a clear tutorial on parallel token modification, non-Gaussian noise models, and token infilling vs. flipping.6:47–9:28 · The hosts as informed peer 6/10 Model Architecture, Attention Masks, and Guidance in Diffusion Swyx probes whether existing pretrained causal backbones can be adapted for diffusion or if models must be trained from scratch. Stefano explains why non-causal bidirectional attention makes causal weights difficult to adapt despite architectural similarities.9:28–12:35 · The hosts as informed peer 5/10 Theoretical Breakthroughs and Comparison with Google's Gemini Diffusion Swyx asks about theoretical breakthroughs and Google's Gemini Diffusion, assuming it is inaccessible for testing. Stefano educates Swyx on discrete score matching mathematics and notes the availability of a public Gemini Diffusion playground.12:35–17:24 · The hosts as informed peer 5/10 Post-Training Alignment and Diffusion Preference Optimization Alessio asks about post-training alignment and compute trade-offs. Stefano details how his original work co-authoring DPO was adapted to diffusion text models and explains how diffusion Pareto dominates autoregressive models across latency and throughput.17:25–20:25 · The hosts as informed peer 5/10 Bidirectional Reasoning, Error Correction, and Architecture Scaling Alessio queries whether diffusion architectures face inherent capability limitations compared to autoregressive reasoning models. Stefano highlights diffusion's built-in bidirectional error correction as superior to sequential token chains.20:26–23:15 · The hosts as informed peer 7/10 Proprietary Serving Engines and Systems Engineering Hiring Swyx demonstrates strong systems awareness regarding serving engines, continuous batching, and KV-cache alternatives. Stefano discusses why custom proprietary serving infrastructure is required to serve discrete diffusion in production.4:04–6:47 · Guest teaching 6/10 Mechanics of Discrete Diffusion versus Autoregressive Text Generation Alessio asks foundational questions about discrete diffusion mechanics and dataset generation compared to standard autoregressive LLMs. Stefano provides a clear tutorial on parallel token modification, non-Gaussian noise models, and token infilling vs. flipping.6:47–9:28 · Guest teaching 5/10 Model Architecture, Attention Masks, and Guidance in Diffusion Swyx probes whether existing pretrained causal backbones can be adapted for diffusion or if models must be trained from scratch. Stefano explains why non-causal bidirectional attention makes causal weights difficult to adapt despite architectural similarities.9:28–12:35 · Guest teaching 7/10 Theoretical Breakthroughs and Comparison with Google's Gemini Diffusion Swyx asks about theoretical breakthroughs and Google's Gemini Diffusion, assuming it is inaccessible for testing. Stefano educates Swyx on discrete score matching mathematics and notes the availability of a public Gemini Diffusion playground.12:35–17:24 · Guest teaching 6/10 Post-Training Alignment and Diffusion Preference Optimization Alessio asks about post-training alignment and compute trade-offs. Stefano details how his original work co-authoring DPO was adapted to diffusion text models and explains how diffusion Pareto dominates autoregressive models across latency and throughput.17:25–20:25 · Guest teaching 6/10 Bidirectional Reasoning, Error Correction, and Architecture Scaling Alessio queries whether diffusion architectures face inherent capability limitations compared to autoregressive reasoning models. Stefano highlights diffusion's built-in bidirectional error correction as superior to sequential token chains.20:26–23:15 · Guest teaching 5/10 Proprietary Serving Engines and Systems Engineering Hiring Swyx demonstrates strong systems awareness regarding serving engines, continuous batching, and KV-cache alternatives. Stefano discusses why custom proprietary serving infrastructure is required to serve discrete diffusion in production.4:04–6:47 · Guest disagreement 1/10 Mechanics of Discrete Diffusion versus Autoregressive Text Generation Alessio asks foundational questions about discrete diffusion mechanics and dataset generation compared to standard autoregressive LLMs. Stefano provides a clear tutorial on parallel token modification, non-Gaussian noise models, and token infilling vs. flipping.6:47–9:28 · Guest disagreement 2/10 Model Architecture, Attention Masks, and Guidance in Diffusion Swyx probes whether existing pretrained causal backbones can be adapted for diffusion or if models must be trained from scratch. Stefano explains why non-causal bidirectional attention makes causal weights difficult to adapt despite architectural similarities.9:28–12:35 · Guest disagreement 1/10 Theoretical Breakthroughs and Comparison with Google's Gemini Diffusion Swyx asks about theoretical breakthroughs and Google's Gemini Diffusion, assuming it is inaccessible for testing. Stefano educates Swyx on discrete score matching mathematics and notes the availability of a public Gemini Diffusion playground.12:35–17:24 · Guest disagreement 1/10 Post-Training Alignment and Diffusion Preference Optimization Alessio asks about post-training alignment and compute trade-offs. Stefano details how his original work co-authoring DPO was adapted to diffusion text models and explains how diffusion Pareto dominates autoregressive models across latency and throughput.17:25–20:25 · Guest disagreement 2/10 Bidirectional Reasoning, Error Correction, and Architecture Scaling Alessio queries whether diffusion architectures face inherent capability limitations compared to autoregressive reasoning models. Stefano highlights diffusion's built-in bidirectional error correction as superior to sequential token chains.20:26–23:15 · Guest disagreement 1/10 Proprietary Serving Engines and Systems Engineering Hiring Swyx demonstrates strong systems awareness regarding serving engines, continuous batching, and KV-cache alternatives. Stefano discusses why custom proprietary serving infrastructure is required to serve discrete diffusion in production.4:04–6:47 · The hosts pushing back 1/10 Mechanics of Discrete Diffusion versus Autoregressive Text Generation Alessio asks foundational questions about discrete diffusion mechanics and dataset generation compared to standard autoregressive LLMs. Stefano provides a clear tutorial on parallel token modification, non-Gaussian noise models, and token infilling vs. flipping.6:47–9:28 · The hosts pushing back 2/10 Model Architecture, Attention Masks, and Guidance in Diffusion Swyx probes whether existing pretrained causal backbones can be adapted for diffusion or if models must be trained from scratch. Stefano explains why non-causal bidirectional attention makes causal weights difficult to adapt despite architectural similarities.9:28–12:35 · The hosts pushing back 1/10 Theoretical Breakthroughs and Comparison with Google's Gemini Diffusion Swyx asks about theoretical breakthroughs and Google's Gemini Diffusion, assuming it is inaccessible for testing. Stefano educates Swyx on discrete score matching mathematics and notes the availability of a public Gemini Diffusion playground.12:35–17:24 · The hosts pushing back 2/10 Post-Training Alignment and Diffusion Preference Optimization Alessio asks about post-training alignment and compute trade-offs. Stefano details how his original work co-authoring DPO was adapted to diffusion text models and explains how diffusion Pareto dominates autoregressive models across latency and throughput.17:25–20:25 · The hosts pushing back 2/10 Bidirectional Reasoning, Error Correction, and Architecture Scaling Alessio queries whether diffusion architectures face inherent capability limitations compared to autoregressive reasoning models. Stefano highlights diffusion's built-in bidirectional error correction as superior to sequential token chains.20:26–23:15 · The hosts pushing back 2/10 Proprietary Serving Engines and Systems Engineering Hiring Swyx demonstrates strong systems awareness regarding serving engines, continuous batching, and KV-cache alternatives. Stefano discusses why custom proprietary serving infrastructure is required to serve discrete diffusion in production.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 22.8% · guest 77.2%0:00 · the hosts 22.8% · guest 77.2%3:00 · the hosts 14.7% · guest 85.3%3:00 · the hosts 14.7% · guest 85.3%6:00 · the hosts 27.9% · guest 72.1%6:00 · the hosts 27.9% · guest 72.1%9:00 · the hosts 24.8% · guest 75.2%9:00 · the hosts 24.8% · guest 75.2%12:00 · the hosts 16.5% · guest 83.5%12:00 · the hosts 16.5% · guest 83.5%15:00 · the hosts 25.4% · guest 74.6%15:00 · the hosts 25.4% · guest 74.6%18:00 · the hosts 18.7% · guest 81.3%18:00 · the hosts 18.7% · guest 81.3%21:00 · the hosts 19.5% · guest 80.5%21:00 · the hosts 19.5% · guest 80.5%24:00 · the hosts 45.4% · guest 54.6%24:00 · the hosts 45.4% · guest 54.6%27:00 · the hosts 16.2% · guest 83.8%27:00 · the hosts 16.2% · guest 83.8%
Sharpest disagreement ▶ 8:48 Correcting assumption about reusing open model backbones

When Swyx assumes nothing can be reused from existing open ecosystems, Stefano directly corrects the premise by explaining how architectures and data pipelines carry over.

Hardest push from the hosts ▶ 15:19 Challenging the latency vs quality trade-off

Alessio pushes back directly to clarify whether diffusion's primary value is merely latency optimization or whether the underlying model quality is fundamentally superior.

Biggest teaching moment ▶ 11:43 Gemini Diffusion playground availability

When Swyx assumes Google's Gemini Diffusion is unreleased and asks how one could test it, Stefano politely informs him that a public playground is already live.

The host holds their own ▶ 20:25 Detailed serving infrastructure probing

Swyx demonstrates deep domain knowledge by pressing on batching dynamics and systems serving constraints specific to non-autoregressive language models.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Mechanics of Discrete Diffusion versus Autoregressive Text Generation 4611 Alessio asks foundational questions about discrete diffusion mechanics and dataset generation compared to standard autoregressive LLMs. Stefano provides a clear tutorial on parallel token modification, non-Gaussian noise models, and token infilling vs. flipping.
Model Architecture, Attention Masks, and Guidance in Diffusion 6522 Swyx probes whether existing pretrained causal backbones can be adapted for diffusion or if models must be trained from scratch. Stefano explains why non-causal bidirectional attention makes causal weights difficult to adapt despite architectural similarities.
Theoretical Breakthroughs and Comparison with Google's Gemini Diffusion 5711 Swyx asks about theoretical breakthroughs and Google's Gemini Diffusion, assuming it is inaccessible for testing. Stefano educates Swyx on discrete score matching mathematics and notes the availability of a public Gemini Diffusion playground.
Post-Training Alignment and Diffusion Preference Optimization 5612 Alessio asks about post-training alignment and compute trade-offs. Stefano details how his original work co-authoring DPO was adapted to diffusion text models and explains how diffusion Pareto dominates autoregressive models across latency and throughput.
Bidirectional Reasoning, Error Correction, and Architecture Scaling 5622 Alessio queries whether diffusion architectures face inherent capability limitations compared to autoregressive reasoning models. Stefano highlights diffusion's built-in bidirectional error correction as superior to sequential token chains.
Proprietary Serving Engines and Systems Engineering Hiring 7512 Swyx demonstrates strong systems awareness regarding serving engines, continuous batching, and KV-cache alternatives. Stefano discusses why custom proprietary serving infrastructure is required to serve discrete diffusion in production.

Statements from this episode (14)

Assertion Supported
Ermon: Inception Labs trained the first commercial-scale diffusion LLMs
“We've been successful in training the first commercial scale diffusion language models. We call this model Mercury.”
Stefano Ermon Aug 4, 2025 ▶ 3:26
Insight
Ermon: Diffusion LLMs gain speed by modifying multiple tokens in parallel
“That's kind of like the reason diffusion, diffusion language models are much faster compared to autoregressive models. Is that each neural network evaluation doesn't just give you one token, like in the typical autoregressive world, but it's able to Output, es…”
Stefano Ermon Aug 4, 2025 ▶ 5:05
Insight
Ermon: Adapting Pretrained Causal LLMs to Diffusion Models Is Difficult
“The challenge is that, yeah, the training objective is quite different because you are training based on denoising as opposed to next token prediction. Diffusion models are not causal and that is also kind of problematic. I mean, it's a big advantage of diffus…”
Stefano Ermon Aug 4, 2025 ▶ 7:29
Insight
Ermon: Diffusion LLMs Can Reuse Standard Architectures and Datasets
“Well, you can use architectures. I think that at least the shapes you, that, that can be leveraged. So you don't have to reinvent and necessarily completely different neural network architectures can, a lot of the data can be used. Like I think perhaps there a…”
Stefano Ermon Aug 4, 2025 ▶ 8:49
Insight
Ermon: Diffusion models excel at infilling due to bidirectional context
“Diffusion models. Not surprisingly, they work pretty well at the infilling where you really need to be able to use context to the left and to the right.”
Stefano Ermon Aug 4, 2025 ▶ 12:01
Assertion Not checkable as stated
Ermon: Google's Gemini Diffusion benchmark numbers match early Mercury Coder results
“They've released some benchmark numbers. They seem to be pretty close to the numbers that we were getting with the Mercury Coder back in some, you know, back in early this year.”
Stefano Ermon Aug 4, 2025 ▶ 12:20
Disclosure
Inception Labs develops specialized DPO algorithm for diffusion language models
“We have a DPO algorithm specialized for diffusion language models.”
Stefano Ermon Aug 4, 2025 ▶ 13:05
Assertion Supported
Ermon: Diffusion LLMs Pareto-dominate autoregressive models on inference efficiency
“On the inference side, what we're seeing is that diffusion models are much more efficient. We're actually able to Pareto dominate autoregressive models. If you think about the typical trade-off between throughput versus latency, which you kind of like cannot, …”
Stefano Ermon Aug 4, 2025 ▶ 14:30
Assertion Partly supported
Inception generalist model matches Claude Haiku quality at 5-10x speed
“We had our generalist model evaluated by artificial analysis and the intelligence score from AA artificial analysis around 40. So it's comparable to GPT, 4.1 nano, cloud haiku, kind of like Close source speed optimized models. It's roughly comparable in terms …”
Stefano Ermon Aug 4, 2025 ▶ 16:55
Prediction Not checkable as stated
Ermon: Diffusion models could become the dominant architecture over autoregressive models
“I'm pretty optimistic about a future where diffusion models Can become the dominant solution. I've seen it happen before with GANs a few years ago, so I wouldn't be surprised if that's the case also here.”
Stefano Ermon Aug 4, 2025 ▶ 18:35
Insight
Ermon: Diffusion models naturally self-correct errors during generation unlike autoregressive LLMs
“The fact that you have error correction that is built in. So you think about an autoregressive model. Once you output something, you can never take it back. And so if you want to do, you know, if you want to fix mistakes, maybe you can do a reasoning chain. Ma…”
Stefano Ermon Aug 4, 2025 ▶ 18:52
Disclosure
Ermon: Inception Labs has no plans to open-source models
“So we don't have a plan at the moment to release models or to open source any model.”
Stefano Ermon Aug 4, 2025 ▶ 20:51
Assertion Not checkable as stated
Ermon: Inception Labs built proprietary engine for production inference traffic
“So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine. And so we are supporting production traffic already with our own inference engine. We support continuous batch and quantization.”
Stefano Ermon Aug 4, 2025 ▶ 21:33
Prediction Not checkable as stated
Ermon: Power constraints will drive diffusion models to replace frontier LLMs
“If it happens, it's gonna be driven by efficiency. Like we're all constrained by essentially power. And if you have, I mean, at the end of the day, it's all an inference game, right? Okay. Training is expensive, but then the thing that matters is being able to…”
Stefano Ermon Aug 4, 2025 ▶ 23:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.