Jan 4, 2025 · 52m · latent-space

AI Engineering for Art - with comfyanonymous

comfyanonymous (Comfy) · 37m spoken Shawn Wang · 5m spoken Alessio Fanelli · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Podcast, hosts Alessio Fanelli and Swix interview Comfyanonymous, the creator of ComfyUI, exploring the technical architecture, memory management optimizations, and open-source foundation model dynamics powering modern generative workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 18.7% of the talking time here. How this is scored →

The hosts as informed peer 4.6 Guest teaching 5.8 Guest disagreement 1.6 The hosts pushing back 1.2
05100:0015:0030:0045:000:49–5:37 · The hosts as informed peer 4/10 The Origin and Motivation Behind ComfyUI Swyx probes Comfy's technical background, asking if he previously worked on distributed systems or GPUs. Comfy explains he was just a regular web developer who got hooked on image generation and hacked Automatic1111's high-res fix to allow multiple passes and models.5:37–8:25 · The hosts as informed peer 5/10 Area Conditioning and Latent Space Composition Swyx asks if area conditioning was mainly designed for fixing hands, but Comfy corrects him, explaining it was developed for spatial compositing across prompts before MultiDiffusion formalised it. Alessio inquires whether cross-model composition is possible, and Comfy details the constraint of needing a shared latent space.8:25–14:12 · The hosts as informed peer 6/10 Working at Stability AI and SDXL Architecture Swyx demonstrates familiarity with the Stability AI research staff and pipeline architectures. Comfy shares insider context about why Stability hired him to pipeline the base and refiner models for SDXL, and how internal red-teaming delays derailed Stable Cascade's release momentum.14:12–16:55 · The hosts as informed peer 4/10 Model Evaluation Methods and Aesthetic Subjectivity Swyx suggests the community rapidly abandons models without thorough evaluation. Comfy pushes back, clarifying that users only upgrade when an obvious leap in capability occurs, and highlights that formal benchmark evaluations matter less than subjective artistic taste and aesthetic vibe checks.16:55–20:16 · The hosts as informed peer 5/10 Fine-Tuning Techniques: Textual Inversions and Encoders Swyx connects textual inversion to representation engineering, prompting Comfy to explain the technical mechanics of training pseudo-word vectors into the text encoder. Comfy details how textual inversions behave differently across SD 1.5, SDXL, and SD3 due to varying numbers of text encoders.20:16–23:47 · The hosts as informed peer 5/10 Text Conditioning: Token Chunking, Context, and Weighting Comfy explains the engineering workarounds used for text prompts exceeding CLIP's 77-token ceiling by breaking text into chunks and concatenating outputs. He also breaks down prompt weighting via vector interpolation and explains why this technique fails on deeper encoders like T5.23:47–25:57 · The hosts as informed peer 4/10 Mechanics of Low-Rank Adaptation (LoRA, LoCon, LoHa) Alessio and Swyx ask about low-rank adaptation methods, noting they have not heard of variants like LoCon and LoHa. Comfy educates them on low-rank matrix decomposition and how different algorithms represent weight deltas during inference.25:57–31:00 · The hosts as informed peer 5/10 UI Architecture: Node Graphs vs. Monolithic Web Frameworks Swyx suggests standard Python UI frameworks like Gradio or Streamlit could have served the project. Comfy strongly criticizes Gradio for coupling UI state with backend logic, arguing that production-grade architecture requires a clean separation between backend execution and frontend graphs.31:00–35:08 · The hosts as informed peer 4/10 GPU Memory Management and Hardware Optimization Comfy details ComfyUI's VRAM management, explaining how proactive model unloading avoids Windows GPU driver fallback paging to system RAM. He also describes the state of AMD ROCm support on Windows versus Linux.35:08–38:34 · The hosts as informed peer 5/10 Node Granularity and Core Sampling Parameters Alessio asks how users learn sampler parameters like CFG, step count, and schedulers. Comfy explains the math behind classifier-free guidance as positive minus negative prompt vectors, while advising that intuitive visual experimentation is more effective than theoretical reading.38:34–41:36 · The hosts as informed peer 4/10 Ecosystem Innovations: Plugins, Custom Nodes, and Node Registry Alessio and Comfy discuss extreme community workflows, such as dynamic game texture streaming and YouTube download pipelines. Comfy notes that custom nodes were made intentionally simple to develop, leading to widespread adoption alongside node duplication challenges.41:36–44:30 · The hosts as informed peer 5/10 Evolution of Open Video Generation: SVD to Mochi Comfy clarifies what constitutes true video generation versus pseudo-video, contrasting 2D spatial models like Stable Video Diffusion and AnimateDiff with true 3D spatio-temporal VAE architectures like Mochi.44:30–48:07 · The hosts as informed peer 4/10 The SDXL Release and the Inflection Point of ComfyUI Comfy explains the viral tipping point of ComfyUI: when SDXL 0.9 leaked before general release, ComfyUI was the only efficient implementation capable of running the dual-model pipeline on consumer VRAM.0:49–5:37 · Guest teaching 5/10 The Origin and Motivation Behind ComfyUI Swyx probes Comfy's technical background, asking if he previously worked on distributed systems or GPUs. Comfy explains he was just a regular web developer who got hooked on image generation and hacked Automatic1111's high-res fix to allow multiple passes and models.5:37–8:25 · Guest teaching 5/10 Area Conditioning and Latent Space Composition Swyx asks if area conditioning was mainly designed for fixing hands, but Comfy corrects him, explaining it was developed for spatial compositing across prompts before MultiDiffusion formalised it. Alessio inquires whether cross-model composition is possible, and Comfy details the constraint of needing a shared latent space.8:25–14:12 · Guest teaching 5/10 Working at Stability AI and SDXL Architecture Swyx demonstrates familiarity with the Stability AI research staff and pipeline architectures. Comfy shares insider context about why Stability hired him to pipeline the base and refiner models for SDXL, and how internal red-teaming delays derailed Stable Cascade's release momentum.14:12–16:55 · Guest teaching 6/10 Model Evaluation Methods and Aesthetic Subjectivity Swyx suggests the community rapidly abandons models without thorough evaluation. Comfy pushes back, clarifying that users only upgrade when an obvious leap in capability occurs, and highlights that formal benchmark evaluations matter less than subjective artistic taste and aesthetic vibe checks.16:55–20:16 · Guest teaching 6/10 Fine-Tuning Techniques: Textual Inversions and Encoders Swyx connects textual inversion to representation engineering, prompting Comfy to explain the technical mechanics of training pseudo-word vectors into the text encoder. Comfy details how textual inversions behave differently across SD 1.5, SDXL, and SD3 due to varying numbers of text encoders.20:16–23:47 · Guest teaching 7/10 Text Conditioning: Token Chunking, Context, and Weighting Comfy explains the engineering workarounds used for text prompts exceeding CLIP's 77-token ceiling by breaking text into chunks and concatenating outputs. He also breaks down prompt weighting via vector interpolation and explains why this technique fails on deeper encoders like T5.23:47–25:57 · Guest teaching 6/10 Mechanics of Low-Rank Adaptation (LoRA, LoCon, LoHa) Alessio and Swyx ask about low-rank adaptation methods, noting they have not heard of variants like LoCon and LoHa. Comfy educates them on low-rank matrix decomposition and how different algorithms represent weight deltas during inference.25:57–31:00 · Guest teaching 6/10 UI Architecture: Node Graphs vs. Monolithic Web Frameworks Swyx suggests standard Python UI frameworks like Gradio or Streamlit could have served the project. Comfy strongly criticizes Gradio for coupling UI state with backend logic, arguing that production-grade architecture requires a clean separation between backend execution and frontend graphs.31:00–35:08 · Guest teaching 7/10 GPU Memory Management and Hardware Optimization Comfy details ComfyUI's VRAM management, explaining how proactive model unloading avoids Windows GPU driver fallback paging to system RAM. He also describes the state of AMD ROCm support on Windows versus Linux.35:08–38:34 · Guest teaching 6/10 Node Granularity and Core Sampling Parameters Alessio asks how users learn sampler parameters like CFG, step count, and schedulers. Comfy explains the math behind classifier-free guidance as positive minus negative prompt vectors, while advising that intuitive visual experimentation is more effective than theoretical reading.38:34–41:36 · Guest teaching 5/10 Ecosystem Innovations: Plugins, Custom Nodes, and Node Registry Alessio and Comfy discuss extreme community workflows, such as dynamic game texture streaming and YouTube download pipelines. Comfy notes that custom nodes were made intentionally simple to develop, leading to widespread adoption alongside node duplication challenges.41:36–44:30 · Guest teaching 6/10 Evolution of Open Video Generation: SVD to Mochi Comfy clarifies what constitutes true video generation versus pseudo-video, contrasting 2D spatial models like Stable Video Diffusion and AnimateDiff with true 3D spatio-temporal VAE architectures like Mochi.44:30–48:07 · Guest teaching 6/10 The SDXL Release and the Inflection Point of ComfyUI Comfy explains the viral tipping point of ComfyUI: when SDXL 0.9 leaked before general release, ComfyUI was the only efficient implementation capable of running the dual-model pipeline on consumer VRAM.0:49–5:37 · Guest disagreement 1/10 The Origin and Motivation Behind ComfyUI Swyx probes Comfy's technical background, asking if he previously worked on distributed systems or GPUs. Comfy explains he was just a regular web developer who got hooked on image generation and hacked Automatic1111's high-res fix to allow multiple passes and models.5:37–8:25 · Guest disagreement 2/10 Area Conditioning and Latent Space Composition Swyx asks if area conditioning was mainly designed for fixing hands, but Comfy corrects him, explaining it was developed for spatial compositing across prompts before MultiDiffusion formalised it. Alessio inquires whether cross-model composition is possible, and Comfy details the constraint of needing a shared latent space.8:25–14:12 · Guest disagreement 2/10 Working at Stability AI and SDXL Architecture Swyx demonstrates familiarity with the Stability AI research staff and pipeline architectures. Comfy shares insider context about why Stability hired him to pipeline the base and refiner models for SDXL, and how internal red-teaming delays derailed Stable Cascade's release momentum.14:12–16:55 · Guest disagreement 3/10 Model Evaluation Methods and Aesthetic Subjectivity Swyx suggests the community rapidly abandons models without thorough evaluation. Comfy pushes back, clarifying that users only upgrade when an obvious leap in capability occurs, and highlights that formal benchmark evaluations matter less than subjective artistic taste and aesthetic vibe checks.16:55–20:16 · Guest disagreement 1/10 Fine-Tuning Techniques: Textual Inversions and Encoders Swyx connects textual inversion to representation engineering, prompting Comfy to explain the technical mechanics of training pseudo-word vectors into the text encoder. Comfy details how textual inversions behave differently across SD 1.5, SDXL, and SD3 due to varying numbers of text encoders.20:16–23:47 · Guest disagreement 1/10 Text Conditioning: Token Chunking, Context, and Weighting Comfy explains the engineering workarounds used for text prompts exceeding CLIP's 77-token ceiling by breaking text into chunks and concatenating outputs. He also breaks down prompt weighting via vector interpolation and explains why this technique fails on deeper encoders like T5.23:47–25:57 · Guest disagreement 1/10 Mechanics of Low-Rank Adaptation (LoRA, LoCon, LoHa) Alessio and Swyx ask about low-rank adaptation methods, noting they have not heard of variants like LoCon and LoHa. Comfy educates them on low-rank matrix decomposition and how different algorithms represent weight deltas during inference.25:57–31:00 · Guest disagreement 4/10 UI Architecture: Node Graphs vs. Monolithic Web Frameworks Swyx suggests standard Python UI frameworks like Gradio or Streamlit could have served the project. Comfy strongly criticizes Gradio for coupling UI state with backend logic, arguing that production-grade architecture requires a clean separation between backend execution and frontend graphs.31:00–35:08 · Guest disagreement 1/10 GPU Memory Management and Hardware Optimization Comfy details ComfyUI's VRAM management, explaining how proactive model unloading avoids Windows GPU driver fallback paging to system RAM. He also describes the state of AMD ROCm support on Windows versus Linux.35:08–38:34 · Guest disagreement 1/10 Node Granularity and Core Sampling Parameters Alessio asks how users learn sampler parameters like CFG, step count, and schedulers. Comfy explains the math behind classifier-free guidance as positive minus negative prompt vectors, while advising that intuitive visual experimentation is more effective than theoretical reading.38:34–41:36 · Guest disagreement 1/10 Ecosystem Innovations: Plugins, Custom Nodes, and Node Registry Alessio and Comfy discuss extreme community workflows, such as dynamic game texture streaming and YouTube download pipelines. Comfy notes that custom nodes were made intentionally simple to develop, leading to widespread adoption alongside node duplication challenges.41:36–44:30 · Guest disagreement 2/10 Evolution of Open Video Generation: SVD to Mochi Comfy clarifies what constitutes true video generation versus pseudo-video, contrasting 2D spatial models like Stable Video Diffusion and AnimateDiff with true 3D spatio-temporal VAE architectures like Mochi.44:30–48:07 · Guest disagreement 1/10 The SDXL Release and the Inflection Point of ComfyUI Comfy explains the viral tipping point of ComfyUI: when SDXL 0.9 leaked before general release, ComfyUI was the only efficient implementation capable of running the dual-model pipeline on consumer VRAM.0:49–5:37 · The hosts pushing back 1/10 The Origin and Motivation Behind ComfyUI Swyx probes Comfy's technical background, asking if he previously worked on distributed systems or GPUs. Comfy explains he was just a regular web developer who got hooked on image generation and hacked Automatic1111's high-res fix to allow multiple passes and models.5:37–8:25 · The hosts pushing back 1/10 Area Conditioning and Latent Space Composition Swyx asks if area conditioning was mainly designed for fixing hands, but Comfy corrects him, explaining it was developed for spatial compositing across prompts before MultiDiffusion formalised it. Alessio inquires whether cross-model composition is possible, and Comfy details the constraint of needing a shared latent space.8:25–14:12 · The hosts pushing back 2/10 Working at Stability AI and SDXL Architecture Swyx demonstrates familiarity with the Stability AI research staff and pipeline architectures. Comfy shares insider context about why Stability hired him to pipeline the base and refiner models for SDXL, and how internal red-teaming delays derailed Stable Cascade's release momentum.14:12–16:55 · The hosts pushing back 2/10 Model Evaluation Methods and Aesthetic Subjectivity Swyx suggests the community rapidly abandons models without thorough evaluation. Comfy pushes back, clarifying that users only upgrade when an obvious leap in capability occurs, and highlights that formal benchmark evaluations matter less than subjective artistic taste and aesthetic vibe checks.16:55–20:16 · The hosts pushing back 1/10 Fine-Tuning Techniques: Textual Inversions and Encoders Swyx connects textual inversion to representation engineering, prompting Comfy to explain the technical mechanics of training pseudo-word vectors into the text encoder. Comfy details how textual inversions behave differently across SD 1.5, SDXL, and SD3 due to varying numbers of text encoders.20:16–23:47 · The hosts pushing back 1/10 Text Conditioning: Token Chunking, Context, and Weighting Comfy explains the engineering workarounds used for text prompts exceeding CLIP's 77-token ceiling by breaking text into chunks and concatenating outputs. He also breaks down prompt weighting via vector interpolation and explains why this technique fails on deeper encoders like T5.23:47–25:57 · The hosts pushing back 0/10 Mechanics of Low-Rank Adaptation (LoRA, LoCon, LoHa) Alessio and Swyx ask about low-rank adaptation methods, noting they have not heard of variants like LoCon and LoHa. Comfy educates them on low-rank matrix decomposition and how different algorithms represent weight deltas during inference.25:57–31:00 · The hosts pushing back 3/10 UI Architecture: Node Graphs vs. Monolithic Web Frameworks Swyx suggests standard Python UI frameworks like Gradio or Streamlit could have served the project. Comfy strongly criticizes Gradio for coupling UI state with backend logic, arguing that production-grade architecture requires a clean separation between backend execution and frontend graphs.31:00–35:08 · The hosts pushing back 1/10 GPU Memory Management and Hardware Optimization Comfy details ComfyUI's VRAM management, explaining how proactive model unloading avoids Windows GPU driver fallback paging to system RAM. He also describes the state of AMD ROCm support on Windows versus Linux.35:08–38:34 · The hosts pushing back 1/10 Node Granularity and Core Sampling Parameters Alessio asks how users learn sampler parameters like CFG, step count, and schedulers. Comfy explains the math behind classifier-free guidance as positive minus negative prompt vectors, while advising that intuitive visual experimentation is more effective than theoretical reading.38:34–41:36 · The hosts pushing back 1/10 Ecosystem Innovations: Plugins, Custom Nodes, and Node Registry Alessio and Comfy discuss extreme community workflows, such as dynamic game texture streaming and YouTube download pipelines. Comfy notes that custom nodes were made intentionally simple to develop, leading to widespread adoption alongside node duplication challenges.41:36–44:30 · The hosts pushing back 1/10 Evolution of Open Video Generation: SVD to Mochi Comfy clarifies what constitutes true video generation versus pseudo-video, contrasting 2D spatial models like Stable Video Diffusion and AnimateDiff with true 3D spatio-temporal VAE architectures like Mochi.44:30–48:07 · The hosts pushing back 1/10 The SDXL Release and the Inflection Point of ComfyUI Comfy explains the viral tipping point of ComfyUI: when SDXL 0.9 leaked before general release, ComfyUI was the only efficient implementation capable of running the dual-model pipeline on consumer VRAM.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 41% · guest 59%0:00 · the hosts 41% · guest 59%3:00 · the hosts 10.9% · guest 89.1%3:00 · the hosts 10.9% · guest 89.1%6:00 · the hosts 17.3% · guest 82.7%6:00 · the hosts 17.3% · guest 82.7%9:00 · the hosts 21.1% · guest 78.9%9:00 · the hosts 21.1% · guest 78.9%12:00 · the hosts 35% · guest 65%12:00 · the hosts 35% · guest 65%15:00 · the hosts 21.2% · guest 78.8%15:00 · the hosts 21.2% · guest 78.8%18:00 · the hosts 9.8% · guest 90.2%18:00 · the hosts 9.8% · guest 90.2%21:00 · the hosts 17% · guest 83%21:00 · the hosts 17% · guest 83%24:00 · the hosts 20.3% · guest 79.7%24:00 · the hosts 20.3% · guest 79.7%27:00 · the hosts 21.5% · guest 78.5%27:00 · the hosts 21.5% · guest 78.5%30:00 · the hosts 10.4% · guest 89.6%30:00 · the hosts 10.4% · guest 89.6%33:00 · the hosts 30.1% · guest 69.9%33:00 · the hosts 30.1% · guest 69.9%36:00 · the hosts 17.9% · guest 82.1%36:00 · the hosts 17.9% · guest 82.1%39:00 · the hosts 15.6% · guest 84.4%39:00 · the hosts 15.6% · guest 84.4%42:00 · the hosts 21.1% · guest 78.9%42:00 · the hosts 21.1% · guest 78.9%45:00 · the hosts 0.5% · guest 99.5%45:00 · the hosts 0.5% · guest 99.5%48:00 · the hosts 9.3% · guest 90.7%48:00 · the hosts 9.3% · guest 90.7%51:00 · the hosts 17.5% · guest 82.5%51:00 · the hosts 17.5% · guest 82.5%
Sharpest disagreement ▶ 28:55 Rejecting Gradio as Bad Software Architecture

Comfy forcefully rejects Gradio, stating it creates messy codebases by coupling frontend and backend logic, dismissing it as unsuitable for maintainable long-term software.

Hardest push from the hosts ▶ 14:11 Pushing Back on Community Churn Assumption

Comfy refutes Swyx's claim that the open-source community impulsively jumps between models, asserting that users only migrate when substantial capability leaps are demonstrated.

Biggest teaching moment ▶ 20:55 Educating on Token Chunking and Text Interpolation

Comfy explains the hidden technical hacks used to overcome CLIP's 77-token limit and details why prompt weighting degrades on modern deep text encoders like T5.

The host holds their own ▶ 12:27 Host Connecting Cascade Architecture and Researcher Pedigree

Swyx demonstrates insider domain expertise by recognizing the Würstchen research team behind Stable Cascade and discussing their migration from Stability AI.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Origin and Motivation Behind ComfyUI 4511 Swyx probes Comfy's technical background, asking if he previously worked on distributed systems or GPUs. Comfy explains he was just a regular web developer who got hooked on image generation and hacked Automatic1111's high-res fix to allow multiple passes and models.
Area Conditioning and Latent Space Composition 5521 Swyx asks if area conditioning was mainly designed for fixing hands, but Comfy corrects him, explaining it was developed for spatial compositing across prompts before MultiDiffusion formalised it. Alessio inquires whether cross-model composition is possible, and Comfy details the constraint of needing a shared latent space.
Working at Stability AI and SDXL Architecture 6522 Swyx demonstrates familiarity with the Stability AI research staff and pipeline architectures. Comfy shares insider context about why Stability hired him to pipeline the base and refiner models for SDXL, and how internal red-teaming delays derailed Stable Cascade's release momentum.
Model Evaluation Methods and Aesthetic Subjectivity 4632 Swyx suggests the community rapidly abandons models without thorough evaluation. Comfy pushes back, clarifying that users only upgrade when an obvious leap in capability occurs, and highlights that formal benchmark evaluations matter less than subjective artistic taste and aesthetic vibe checks.
Fine-Tuning Techniques: Textual Inversions and Encoders 5611 Swyx connects textual inversion to representation engineering, prompting Comfy to explain the technical mechanics of training pseudo-word vectors into the text encoder. Comfy details how textual inversions behave differently across SD 1.5, SDXL, and SD3 due to varying numbers of text encoders.
Text Conditioning: Token Chunking, Context, and Weighting 5711 Comfy explains the engineering workarounds used for text prompts exceeding CLIP's 77-token ceiling by breaking text into chunks and concatenating outputs. He also breaks down prompt weighting via vector interpolation and explains why this technique fails on deeper encoders like T5.
Mechanics of Low-Rank Adaptation (LoRA, LoCon, LoHa) 4610 Alessio and Swyx ask about low-rank adaptation methods, noting they have not heard of variants like LoCon and LoHa. Comfy educates them on low-rank matrix decomposition and how different algorithms represent weight deltas during inference.
UI Architecture: Node Graphs vs. Monolithic Web Frameworks 5643 Swyx suggests standard Python UI frameworks like Gradio or Streamlit could have served the project. Comfy strongly criticizes Gradio for coupling UI state with backend logic, arguing that production-grade architecture requires a clean separation between backend execution and frontend graphs.
GPU Memory Management and Hardware Optimization 4711 Comfy details ComfyUI's VRAM management, explaining how proactive model unloading avoids Windows GPU driver fallback paging to system RAM. He also describes the state of AMD ROCm support on Windows versus Linux.
Node Granularity and Core Sampling Parameters 5611 Alessio asks how users learn sampler parameters like CFG, step count, and schedulers. Comfy explains the math behind classifier-free guidance as positive minus negative prompt vectors, while advising that intuitive visual experimentation is more effective than theoretical reading.
Ecosystem Innovations: Plugins, Custom Nodes, and Node Registry 4511 Alessio and Comfy discuss extreme community workflows, such as dynamic game texture streaming and YouTube download pipelines. Comfy notes that custom nodes were made intentionally simple to develop, leading to widespread adoption alongside node duplication challenges.
Evolution of Open Video Generation: SVD to Mochi 5621 Comfy clarifies what constitutes true video generation versus pseudo-video, contrasting 2D spatial models like Stable Video Diffusion and AnimateDiff with true 3D spatio-temporal VAE architectures like Mochi.
The SDXL Release and the Inflection Point of ComfyUI 4611 Comfy explains the viral tipping point of ComfyUI: when SDXL 0.9 leaked before general release, ComfyUI was the only efficient implementation capable of running the dual-model pipeline on consumer VRAM.

Statements from this episode (22)

Disclosure
Comfyanonymous had never written PyTorch before October 2022
“So basically October, 20, 22, just like I hadn't written a line of PyTorch before that. So it's completely new.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 2:45
What-if
Comfy: Extending Automatic1111 Was Harder Than Building ComfyUI From Scratch
“It would have been harder to implement that in the auto interface than to create my own interface, so that's when I decided to create my own.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 4:26
Assertion Supported
Comfyanonymous built and released ComfyUI in just fifteen days
“So I started writing the code January one, 20, 23. And then I released the first version on GitHub January, 1620, 23.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 4:58
Assertion Supported
MultiDiffusion researchers published area conditioning a month after ComfyUI
“And then a month later, there was a paper that came out called multi diffusion, which was the same thing, but yeah, that's .”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 7:11
Insight
Comfyanonymous: Latent space efficiency drove Stable Diffusion's popularity
“That's the problem that that's the reason why stable diffusion actually became like popular, like, cause was because of the latent space.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 8:06
Opinion
Comfy: Flux is better for consistency while SD 3.5 is better for creativity
“If you want some to make something more like creative, maybe as the 3.5, if you want to make something more consistent and flux is. Probably better.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 10:56
Assertion Not checkable as stated
Stability AI delayed Stable Cascade three months for red teaming
“And inside stability, actually that model was ready like three months before, but it got stuck in red teaming.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 12:28
What-if
Comfy: Stable Cascade would have been popular if SD3 hadn't stolen momentum
“If that model had released was Supposed to be released by the authors, then it would probably have gotten very popular since it's a step up from SDXL, but it got all of its momentum stolen by the SD three announcement.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 12:38
Assertion Supported
Comfy: Stable Cascade researchers left Stability AI right after its release
“They worked at stability for a bit and they left right after the Cascade release.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 13:31
Insight
Comfy: Training on all internet images yields ugly outputs
“You can create a very good model that doesn't generate nice images, because most images on the internet are ugly, so if you, if that's like, if you just, oh, I have the best model that can, like, it's super smart, I create it on all the, like, I turn it on jus…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 16:11
Disclosure
Comfyanonymous: Stability AI internally trained effective T5-XXL textual inversions
“Back as like when I was at stability, we actually did train internally some, Like textual versions on like T five XXL actually worked pretty well, but for some reason, yeah, people don't use them.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 18:49
Assertion Supported
Comfy: Prompt weighting does not work on deep text encoders like T5-XXL
“This stops working the deeper your text encoder is. So, on T-Five XSL, it doesn't work at all, so.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 23:06
Opinion
Comfyanonymous criticizes Gradio for coupling interface logic with backends
“Yeah, Gradio, I don't like Gradio. It's bad. Like, the, that's one of the reasons why, like, Automatic was very bad. It's great because the problem with Gradio, it forces you to, well, not forces you, but it kind of makes your interface logic and your back-end…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 28:56
Insight
Comfyanonymous: PyTorch Lacks Fine-Grained Memory Control for Complex Pipelines
“The problem with PyTorch is it's high levels. Don't have that much fine-grained control over, like specific memory stuff, so kind of have to leave, like, the memory freeing to Python and PyTorch, which is, can be annoying sometimes.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 33:13
Opinion
AMD GPUs perform horribly on Windows for generative AI tasks
“AMD works horribly on Windows. Like, on Linux, it works fine. It's lower than the price equivalent NVIDIA GPU. But it works, like, you can use it, generate images, everything works. On Linux, on Windows, You might have a hard time”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 34:09
Insight
Comfy: Lowering Custom Node Barriers Drove Ecosystem Growth but Created Duplication
“Something I think what I did is I made it easy to make custom notes. So I think that, that helped a lot for the ecosystem, because it is very easy to just make a note. So yeah, a bit too easy sometimes. Then we have the issue where there's a lot of Custom note…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 40:58
Insight
Comfyanonymous: True video models use 3D latents rather than 2D
“Why I say it's not a true video model is that you still have, like, the two D latents. Like, a true video model, like Mochi, for example, would have three D latents.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 42:22
Opinion
Comfyanonymous considers Genmo's Mochi the best open video generation model
“There's, yeah, there's actually a few of them, but the one I've implemented in Comfy is Mochi, because that, that seems to be The best one so far.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 43:05
Assertion Not checkable as stated
Automatic1111's inefficient SDXL implementation drove ComfyUI's viral user adoption
“The big, one point zero release happened, and wow, Confu UI was the only way a lot of people could actually run it on their computers, because it just, like, automatic was so, like, inefficient and bad that most people couldn't act, like, it just wouldn't work…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 47:23
Disclosure
Comfyanonymous remains the sole developer on the ComfyUI core backend
“So right now core, like on the core itself, it's me. But because the reason we're focused, like all the focus has been mostly on the front end right now, because that's the thing that's been neglected for. A long time.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 48:20
Disclosure
ComfyUI plans to monetize via cloud inference and enterprise features
“We're still gonna continue, like, doing the open source, like making MVUI the best way to run Like stable effusion models, like at least the open source side, and like, it's going to be best way to run or models locally, but we will have a few, like a few thin…”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 49:34
Insight
Comfyanonymous: Startups wrapping ComfyUI expand the ecosystem even without contributing
“As long as they use Comfy, it's I think it helps the ecosystem. So because more people, even if they don't, like, even if they don't contribute directly, The fact that they are using Comfy means that like people are more likely to like join the ecosystem.”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 50:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.