Sep 23, 2025 · 32m · latent-space

⚡️ Beyond Transformers with Power Retention

Diego Bachman · 25m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Podcast, Diego Bachman of Manifest AI discusses replacing transformer attention bottlenecks with Power Retention, an architecture utilizing fixed-size memory state. Diego demonstrates how custom CUDA kernel optimization via the Vidrio framework and model metamorphosis enable dramatic inference speedups, reduced memory overhead, and efficient long-context processing.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.4 Guest teaching 5.3 Guest disagreement 1.6 The hosts pushing back 0.5
05100:0010:0020:0030:000:55–5:27 · The hosts as informed peer 3/10 Transformer Bottlenecks and the Power Retention Architecture Alessio invites Diego to explain the founding thesis behind Manifest AI. Diego gives an in-depth critique of standard transformers, dismissing 256k-context claims as windowed Band-Aid solutions that degrade middle-context recall, and introduces the fixed-memory power retention architecture.5:28–8:14 · The hosts as informed peer 5/10 Inference Efficiency and Memory Allocation with Power Retention Alessio demonstrates familiarity with prefix caching economics in inference serving. Diego explains how dynamic KV cache sizing creates complex GPU scheduling problems, contrasting it with power retention's deterministic memory partitioning.8:14–14:27 · The hosts as informed peer 6/10 Vidrio Framework and Just-In-Time CUDA Kernel Optimization Alessio showcases technical understanding of low-level kernel optimization by noting Manifest's faster reimplementation of FlashAttention-2. Diego explains the mechanics of Vidrial, describing its empirical JIT parameter-sweeping approach across GPU configurations.14:27–19:22 · The hosts as informed peer 5/10 Model Metamorphosis: Converting StarCoder into PowerCoder Alessio prompts about the metamorphosis technique and mentions compute partners like SF Compute, but misidentifies StarCoder's base context window as 8k. Diego gently corrects it to 4k and shows training loss convergence curves in Log Cabin.19:22–22:46 · The hosts as informed peer 3/10 Live Demonstration of Manifesto Code Repair Tool Diego presents a live demonstration of the Manifesto code repair tool fixing broken classes inside a NanoGPT repository. Alessio provides supportive facilitation, prompting Diego to zoom in and commenting on training data overlap.22:46–26:04 · The hosts as informed peer 5/10 Expanding the Open-Source Ecosystem and Provider Partnerships Alessio asks whether inference providers like Together or Fireworks should convert every model to power retention and probes for potential downsides. Diego explains the strategic value for chain-of-thought inference workloads.26:04–28:38 · The hosts as informed peer 5/10 Scaling Power Retention and Overcoming Institutional Inertia Alessio asks how far the team is from training frontier-scale GPT-5 models using power retention. Diego draws historical parallels to early Transformer adoption, acknowledging the substantial institutional inertia at major labs.28:38–32:45 · The hosts as informed peer 3/10 Training Benchmarks, Hiring, and Long-Context Data Needs Diego checks back on the live benchmark showing a 5x training iteration speedup, makes recruiting pitches, and explains why standard document packing across short web texts undermines authentic long-context learning.0:55–5:27 · Guest teaching 7/10 Transformer Bottlenecks and the Power Retention Architecture Alessio invites Diego to explain the founding thesis behind Manifest AI. Diego gives an in-depth critique of standard transformers, dismissing 256k-context claims as windowed Band-Aid solutions that degrade middle-context recall, and introduces the fixed-memory power retention architecture.5:28–8:14 · Guest teaching 6/10 Inference Efficiency and Memory Allocation with Power Retention Alessio demonstrates familiarity with prefix caching economics in inference serving. Diego explains how dynamic KV cache sizing creates complex GPU scheduling problems, contrasting it with power retention's deterministic memory partitioning.8:14–14:27 · Guest teaching 6/10 Vidrio Framework and Just-In-Time CUDA Kernel Optimization Alessio showcases technical understanding of low-level kernel optimization by noting Manifest's faster reimplementation of FlashAttention-2. Diego explains the mechanics of Vidrial, describing its empirical JIT parameter-sweeping approach across GPU configurations.14:27–19:22 · Guest teaching 5/10 Model Metamorphosis: Converting StarCoder into PowerCoder Alessio prompts about the metamorphosis technique and mentions compute partners like SF Compute, but misidentifies StarCoder's base context window as 8k. Diego gently corrects it to 4k and shows training loss convergence curves in Log Cabin.19:22–22:46 · Guest teaching 3/10 Live Demonstration of Manifesto Code Repair Tool Diego presents a live demonstration of the Manifesto code repair tool fixing broken classes inside a NanoGPT repository. Alessio provides supportive facilitation, prompting Diego to zoom in and commenting on training data overlap.22:46–26:04 · Guest teaching 4/10 Expanding the Open-Source Ecosystem and Provider Partnerships Alessio asks whether inference providers like Together or Fireworks should convert every model to power retention and probes for potential downsides. Diego explains the strategic value for chain-of-thought inference workloads.26:04–28:38 · Guest teaching 5/10 Scaling Power Retention and Overcoming Institutional Inertia Alessio asks how far the team is from training frontier-scale GPT-5 models using power retention. Diego draws historical parallels to early Transformer adoption, acknowledging the substantial institutional inertia at major labs.28:38–32:45 · Guest teaching 6/10 Training Benchmarks, Hiring, and Long-Context Data Needs Diego checks back on the live benchmark showing a 5x training iteration speedup, makes recruiting pitches, and explains why standard document packing across short web texts undermines authentic long-context learning.0:55–5:27 · Guest disagreement 5/10 Transformer Bottlenecks and the Power Retention Architecture Alessio invites Diego to explain the founding thesis behind Manifest AI. Diego gives an in-depth critique of standard transformers, dismissing 256k-context claims as windowed Band-Aid solutions that degrade middle-context recall, and introduces the fixed-memory power retention architecture.5:28–8:14 · Guest disagreement 1/10 Inference Efficiency and Memory Allocation with Power Retention Alessio demonstrates familiarity with prefix caching economics in inference serving. Diego explains how dynamic KV cache sizing creates complex GPU scheduling problems, contrasting it with power retention's deterministic memory partitioning.8:14–14:27 · Guest disagreement 1/10 Vidrio Framework and Just-In-Time CUDA Kernel Optimization Alessio showcases technical understanding of low-level kernel optimization by noting Manifest's faster reimplementation of FlashAttention-2. Diego explains the mechanics of Vidrial, describing its empirical JIT parameter-sweeping approach across GPU configurations.14:27–19:22 · Guest disagreement 1/10 Model Metamorphosis: Converting StarCoder into PowerCoder Alessio prompts about the metamorphosis technique and mentions compute partners like SF Compute, but misidentifies StarCoder's base context window as 8k. Diego gently corrects it to 4k and shows training loss convergence curves in Log Cabin.19:22–22:46 · Guest disagreement 0/10 Live Demonstration of Manifesto Code Repair Tool Diego presents a live demonstration of the Manifesto code repair tool fixing broken classes inside a NanoGPT repository. Alessio provides supportive facilitation, prompting Diego to zoom in and commenting on training data overlap.22:46–26:04 · Guest disagreement 1/10 Expanding the Open-Source Ecosystem and Provider Partnerships Alessio asks whether inference providers like Together or Fireworks should convert every model to power retention and probes for potential downsides. Diego explains the strategic value for chain-of-thought inference workloads.26:04–28:38 · Guest disagreement 2/10 Scaling Power Retention and Overcoming Institutional Inertia Alessio asks how far the team is from training frontier-scale GPT-5 models using power retention. Diego draws historical parallels to early Transformer adoption, acknowledging the substantial institutional inertia at major labs.28:38–32:45 · Guest disagreement 2/10 Training Benchmarks, Hiring, and Long-Context Data Needs Diego checks back on the live benchmark showing a 5x training iteration speedup, makes recruiting pitches, and explains why standard document packing across short web texts undermines authentic long-context learning.0:55–5:27 · The hosts pushing back 0/10 Transformer Bottlenecks and the Power Retention Architecture Alessio invites Diego to explain the founding thesis behind Manifest AI. Diego gives an in-depth critique of standard transformers, dismissing 256k-context claims as windowed Band-Aid solutions that degrade middle-context recall, and introduces the fixed-memory power retention architecture.5:28–8:14 · The hosts pushing back 1/10 Inference Efficiency and Memory Allocation with Power Retention Alessio demonstrates familiarity with prefix caching economics in inference serving. Diego explains how dynamic KV cache sizing creates complex GPU scheduling problems, contrasting it with power retention's deterministic memory partitioning.8:14–14:27 · The hosts pushing back 0/10 Vidrio Framework and Just-In-Time CUDA Kernel Optimization Alessio showcases technical understanding of low-level kernel optimization by noting Manifest's faster reimplementation of FlashAttention-2. Diego explains the mechanics of Vidrial, describing its empirical JIT parameter-sweeping approach across GPU configurations.14:27–19:22 · The hosts pushing back 1/10 Model Metamorphosis: Converting StarCoder into PowerCoder Alessio prompts about the metamorphosis technique and mentions compute partners like SF Compute, but misidentifies StarCoder's base context window as 8k. Diego gently corrects it to 4k and shows training loss convergence curves in Log Cabin.19:22–22:46 · The hosts pushing back 0/10 Live Demonstration of Manifesto Code Repair Tool Diego presents a live demonstration of the Manifesto code repair tool fixing broken classes inside a NanoGPT repository. Alessio provides supportive facilitation, prompting Diego to zoom in and commenting on training data overlap.22:46–26:04 · The hosts pushing back 1/10 Expanding the Open-Source Ecosystem and Provider Partnerships Alessio asks whether inference providers like Together or Fireworks should convert every model to power retention and probes for potential downsides. Diego explains the strategic value for chain-of-thought inference workloads.26:04–28:38 · The hosts pushing back 1/10 Scaling Power Retention and Overcoming Institutional Inertia Alessio asks how far the team is from training frontier-scale GPT-5 models using power retention. Diego draws historical parallels to early Transformer adoption, acknowledging the substantial institutional inertia at major labs.28:38–32:45 · The hosts pushing back 0/10 Training Benchmarks, Hiring, and Long-Context Data Needs Diego checks back on the live benchmark showing a 5x training iteration speedup, makes recruiting pitches, and explains why standard document packing across short web texts undermines authentic long-context learning.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 2:20 Calling out fake long-context transformer claims

Diego dismisses models advertising 256k context as untrue transformers relying on windowed attention band-aids that discard critical information.

Hardest push from the hosts ▶ 24:53 Questioning why providers wouldn't convert everything

Alessio challenges the narrative by pressing Diego on whether there are hidden downsides preventing every major inference provider from immediately converting all models.

Biggest teaching moment ▶ 29:45 Exposing the illusion of pretraining context length

Diego explains to listeners and the host that 90% of pretraining web documents are under 2k tokens, meaning long-context compute is largely wasted through document packing.

The host holds their own ▶ 8:14 Demonstrating deep kernel engineering knowledge

Alessio showcases host expertise by bringing up low-level CUDA kernel execution and highlighting Manifest's performance gains over Tri Dao's original FlashAttention-2 implementation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Transformer Bottlenecks and the Power Retention Architecture 3750 Alessio invites Diego to explain the founding thesis behind Manifest AI. Diego gives an in-depth critique of standard transformers, dismissing 256k-context claims as windowed Band-Aid solutions that degrade middle-context recall, and introduces the fixed-memory power retention architecture.
Inference Efficiency and Memory Allocation with Power Retention 5611 Alessio demonstrates familiarity with prefix caching economics in inference serving. Diego explains how dynamic KV cache sizing creates complex GPU scheduling problems, contrasting it with power retention's deterministic memory partitioning.
Vidrio Framework and Just-In-Time CUDA Kernel Optimization 6610 Alessio showcases technical understanding of low-level kernel optimization by noting Manifest's faster reimplementation of FlashAttention-2. Diego explains the mechanics of Vidrial, describing its empirical JIT parameter-sweeping approach across GPU configurations.
Model Metamorphosis: Converting StarCoder into PowerCoder 5511 Alessio prompts about the metamorphosis technique and mentions compute partners like SF Compute, but misidentifies StarCoder's base context window as 8k. Diego gently corrects it to 4k and shows training loss convergence curves in Log Cabin.
Live Demonstration of Manifesto Code Repair Tool 3300 Diego presents a live demonstration of the Manifesto code repair tool fixing broken classes inside a NanoGPT repository. Alessio provides supportive facilitation, prompting Diego to zoom in and commenting on training data overlap.
Expanding the Open-Source Ecosystem and Provider Partnerships 5411 Alessio asks whether inference providers like Together or Fireworks should convert every model to power retention and probes for potential downsides. Diego explains the strategic value for chain-of-thought inference workloads.
Scaling Power Retention and Overcoming Institutional Inertia 5521 Alessio asks how far the team is from training frontier-scale GPT-5 models using power retention. Diego draws historical parallels to early Transformer adoption, acknowledging the substantial institutional inertia at major labs.
Training Benchmarks, Hiring, and Long-Context Data Needs 3620 Diego checks back on the live benchmark showing a 5x training iteration speedup, makes recruiting pitches, and explains why standard document packing across short web texts undermines authentic long-context learning.

Statements from this episode (14)

Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer With a context length of, you know, 256,000 or more, they're not using a true transformer. What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Diego Bachman Sep 23, 2025 ▶ 2:58
Disclosure
Bachman: Manifest AI is releasing Power Retention architecture with fixed-size memory
“Power retention is the specific variant that we're about to release. And it basically works by instead of the memory constantly growing, this constantly Growing KVCache. You have a memory that is a fixed size and each new token simply gets compressed into this…”
Diego Bachman Sep 23, 2025 ▶ 4:17
Assertion Open · timeframe Sep 2026
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”
Diego Bachman Sep 23, 2025 ▶ 7:46
Disclosure
Bachman: Manifest Switched Power Retention from Triton to Custom CUDA
“Actually, our initial version of power retention was written in Triton, but we realized quickly that it just didn't offer the flexibility to really squeeze the performance that we wanted out of the GPU. So we took a step back and dove into CUDA.”
Diego Bachman Sep 23, 2025 ▶ 9:33
Assertion Open · timeframe Sep 2028
Bachman: StarCoder-3B converted to Power Retention matches baseline loss in two hours
“After just 10,000 steps of training, which this training one took about two hours, this orange curve, you see that it fully matches the original loss.”
Diego Bachman Sep 23, 2025 ▶ 16:11
Assertion Open · timeframe Sep 2028
Bachman: PowerCoder-3B reaches 35% HumanEval accuracy versus StarCoder's 30%
“In the end, this converges to, I believe, about 35% accuracy on human eval, whereas the star coder baseline was about 30%.”
Diego Bachman Sep 23, 2025 ▶ 17:33
Assertion Supported
Bachman: Power Retention avoids quadratic compute scaling during long-context training
“So yeah, but we don't pay a quadratic cost. If you were looking at the star coder baseline, it would get even more, more expensive way more quickly.”
Diego Bachman Sep 23, 2025 ▶ 18:34
Disclosure
Bachman: Manifest AI built Manifesto to repair repositories using PowerCoder-3B
“We made a tool called Manifesto, where basically you can point it at any repo, including a very large repo, and it will basically fix any in that repo, or do its best to, of course. It's only a three billion parameter model, so, you know, it has its limitation…”
Diego Bachman Sep 23, 2025 ▶ 20:19
Disclosure
Bachman: Manifest plans 30-billion-parameter foundation model using Power Retention
“That's something we plan on doing at the like, thirty billion scale in the coming months.”
Diego Bachman Sep 23, 2025 ▶ 23:07
Disclosure
Bachman: Manifest AI will open-source all tools for transformer metamorphosis
“We're going to be completely open sourcing all of the pieces that you need to do this metamorphosis yourself”
Diego Bachman Sep 23, 2025 ▶ 23:16
Assertion Open · timeframe Sep 2028
Bachman: Power Retention models match original base model performance
“They'll come out with a nice shiny new, a power retention architecture that has the same performance on whatever data set they want as the original base model did.”
Diego Bachman Sep 23, 2025 ▶ 24:00
Prediction Open · timeframe Sep 2026
Bachman: Big Foundation Models Will Train on Power Retention Within a Year
“After that, I think, you know, probably within six months to a year, we're going to start to see the really big foundation models being trained in this way.”
Diego Bachman Sep 23, 2025 ▶ 28:28
Assertion Supported
Bachman: 90% of documents in web pre-training datasets are under 2k context
“If you take like any old, like the pile or one of the more modern ones, basically any common crawl derived, scrape, open web text, any of these things, you're going to find that almost all documents are very short. There's some long context documents, but the …”
Diego Bachman Sep 23, 2025 ▶ 30:30
Insight
Bachman: Compute-optimal models on internet text don't need long context
“In general, most internet text has mostly short-term structure. There's just not that much value in capturing long-term structure, and so compute optimal models on internet text actually don't have that long context, and so, of course, you're perfectly fine us…”
Diego Bachman Sep 23, 2025 ▶ 31:20
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.