Prediction Open AI assessment confidence: 85% certainty 3/5 debate potential 3/5

Bachman: Big Foundation Models Will Train on Power Retention Within a Year

Diego Bachman · ⚡️ Beyond Transformers with Power Retention · Sep 23, 2025 · at 28:28

Diego Bachman of Manifest AI discusses the timeline and community adoption trajectory for scaling the Power Retention architecture to large foundation models.

0:00 / 0:07exact quote · 7.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“After that, I think, you know, probably within six months to a year, we're going to start to see the really big foundation models being trained in this way.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Diego Bachman

Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer With a context length of, you know, 256,000 or more, they're not using a true transformer. What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Diego Bachman Sep 23, 2025 ▶ 2:58 ⚡️ Beyond Transformers with Power Retention
Assertion Open · timeframe Sep 2026
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”
Diego Bachman Sep 23, 2025 ▶ 7:46 ⚡️ Beyond Transformers with Power Retention
Assertion Open · timeframe Sep 2028
Bachman: Power Retention models match original base model performance
“They'll come out with a nice shiny new, a power retention architecture that has the same performance on whatever data set they want as the original base model did.”
Diego Bachman Sep 23, 2025 ▶ 24:00 ⚡️ Beyond Transformers with Power Retention
Assertion Open · timeframe Sep 2028
Bachman: StarCoder-3B converted to Power Retention matches baseline loss in two hours
“After just 10,000 steps of training, which this training one took about two hours, this orange curve, you see that it fully matches the original loss.”
Diego Bachman Sep 23, 2025 ▶ 16:11 ⚡️ Beyond Transformers with Power Retention
Insight
Bachman: Compute-optimal models on internet text don't need long context
“In general, most internet text has mostly short-term structure. There's just not that much value in capturing long-term structure, and so compute optimal models on internet text actually don't have that long context, and so, of course, you're perfectly fine us…”
Diego Bachman Sep 23, 2025 ▶ 31:20 ⚡️ Beyond Transformers with Power Retention
Disclosure
Bachman: Manifest AI is releasing Power Retention architecture with fixed-size memory
“Power retention is the specific variant that we're about to release. And it basically works by instead of the memory constantly growing, this constantly Growing KVCache. You have a memory that is a fixed size and each new token simply gets compressed into this…”
Diego Bachman Sep 23, 2025 ▶ 4:17 ⚡️ Beyond Transformers with Power Retention
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.