People, every show

Diego Bachman

Co-founder & CEO, Manifest AI. On 1 show, 1 appearance. The Shows tab opens the full record on each.

founderexecutivescientist@jacobmbuckman ↗jacobbuckman.com ↗

Jacob Buckman is a machine learning researcher and the co-founder and CEO of Manifest AI, an independent research lab addressing AI context-length limits. He is known for developing the Power Retention architecture to eliminate growing KV-cache bottlenecks and the Vidrial GPU kernel framework.

1shows
1appearances
14statements
3resolved
2supported
1contradicted
67%fully supported

Everything Diego Bachman said on any show that made the record, most notable first. Each card names its show and opens the statement there.

LATENT SPACE Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer With a context length of, you know, 256,000 or more, they're not using a true transformer. What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Diego Bachman Sep 23, 2025 ▶ 2:58 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Assertion Open · timeframe Sep 2026
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”
Diego Bachman Sep 23, 2025 ▶ 7:46 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Assertion Open · timeframe Sep 2028
Bachman: Power Retention models match original base model performance
“They'll come out with a nice shiny new, a power retention architecture that has the same performance on whatever data set they want as the original base model did.”
Diego Bachman Sep 23, 2025 ▶ 24:00 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Assertion Open · timeframe Sep 2028
Bachman: StarCoder-3B converted to Power Retention matches baseline loss in two hours
“After just 10,000 steps of training, which this training one took about two hours, this orange curve, you see that it fully matches the original loss.”
Diego Bachman Sep 23, 2025 ▶ 16:11 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Prediction Open · timeframe Sep 2026
Bachman: Big Foundation Models Will Train on Power Retention Within a Year
“After that, I think, you know, probably within six months to a year, we're going to start to see the really big foundation models being trained in this way.”
Diego Bachman Sep 23, 2025 ▶ 28:28 ⚡️ Beyond Transformers with Power Retention
Bachman: Compute-optimal models on internet text don't need long context
“In general, most internet text has mostly short-term structure. There's just not that much value in capturing long-term structure, and so compute optimal models on internet text actually don't have that long context, and so, of course, you're perfectly fine us…”
Diego Bachman Sep 23, 2025 ▶ 31:20 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Disclosure
Bachman: Manifest AI is releasing Power Retention architecture with fixed-size memory
“Power retention is the specific variant that we're about to release. And it basically works by instead of the memory constantly growing, this constantly Growing KVCache. You have a memory that is a fixed size and each new token simply gets compressed into this…”
Diego Bachman Sep 23, 2025 ▶ 4:17 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Disclosure
Bachman: Manifest Switched Power Retention from Triton to Custom CUDA
“Actually, our initial version of power retention was written in Triton, but we realized quickly that it just didn't offer the flexibility to really squeeze the performance that we wanted out of the GPU. So we took a step back and dove into CUDA.”
Diego Bachman Sep 23, 2025 ▶ 9:33 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Assertion Open · timeframe Sep 2028
Bachman: PowerCoder-3B reaches 35% HumanEval accuracy versus StarCoder's 30%
“In the end, this converges to, I believe, about 35% accuracy on human eval, whereas the star coder baseline was about 30%.”
Diego Bachman Sep 23, 2025 ▶ 17:33 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Assertion Supported
Bachman: Power Retention avoids quadratic compute scaling during long-context training
“So yeah, but we don't pay a quadratic cost. If you were looking at the star coder baseline, it would get even more, more expensive way more quickly.”
Diego Bachman Sep 23, 2025 ▶ 18:34 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Assertion Supported
Bachman: 90% of documents in web pre-training datasets are under 2k context
“If you take like any old, like the pile or one of the more modern ones, basically any common crawl derived, scrape, open web text, any of these things, you're going to find that almost all documents are very short. There's some long context documents, but the …”
Diego Bachman Sep 23, 2025 ▶ 30:30 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Disclosure
Bachman: Manifest AI built Manifesto to repair repositories using PowerCoder-3B
“We made a tool called Manifesto, where basically you can point it at any repo, including a very large repo, and it will basically fix any in that repo, or do its best to, of course. It's only a three billion parameter model, so, you know, it has its limitation…”
Diego Bachman Sep 23, 2025 ▶ 20:19 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Disclosure
Bachman: Manifest plans 30-billion-parameter foundation model using Power Retention
“That's something we plan on doing at the like, thirty billion scale in the coming months.”
Diego Bachman Sep 23, 2025 ▶ 23:07 ⚡️ Beyond Transformers with Power Retention
LATENT SPACE Disclosure
Bachman: Manifest AI will open-source all tools for transformer metamorphosis
“We're going to be completely open sourcing all of the pieces that you need to do this metamorphosis yourself”
Diego Bachman Sep 23, 2025 ▶ 23:16 ⚡️ Beyond Transformers with Power Retention

One line per show, most statements first. The link opens Diego's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
LATENT SPACELEDGER Co-founder & CEO, Manifest AI 1 14 67% 2/3 full record on Latent Space →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.