transformer architecture

also referred to as: transformer architectures

5 statements across 5 episodes · 2 bullish · 1 bearish · 5 people on the record · first statement Nov 11, 2024 by Stanislas Polu · across every show →

Everything said about transformer architecture, oldest first

Nov 11, 2024
Assertion Not checkable as stated
Polu: OpenAI Believed in Transformer Scaling Pre-Kaplan Paper
“Before that, there really was a strong belief in, in scale. I think it was just the belief that the transformer was a generic enough architecture that you could learn anything, and that this was just a question of scaling.”
Stanislas Polu Nov 11, 2024 ▶ 14:17 Agents @ Work: Dust.tt — with Stanislas Polu
Mar 13, 2025 neutral
Insight
Shankar: LLM failure modes and evaluation techniques have largely stabilized
“I think techniques have stabilized. I think the kinds of failure modes of LLMs, I mean, they're still there, but it's not like changing every single day. We know that LLMs are bad at certain things. We know a little bit more about say limitations of the transf…”
Shreya Shankar Mar 13, 2025 ▶ 11:21 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
Oct 1, 2025 positive
Prediction Not checkable as stated
Feldman: Transformer architecture has several more years of viability
“And finally, I think the transformer, the current architecture has a way still to run. I think we will see that for several years more.”
Andrew Feldman Oct 1, 2025 ▶ 28:32 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Jun 18, 2026 positive
Opinion
Midha: Anthropic's velocity came from standardizing on the transformer architecture
“Like, one of the reasons Anthropic has had extraordinary sort of velocity is because they picked the transform architecture and said, this is simple, let's double down on it, right? And now, luckily, there's enough investment going into space that we can affor…”
Anjney Midha Jun 18, 2026 ▶ 26:35 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Sep 4, 2026 negative
Insight
Anandkumar: Standard Transformers cannot scale to 5-trillion context lengths for physics
“On the other hand, if you think about using transformer architectures that have worked so well for language, that just wouldn't be able to support a five trillion context length. No matter all the compute in the world is thrown at it. So that kind of quadratic…”
Anima Anandkumar Sep 4, 2026 ▶ 10:58 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.