Transformer Architecture

topic on 9 shows · 20 statements across 20 episodes

More or Less Latent Space Lenny's Podcast No Priors WTF is with Nikhil Kamath the MAD Podcast Big Technology All-In 20VC

20 statements about Transformer Architecture, every show

Anandkumar: Standard Transformers cannot scale to 5-trillion context lengths for physics
“On the other hand, if you think about using transformer architectures that have worked so well for language, that just wouldn't be able to support a five trillion context length. No matter all the compute in the world is thrown at it. So that kind of quadratic…”
Anima Anandkumar Sep 4, 2026 ▶ 10:58 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Midha: Anthropic's velocity came from standardizing on the transformer architecture
“Like, one of the reasons Anthropic has had extraordinary sort of velocity is because they picked the transform architecture and said, this is simple, let's double down on it, right? And now, luckily, there's enough investment going into space that we can affor…”
Anjney Midha Jun 18, 2026 ▶ 26:35 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Ries: Corporations are the oldest form of artificial intelligence
“Corporations, organizations are the oldest form of artificial intelligence on the planet. They are an example of this emergent intelligence, the same scientific principle that makes the transformer architecture work and appear intelligent. That same principle …”
Eric Ries May 10, 2026 ▶ 1:31:20 How Anthropic, Costco, and Patagonia all build incorruptible companies | Eric Ries
MAD What-if
AI capabilities would have been achieved even without inventing transformers
“I think if we hadn't invented the transformer, we would have gotten there with whatever LSTM you know, state space model, whatever, anything else people were developing, we would have gotten there.”
Zico Kolter May 7, 2026 ▶ 1:08:19 OpenAI Board Member Zico Kolter: Modern AI Is Just 200 Lines of Code
MAD Insight
Vision Transformers benefited heavily from software infrastructure built for language models
“There was also benefit of doing that simply because the rest of the, that the machine learning field, which was working on, on, on language, they were using this, like architecture. So they were building infra for it, making it faster. And, you know, like the,…”
Mostafa Dehghani Apr 2, 2026 ▶ 41:14 AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
MAD Assertion Supported
Transformer remains state of the art for LLM performance
“I would say right now, yes, because it's still the state of the art. So there is nothing really better in terms of state of the art performance, getting better quality results.”
Sebastian Raschka Jan 29, 2026 ▶ 2:10 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
LATENT SPACE Prediction Not checkable as stated
Feldman: Transformer architecture has several more years of viability
“And finally, I think the transformer, the current architecture has a way still to run. I think we will see that for several years more.”
Andrew Feldman Oct 1, 2025 ▶ 28:32 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
20VC Assertion Supported
Frosst: Core transformer architecture has remained largely unchanged for a decade
“The models themselves, like transformer architecture, which is the original model that, yeah, that was introduced in Hasn't changed very much, right? Like all, the whole industry is still using transformers. We've changed the way we train them, but the model a…”
Nick Frosst Sep 1, 2025 ▶ 4:41 Cohere Founder, Nick Frosst: How To Compete with OpenAI & Anthropic, and Sam Altman’s AI Disservice · 20VC with Harry Stebbings
MAD Assertion Supported
Gomez: Google quickly integrated Transformer architecture into Search and Translate
“No, they jumped all over it. So it went to production inside of search, inside of translate, like the existing product suite. And so they, to say that they didn't adopt the transformer architecture would not be correct.”
Aidan Gomez Jun 5, 2025 ▶ 11:35 Inside the Paper That Changed AI Forever - Cohere CEO Aidan Gomez on 2025 Agents
Shankar: LLM failure modes and evaluation techniques have largely stabilized
“I think techniques have stabilized. I think the kinds of failure modes of LLMs, I mean, they're still there, but it's not like changing every single day. We know that LLMs are bad at certain things. We know a little bit more about say limitations of the transf…”
Shreya Shankar Mar 13, 2025 ▶ 11:21 [Lightning Pod] Evals: How to Improve AI Consistently — with Hamel Husain and Shreya Shankar
MAD Opinion
Kiela: Attention mechanism, not Transformers, was the real AI breakthrough
“So I would say, and maybe I'm biased because one of my best friends is, is on the original attention paper, but that was the real breakthrough. It's just like figuring out that you have this attention mechanism that actually allows you to yeah, to do a much be…”
Douwe Kiela Mar 6, 2025 ▶ 20:19 Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
WTF Opinion
LLMs primarily perform data retrieval and possess very little actual reasoning
“If it's text, they will regurgitate solutions to puzzles. They will, you know, give you answers to questions you may have. It's mostly retrieval. There's a very tiny bit of reasoning, but really not much and that's an important limitation.”
Yann LeCun Nov 27, 2024 ▶ 57:09 WTF is Artificial Intelligence Really? | Yann LeCun x Nikhil Kamath | People by WTF Ep #4 · Nikhil Kamath
LATENT SPACE Assertion Not checkable as stated
Polu: OpenAI Believed in Transformer Scaling Pre-Kaplan Paper
“Before that, there really was a strong belief in, in scale. I think it was just the belief that the transformer was a generic enough architecture that you could learn anything, and that this was just a question of scaling.”
Stanislas Polu Nov 11, 2024 ▶ 14:17 Agents @ Work: Dust.tt — with Stanislas Polu
MORE OR LESS Prediction Not checkable as stated
Dave Morin: AI progress will probably plateau longer than people think
“A simple way to think of it is like, we're probably gonna be kind of where we are right now for a lot longer than people think. If the current architecture and the current energy and the current capital needs are the gating factors.”
Dave Morin Oct 11, 2024 ▶ 42:17 #68: Google's Antitrust Saga Continues… · More or Less Podcast
NO PRIORS Opinion
Sutskever: Transformers can reach AGI; alternatives only offer compute efficiency
“So it's better to think about it in terms of compute efficiency rather than in terms of, can it get there at all? I think at this point, the answer is obviously yes.”
Ilya Sutskever Nov 2, 2023 ▶ 28:00 No Priors Ep. 39 | With OpenAI Co-Founder & Chief Scientist Ilya Sutskever
BIG TECHNOLOGY Assertion Not checkable as stated
Nemade: Google teams only noticed Transformer paper after citations grew
“I think once it started garnering a lot of citations over the months and quarters, that's when people started paying attention, even inside of Google that, Hey, this transformer architecture seems amazing.”
Gaurav Nemade Oct 30, 2023 ▶ 7:49 Why Google Never Shipped LaMDA Its ChatGPT Predecessor
NO PRIORS Opinion
Gil: GPU and Transformer lock-in is stronger than 1990s Wintel
“And so this is, I feel like, a stronger version of that in some sense, where you have the underlying compute architecture and the most important model reinforcing each other in a way that kind of locks both of them in.”
Elad Gil Sep 14, 2023 ▶ 40:04 No Priors Ep. 32 | With NEAR’s Illia Polosukhin
NO PRIORS Insight
Uszkoreit: Community optimism drove Transformer adoption and success
“The other main contributor, I think, to the success of this architecture was optimism and hope. So suddenly you were in a situation where, for whatever reason, a bunch of things that people tried with this started to work, and then more started to work, and th…”
Jakob Uszkoreit Aug 24, 2023 ▶ 7:00 No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
20VC Assertion Not checkable as stated
LeCun: Transformer scaling with self-supervised learning surprised AI researchers
“The fact that self-supervised learning methods applied to transformer architectures Work amazingly well, and they work, you know, way beyond what we could have expected. The fact that we can do basically train systems to understand language, translate language…”
Yann LeCun May 15, 2023 ▶ 7:03 Yann LeCun: Meta’s New AI Model LLaMA; Why Elon is Wrong about AI; Open-source AI Models | E1014 · 20VC with Harry Stebbings
ALL-IN Assertion Contradicted
Palihapitiya claims transformer LLMs trained on identical data yield identical answers
“I'm saying when you look at transformer architecture today, every LLM that you write on the same corpus of underlying data for training will get to the same answer.”
Chamath Palihapitiya Feb 11, 2023 ▶ 56:28 E115: The AI Search Wars: Google vs. Microsoft, Nordstream report, State of the Union

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.