Transformer

also referred to as: transformers

41 statements across 25 episodes · 14 bullish · 10 bearish · 24 people on the record · first statement Jun 20, 2023 by George Hotz · said 5 times in 3 episodes since 2025 · across every show →

Mentions by year

brought up most by Yi Tay (2), Elie Bakouch (1), Chris Manning (1)

tap a year for its mentions
00214220252026episodesmentions
01220252026episodes it came up in
00112220252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Transformer, oldest first

Jun 20, 2023 negative
Assertion Supported
Hotz: Qualcomm SNPE cannot run transformers due to missing outer product ops
“Qualcomm's S and PE can't run transformers for this reason. So most matrix multiplies in neural networks are weights times values, right? Whereas you know, when you get to the outer product in in transformers, well, it's waste times weight. It's a, it's values…”
George Hotz Jun 20, 2023 ▶ 1:15:46 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Aug 3, 2023 positive
Opinion
The 14-billion-parameter RWKV model is competitive with Transformers
“I think the RWKV scale up to They have a model at fourteen billion that seems pretty competitive with transformers.”
Tri Dao Aug 3, 2023 ▶ 46:51 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023 bullish
Prediction Not checkable as stated
RNNs will outperform Transformers in batch generation and long sequences
“I am personally bullish on, on, on RNNs. I think RNNs they don't, they essentially summarize the past into a state vector. They have fixed size, so the size doesn't grow with the history. So that means that you don't need as much memory to keep around all the …”
Tri Dao Aug 3, 2023 ▶ 49:14 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 3, 2023
Insight
Hardware and software co-evolve to favor dominant AI architectures
“There is this feedback loop where somehow The model architectures that take advantage of hardware become popular, and the hardware will also kind of evolve to optimize a little bit for that kind of architecture, and software framework software frameworks will …”
Tri Dao Aug 3, 2023 ▶ 36:22 FlashAttention-2: Making Transformers 800% faster AND exact
Aug 31, 2023 neutral
Assertion Supported
Bo Peng Created RWKV on EleutherAI Forum to Parallelize RNNs
“Blink, or Bopeng is the actual name decided basically as an individual, literally at the illiterate AI forum, decided that, hey I think we can modify recurrent neural networks, no, neural networks, based on the Apple paper, the light attention that I showed pr…”
Eugene Cheah Aug 31, 2023 ▶ 53:46 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 positive
Insight
Fine-Tuning Transformers Requires Only 1,000 to 10,000 Data Examples
“Post-Transformer it's literally, you probably need only, like, a thousand or 10,000, enough, like, data that you can literally get an intern a few weeks to just get it done, and you have a working model. It may not be that great, but frankly, every piece of da…”
Eugene Cheah Aug 31, 2023 ▶ 10:01 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 bullish
Assertion Supported
RWKV Architecture Is Proven to Scale to Any Parameter Size
“What we have already proven is that it can be scaled and trained by a transformer. How I do so, we'll cover later. And this can be scaled to as many parameters as we want.”
Eugene Cheah Aug 31, 2023 ▶ 37:52 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 bullish
Opinion
Cheah: RWKV Achieves Linear Scaling With No Trade-Offs in Reasoning
“So, so this is like literally us saying, there's no trade-offs. Yeah, you don't lose out in that process.”
Eugene Cheah Aug 31, 2023 ▶ 1:07:18 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 bearish
Prediction Didn’t hold up
Cheah: Standard Transformers Will Never Scale to Ten Million Tokens
“I think what was quick, I think it was rather quick after I concluded that transformer as it is will not scale to ten million tokens.”
Eugene Cheah Aug 31, 2023 ▶ 20:25 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 negative
Assertion Supported
Training LLMs Beyond Two Epochs Causes Overfitting and Degradation
“Anything beyond that, and we can confirm, even for our model, ours is more like closer to two, but the idea is still there, that it starts to overfit, and it starts to degrade in a lot of things.”
Eugene Cheah Aug 31, 2023 ▶ 1:26:51 RWKV: Reinventing RNNs for the Transformer Era
Aug 31, 2023 negative
Insight
Cheah: Pre-Transformer Academic Neural Network Research Is No Longer Relevant
“Frankly, almost everything that is, that matters, Ah, was basically in the past four years. Like, there were a lot of things that fit in academics that were before that, and you know, and they were mostly dealing with models that were under a billion parameter…”
Eugene Cheah Aug 31, 2023 ▶ 1:37:51 RWKV: Reinventing RNNs for the Transformer Era
Feb 8, 2024 bearish
Opinion
Prakash: Transformer architectures will not reach 5,000 tokens per second inference
“We are running into you know, the limits of how fast you can make transformers and you know, we want inference at 5000 tokens per second. And I don't think we will get there with transformers”
Vipul Ved Prakash Feb 8, 2024 ▶ 16:34 Building an open AI company - with Ce and Vipul of Together AI
Mar 14, 2024
Disclosure
Shulman: Suno focuses primarily on audio tokenization to leverage text transformers
“What we do is we benefit from all of the beautiful things people do with transformers and text, and we focus very hard basically on how do I tokenize audio in the right way.”
Mikey Shulman Mar 14, 2024 ▶ 6:54 Making Transformers Sing - with Mikey Shulman of Suno
Apr 24, 2024 neutral
Assertion Supported
Liu: Stitch Fix used Transformer models for recommendations before GPT-3
“We actually were using Transformers at Stitch Fix, like, before the GPT-III model, so we were just using Transformers for recommendation systems.”
Jason Liu Apr 24, 2024 ▶ 1:42 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Apr 27, 2024 neutral
Insight
Bach: The Transformer Architecture Is Not Its Own Meta-Algorithm
“The transformer itself is not its own meta-algorithm. Probably the person inventing the transformer didn't have a transformer running on their brain. There's something more general going on.”
Joscha Bach Apr 27, 2024 ▶ 1:37:29 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
May 31, 2024 positive
Opinion
Huang: Google's internal AI tooling was far superior to competitors
“Google was using AI for systems before everybody else too, right? They invented a transformer, and their internal set of tooling was just so far superior to everything else. Like, it's really hard for people to go back after seeing that.”
Mark Huang May 31, 2024 ▶ 7:41 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Jul 23, 2024 neutral
Insight
Scialom: Fixed compute per token in transformers forces models to generate extra tokens to think
“We are lacking of flexibility for pre-training architecture, transformers, where We spend the same amount of compute per token. And so, because of that, how can you, like, mitigate this? By generating more tokens, so more thoughts, more compute, because you ha…”
Thomas Scialom Jul 23, 2024 ▶ 51:19 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Sep 21, 2024 positive
Disclosure
Karpathy: llm.c trains transformers in C with minimal C++
“We're training transformers in C at a pinch of C++.”
Andrej Karpathy Sep 21, 2024 ▶ 1:38 llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
Nov 11, 2024 negative
Assertion Not checkable as stated
Polu: Transformers failed at code fuzzing because they are too slow
“We're trying to apply transformers to code fuzzing. So code fuzzing, you have kind of an, sorry, an algorithm that goes really fast and tries to mutate the inputs of a library to find bugs. And we try to apply a transformer to that and do reinforcement learnin…”
Stanislas Polu Nov 11, 2024 ▶ 3:46 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 11, 2024 positive
Insight
Polu: Combining LLMs with formal math pairs creativity with proof verification
“Transformers are very creative, but yet they do mistakes. And formal math systems are the ability to verify a proof. And the tactics they can use to solve problems are very mechanical. So you miss the creativity. And so the idea was to try to explore both toge…”
Stanislas Polu Nov 11, 2024 ▶ 6:30 Agents @ Work: Dust.tt — with Stanislas Polu
Dec 24, 2024 positive
Insight
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Eugene Cheah Dec 24, 2024 ▶ 28:15 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Jul 2, 2025 neutral
What-if
Morris: ChatGPT could likely have been built using RNNs instead of Transformers
“And I think like, we honestly probably could have gotten this with RNNs. I know like the scaling laws paper shows that RNNs have worse curves for scaling, but probably people would have been like, I bet you could have built chat GPT with a very sophisticated R…”
Jack Morris Jul 2, 2025 ▶ 1:10:28 Information Theory for Language Models: Jack Morris
Jul 2, 2025 neutral
Assertion Supported
Morris: 32-bit transformer models store only 3.6 to 3.9 bits per parameter
“Transformers that are trained in 32 bit precision, we approximate can store about 3.6 bits of information to maybe 3.9 bits somewhere in there per parameter.”
Jack Morris Jul 2, 2025 ▶ 56:00 Information Theory for Language Models: Jack Morris
Jul 28, 2025 negative
Insight
Mohan: Off-the-shelf serving frameworks leave significant FLOP utilization on the table
“The open source serving. Offerings are just, I will say not great in that they aren't customized to transformers and these kinds of workloads where I have high latency and I want to like batch requests and I want to batch requests while keeping latency low. Bu…”
Varun Mohan Jul 28, 2025 ▶ 14:18 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Aug 6, 2025 neutral
Assertion Not checkable as stated
Swyx: Most AI researchers are scaling Transformers, not alternative architectures
“I think most people that I talk to are not seriously pursuing alternative architectures. There are some notable exceptions, primarily together with the Mombard architecture recursal with RWKV. There's like the XLSTM that was created by a separate Hogwriter. An…”
Shawn Wang Aug 6, 2025 ▶ 1:10:49 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Aug 18, 2025
Insight
Sohmers: Transformer inference is memory-bound with a 1:1 FLOP-to-byte ratio
“And on the other side of this chart, you have the case of a transformer where when you're actually, you know, doing attention, Or really just any case where you're doing, you're fundamentally doing matrix vector multiplication rather than matrix matrix multipl…”
Thomas Sohmers Aug 18, 2025 ▶ 9:27 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Aug 29, 2025 neutral
Opinion
Morcos: The Transformer is just one of many equivalently good architectures
“And one of my like more controversial viewpoints, I think, is that I think the transformer is a great advance to be sure, but I think it's one of a very large Set of equivalently good architectures that we could have found. And there are many, many ways we cou…”
Ari Morcos Aug 29, 2025 ▶ 12:35 Better Data is All You Need — Ari Morcos, Datology
Sep 23, 2025 negative
Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer With a context length of, you know, 256,000 or more, they're not using a true transformer. What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Diego Bachman Sep 23, 2025 ▶ 2:58 ⚡️ Beyond Transformers with Power Retention
Sep 23, 2025 neutral
Insight
Bachman: Compute-optimal models on internet text don't need long context
“In general, most internet text has mostly short-term structure. There's just not that much value in capturing long-term structure, and so compute optimal models on internet text actually don't have that long context, and so, of course, you're perfectly fine us…”
Diego Bachman Sep 23, 2025 ▶ 31:20 ⚡️ Beyond Transformers with Power Retention
Nov 25, 2025 neutral
Insight
Johnson: Transformers are natively models of sets, not sequences
“Transformers are actually not a model of sequences. A transformer is natively a model of sets.”
Justin Johnson Nov 25, 2025 ▶ 56:45 After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
Jan 23, 2026 neutral
Opinion
Yi Tay: Today's Models Likely Couldn't Invent the Transformer from Pre-2015 Data
“Even today's models, they might not even be able to invent the transformer. Like, if you freeze the time at a certain time, and even you bring the time, I mean, the model is a transformer, so I just say there's no, assuming there's no leakage.”
Yi Tay Jan 23, 2026 ▶ 31:59 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Jan 23, 2026 bullish
Prediction Not checkable as stated
Yi Tay: The Architecture That Achieves AGI Will Still Be a Transformer
“It will be a transformer, I think. Like people, it depends on what you call it, but I think unless the paradigm shifts completely, which is, I mean, as a scientist, you cannot like completely say no to like that, this would never happen. But my feeling is that…”
Yi Tay Jan 23, 2026 ▶ 46:26 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Jan 23, 2026 neutral
Opinion
Yi Tay: The AI Industry Is Stuck in a Transformer Local Minimum
“So now we are like in this local minima of like transformers, everything, everything, right? Maybe it's not easy to like get totally out Of this, because also a lot of people's investment optimization have been done. So the things that play well needs to play …”
Yi Tay Jan 23, 2026 ▶ 50:43 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Jan 23, 2026
Insight
Yi Tay: Gradient descent learning paradigm is AI's bottleneck, not architecture
“It's not architecture itself. That's, that there's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the to…”
Yi Tay Jan 23, 2026 ▶ 49:07 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Feb 12, 2026 positive
Assertion Supported
Dean: Transformers delivered 10x to 100x compute efficiency over LSTMs
“Transformers similarly gave you a 10 X to a hundred X improvement in, you know compute cost to a given quality level versus say LSTMs at the time.”
Jeff Dean Feb 12, 2026 ▶ 1:02:15 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Apr 2, 2026 positive
Opinion
Manning: Transformer internal weights can act as joint representations for world models
“I'm not actually convinced that's right, because although the token production is this autoregressive process that's heading, you know, left to right, I guess don't have to be left or right, but anyway, in sequence of tokens, we could have right to left Arabic…”
Chris Manning Apr 2, 2026 ▶ 20:57 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Apr 3, 2026 positive
Insight
Andreessen: AlexNet and transformers were the true inflection points of modern AI
“I think the real story is it was the Alex net basically breakthrough in like 2013. That was the real knee in the curve. And then it was obviously the transformer breakthrough in 17.”
Marc Andreessen Apr 3, 2026 ▶ 3:54 Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
Jun 1, 2026 neutral
Insight
Ethan He: Training transformers directly on raw image pixels is impossible
“If you're trying, if you can technically, theoretically train image or video models on pure pixels, but the problem is that the, it's a lot of tokens. So like one image, like it's a thousand by a thousand is like one million tokens, one million pixels. It's im…”
Ethan He Jun 1, 2026 ▶ 15:30 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 negative
Insight
Ethan He: Training models directly on MP4 tokens is extremely difficult
“So people actually have tried that, but the main challenge is the latent space for the MP four tokens are not, we're not very comprehensible for the models. It's extremely hard to train on that.”
Ethan He Jun 1, 2026 ▶ 20:57 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Aug 26, 2026 positive
Insight
Fourier neural operators scale quasi-linearly, avoiding transformers' quadratic complexity
“If we were to use transformers and we require a very high resolution, it would become untenable because of the quadratic complexity and all to all connections. On the other hand, if you did that with Fourier transforms, we have like quasi linear complexity and…”
Anima Anandkumar Aug 26, 2026 ▶ 21:51 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Aug 26, 2026 bearish
Prediction Not checkable as stated
Transformers will never scale to high-resolution 4D physics simulations
“So forget ever having a transformer for anything of this scale. All of the world's compute will not be enough. And first of all, they all have to be co-located to be able to ever do this. So that's why we need other architectures.”
Anima Anandkumar Aug 26, 2026 ▶ 29:41 🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.