Transformer

also referred to as: transformers

28 statements across 20 episodes · 17 bullish · 5 bearish · 19 people on the record · first statement Apr 25, 2023 by Dr. Percy Liang · said 13 times in 6 episodes since 2023 · across every show →

Mentions by year

brought up most by Sarah Guo (4), Jensen Huang (4), Elad Gil (2), Liam Fedus (1), Illia Polosukhin (1), Andrej Karpathy (1)

tap a year for its mentions
00821542023202420252026episodesmentions
0242023202420252026episodes it came up in
001.52342023202420252026episodesmentions per episode
2026 1 mention in 1 episode
2024 1 mention in 1 episode
2023 11 mentions in 4 episodes 3 per episode

every mention, scene by scene, with the transcript →

Everything said about Transformer, oldest first

Apr 25, 2023 positive
Assertion Supported
Liang: Stanford researchers developed attention-free architectures competitive with transformers
“So one of my colleagues, Chris Ray and his students have developed other architectures, which are actually at smaller scales, competitive with transformers. And actually don't require the central operation of attention.”
Dr. Percy Liang Apr 25, 2023 ▶ 28:26 No Priors Ep. 7 | With Stanford Professor Dr. Percy Liang
Apr 25, 2023 positive
Insight
Srinivas: Academic AI researchers should pursue radical alternatives to transformers
“I think it's best to look for alternatives to the transformer, alternatives to like language models, that sort of radical directions, then Trying to improve them because there's so much incentive for the existing companies to do that.”
Aravind Srinivas Apr 25, 2023 ▶ 36:14 No Priors Ep. 9 | With Perplexity AI’s Aravind Srinivas and Denis Yarats
Apr 25, 2023 positive
Insight
Self-Attention Brought GPU Parallelism to Sequence Modeling
“The insight here was, Hey, you can use the same attention thing to like, look back at the past of the sequence that you're trying to produce. And you know, the beauty is that the, that It runs great on on GPUs and CPUs, and it's kind of parallel to, like, how …”
Noam Shazeer Apr 25, 2023 ▶ 6:21 No Priors Ep. 12 | With Noam Shazeer
Apr 25, 2023 positive
Insight
Transformers Beat Recurrent Models by Processing Whole Sequences at Once
“The magic of transformer kind of like convolutions is that you get to process the whole sequence at once. I mean, it still talks, you know, it's still a function of like the, you know, the predictions for the later words are dependent on what the earlier words…”
Noam Shazeer Apr 25, 2023 ▶ 4:50 No Priors Ep. 12 | With Noam Shazeer
May 1, 2023
Insight
Valenzuela: It takes 12-24 months to understand new AI breakthroughs
“The moment something gets released, like, let's say transformers or a particular piece of technology that you think would be interesting or could be worth experimenting with, I think it takes a collective set of months, like, 12, 24 months sometimes to underst…”
Cristobal Valenzuela May 1, 2023 ▶ 13:11 No Priors Ep. 2 | With Runway ML’s Cristobal Valenzuela
Aug 24, 2023 neutral
Insight
Uszkoreit: Transformer breakthrough was driven by accelerator hardware fit
“And if you want to look at, say, the biggest differences, for example, between the transformer, as it was described in the attention is all you need paper, and some of its ancestors, like this decomposable attention model, the big difference is just that the t…”
Jakob Uszkoreit Aug 24, 2023 ▶ 3:59 No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
Sep 14, 2023 positive
Prediction Held up
Polosukhin: Market full of Transformer-optimized accelerators will launch by 2024
“And so like, we're going to have a, you know, a market full of hardware accelerators, which are still optimized for transformers, or at least like similar structured architectures hitting the market like this year and next year.”
Illia Polosukhin Sep 14, 2023 ▶ 41:30 No Priors Ep. 32 | With NEAR’s Illia Polosukhin
Sep 14, 2023 bullish
Prediction Not checkable as stated
Polosukhin: Transformers will be very hard for alternative architectures to match
“The simplicity of this architecture and like, indeed, like the amount of optimization that's going into this right now is just, it will be really hard to match”
Illia Polosukhin Sep 14, 2023 ▶ 34:48 No Priors Ep. 32 | With NEAR’s Illia Polosukhin
Oct 26, 2023 bullish
Insight
Gil: AI capabilities represent a major discontinuity, not a linear progression
“I feel like they're in general, people are viewing AI as this continuum where it's like, it's a CNN, RNN, and now we have transformers and it's just a straight line. And instead, obviously it's a big discontinuity in terms of capabilities. And I think most peo…”
Elad Gil Oct 26, 2023 ▶ 19:26 No Priors Ep. 38 | With Material Security Co-Founder Ryan Noon
Nov 2, 2023 positive
Assertion Not checkable as stated
Sutskever: The AI formula is training larger transformers on more data
“There is one specific formula right now that everyone is doing. And this formula is train a larger and larger transformer on more and more data.”
Ilya Sutskever Nov 2, 2023 ▶ 13:16 No Priors Ep. 39 | With OpenAI Co-Founder & Chief Scientist Ilya Sutskever
Dec 14, 2023 bearish
Opinion
Guo: Confidence is dropping that transformers will remain dominant AI architecture
“If you'd asked me how committed are people to transformers as the dominant architecture at the end of 20, 23, I'd say very committed and I feel less confident about that today.”
Sarah Guo Dec 14, 2023 ▶ 31:51 No Priors Ep. 44 | With Former Square CEO Alyssa Henry
Dec 21, 2023 bearish
Prediction Not checkable as stated
Hoffman: Novel AI architectures replacing transformers are over two years away
“I think the new architectures, there's a bunch of stuff that I've been trying to work with, but I don't think new architectures will be one to two years. I'd be surprised if it was.”
Reid Hoffman Dec 21, 2023 ▶ 39:19 No Priors Ep. 45 | With Reid Hoffman
Apr 18, 2024 negative
Opinion
Doshi: DiT Is Likely Not the Right Architecture for Generative Vision
“So I think the architecture is like mostly going, is most likely going to change. I don't think that DIT is like the right architecture, but transformers certainly.”
Suhail Doshi Apr 18, 2024 ▶ 18:12 No Priors Ep. 60 | With Playground AI Founder Suhail Doshi
May 16, 2024 positive
Disclosure
Suno builds on transformers and focuses innovation on audio tokenization
“We don't make it a secret that these are just transformers. This is somewhat our backgrounds doing text before, but also transformers scale nicely. A lot of work ends up being done for you by the open source text community, which is always really nice. We can …”
Mikey Shulman May 16, 2024 ▶ 6:07 No Priors Ep. 64 | With Suno CEO and Co-Founder Mikey Shulman
Jun 27, 2024 bearish
Opinion
Gu: Transformers struggle significantly on raw pixel or audio waveform data
“People think that like you can throw a transformer at like anything and it just works. Actually it doesn't really like if you try to throw it at like the raw pixel level or the raw sample level and in audio waveforms I think it doesn't work nearly as well.”
Albert Gu Jun 27, 2024 ▶ 8:16 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu
Jun 27, 2024 positive
Insight
Gu: State Space Models can be applied to almost all data types
“So it really can be applied to pretty much everything. So just like kind of Transformers, these are applied to everything. So can these sort of models over the course of research over a few years, we kind of realized that there are different advantages for dif…”
Albert Gu Jun 27, 2024 ▶ 6:51 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu
Jun 27, 2024 neutral
Assertion Supported
Gu: Early SSMs excelled at raw signals but lagged Transformers on text
“The first types of models we were looking at were really good actually at modeling kind of these raw waveforms raw pixels, things like that, but not as good at modeling text, and transformers are way better there.”
Albert Gu Jun 27, 2024 ▶ 7:49 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu
Aug 30, 2024 positive
Insight
Steinberger: In-context learning functions as an online optimizer
“I think of that as some sort of a, as an online optimizer in a sense that instead of compressing a set of data, you're trying to learn an optimizer.”
Eric Steinberger Aug 30, 2024 ▶ 5:11 No Priors Ep. 79 | With Magic.dev CEO and Co-Founder Eric Steinberger
Sep 5, 2024 neutral
Assertion Not checkable as stated
Karpathy: RoPE is the only major transformer architecture change in five years
“The transformer hasn't changed that much. You know, we've added the rope positional and the rope relative positional encodings. That's like the major change. Everything else doesn't really matter too much. It's like plus three percent on a small few things. Bu…”
Andrej Karpathy Sep 5, 2024 ▶ 16:51 No Priors Ep. 80 | With Andrej Karpathy from OpenAI and Tesla
Sep 5, 2024 positive
Opinion
Karpathy: Transformers are a more efficient system than the human brain
“I think transformers are actually better than the human brain in a bunch of ways. I think they're actually a lot more efficient system. And the reason they don't work as good as the human brain is mostly data issue, roughly speaking, is the first order approxi…”
Andrej Karpathy Sep 5, 2024 ▶ 20:48 No Priors Ep. 80 | With Andrej Karpathy from OpenAI and Tesla
Sep 5, 2024 positive
Insight
Karpathy: Clean AI scaling laws are a property of transformers, not LSTMs
“When people talk about the scaling loss in neural networks, the scaling laws are actually a to a large extent of a property of the transformer. Before the transformer, people were playing with LSTMs and stacking them, etc. You don't actually get like clean sca…”
Andrej Karpathy Sep 5, 2024 ▶ 14:59 No Priors Ep. 80 | With Andrej Karpathy from OpenAI and Tesla
Sep 5, 2024 bullish
Opinion
Karpathy: No single technical bottleneck is blocking robotics, just data grunt work
“I don't know that there's, like, any individual impediments that I'm, like, really familiar with. I just think it's a lot of grunt work. A lot of, like, the tools are available. Transformers are this beautiful, like, blob of tissue. You can just get just arbit…”
Andrej Karpathy Sep 5, 2024 ▶ 14:18 No Priors Ep. 80 | With Andrej Karpathy from OpenAI and Tesla
Oct 24, 2024 negative
Opinion
Dolgov: Off-the-shelf AI models cannot achieve human-beating driverless safety records
“The power of transformers, the power of realism is mind-blowing, right? So with just a little bit of effort, you get something on the road, and it works. You can, you know, drive, I don't know, 1000 of miles, and we just, it will blow your mind. But then is th…”
Dmitri Dolgov Oct 24, 2024 ▶ 39:52 No Priors Ep. 87 | With Co-CEO of Waymo Dmitri Dolgov
Mar 20, 2025 positive
Disclosure
Physical Intelligence uses transformers and pre-trained VLMs for robot models
“And in terms of the architecture, we're using transformers and we are using pre-trained models, pre-trained vision language models, and that allows you to leverage all of the rich information on the internet.”
Chelsea Finn Mar 20, 2025 ▶ 6:51 No Priors Ep. 107 | With Physical Intelligence Co-Founder Chelsea Finn
May 15, 2025
Disclosure
Glean Used Transformers for Semantic Matching in Version One
“The version, one of our product actually already used transformers for semantic, you know, matching”
Arvind Jain May 15, 2025 ▶ 3:03 No Priors Ep. 115 | With Glean Founder and CEO Arvind Jain
May 15, 2025 positive
Insight
Jain: Enterprise search requires transformer models due to scarce user signals
“On the web, even if you don't have good semantic understanding, there is so much that you're going to learn from people's behavior because, you know, you have a billion people, you know, coming and using your product. In the enterprise, you don't have that lux…”
Arvind Jain May 15, 2025 ▶ 8:23 No Priors Ep. 115 | With Glean Founder and CEO Arvind Jain
Oct 31, 2025 bullish
Opinion
Guo: Reasoning is the biggest AI architecture shift since the transformer
“Reasoning is the biggest paradigm shift in AI architecture since the transformer.”
Sarah Guo Oct 31, 2025 ▶ 10:38 No Priors Ep. 138 | The Best of 2025 (So Far) with Sarah Guo and Elad Gil
Sep 3, 2026
Insight
Haas: AI chip startups need supply chain access, not just great design
“So long as the transformer is the unit of energy relative to how you generate AI training and AI inference by design, it is a, it is, it's very compute intensive, it's very memory intensive. So if you think about that, that's going to drive a lot of demand on …”
Rene Haas Sep 3, 2026 ▶ 12:59 Redefining Chip Architecture with Arm CEO Rene Haas
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.