Lubershain: Grid equipment and gas plant costs have doubled or tripled
“Conductor is, like, twice as expensive. So, like, basic aluminum, steel-reinforced conductor. Transformers are two-plus times more expensive. Switchgear is two times more expensive. Gas power plants are two to three times more expensive.”
Haas: AI chip startups need supply chain access, not just great design
“So long as the transformer is the unit of energy relative to how you generate AI training and AI inference by design, it is a, it is, it's very compute intensive, it's very memory intensive. So if you think about that, that's going to drive a lot of demand on …”
Jeffrey: AI will not reduce code volume, just make code disposable
“There isn't going to be less code. The code just becomes less valuable throw away.”
Fourier neural operators scale quasi-linearly, avoiding transformers' quadratic complexity
“If we were to use transformers and we require a very high resolution, it would become untenable because of the quadratic complexity and all to all connections. On the other hand, if you did that with Fourier transforms, we have like quasi linear complexity and…”
Transformers will never scale to high-resolution 4D physics simulations
“So forget ever having a transformer for anything of this scale. All of the world's compute will not be enough. And first of all, they all have to be co-located to be able to ever do this. So that's why we need other architectures.”
The 'Original Sin' of Transformers Is Pairing Memory-Bound and Compute-Bound Layers
“And I would say the original sin of Transformers is that you've taken this extremely fundamentally memory bound layer and juxtaposed it right next to a compute bound layer. It is very difficult to have a single chip that is good at both compute operations and …”
Movva: Transformers scale well because they impose no human priors
“What transformers really did well is that they scaled. Transformers make no such human prior. Transformers just say, well, there's gonna be a pattern in the sequence of data, and if there is a pattern, I'm gonna find it.”
Angelopoulos: Data is the hardest part of model training as algorithms commoditize
“The data is really the hardest part of model training. Because you need to source it. It's so dirty. Nobody wants to do that shit. Nobody wants to hire all these people to generate data and then, you know, turn that into basically data plus GPUs equals model. …”
Krishnan: Nobody expected Transformers and BERT to unlock emergent reasoning or AGI
“At least at the time Transformers came out, BERT came out and all that, nobody really thought this was a this was a path to emergent reasoning capabilities, a path to AGI itself.”
Katti: Turbine and transformer manufacturers face multi-year lead times to expand capacity
“Those industries have historically have not added much capacity for the last decade or so ago, and they've suddenly experienced a demand shock. And it takes years before you can add capacity to produce more turbines and transformers.”
Catanzaro: Combining SSMs and transformers produces smarter AI models than either alone
“Using both of these together was actually better than using either one on their own. And that is independent of the speed benefit. That is just the model is smarter.”
Lubershane: Grid equipment costs 2-3x more and takes 3-5 years to deliver
“Anything you want to order today is probably going to take three plus years in the case of gas turbines, probably more like five years to be able to get your hands on. Anything you order today, whether it's a turbine or a transformer or just the, you know, alu…”
Hinton: Post-Transformer AI Progress Stems Mainly from Hardware and Engineering
“We've also seen new ideas, but mainly since Transformers, it's been much better hardware, many more resources better engineering, and many more talented people.”
Ethan He: Training transformers directly on raw image pixels is impossible
“If you're trying, if you can technically, theoretically train image or video models on pure pixels, but the problem is that the, it's a lot of tokens. So like one image, like it's a thousand by a thousand is like one million tokens, one million pixels. It's im…”
Ethan He: Training models directly on MP4 tokens is extremely difficult
“So people actually have tried that, but the main challenge is the latent space for the MP four tokens are not, we're not very comprehensible for the models. It's extremely hard to train on that.”
Ries: Every Transformer Paper Coauthor Left Google to Commercialize Elsewhere
“And if you look at the coauthors of that paper has a ton of coauthors, not a single one commercialized the transformer at Google. They all, every single one had to leave and do it elsewhere.”
Chaubard: Transformers cannot sort lists longer than layer count in one pass
“In a one-shot basis. It's like literally that we know a theoretical lower bound that for comparison sort, you can't do better than n log n steps. And if I have a list that's 31 characters or elements long, and my transformer is 30, I run out of steps to do com…”
Gupta: Recursive architectures achieve compute depth without parameter depth
“Recursion advantage now gives you a bunch of advantages over transformers where rather than having, you know, 500 or a thousand or a million or whatever transformer layers and having tons and tons of parameters, you get compute depth basically without this par…”
Pichai: BERT and MUM drove Google Search's largest quality leaps
“BERT and MUM, people underestimate how much, because we measure search quality so religiously, some of the biggest jumps in search quality in that period where search went ahead of everyone else was because of BERT and MUM. We built transformers and used it im…”
Ludwig: Transformer advances prompted Applied Intuition's physical AI expansion
“One was transformers. That were originally successful for large language models, like you heard about Anthropic and OpenAI those started to have an impact on self-driving technology and robotics, and we thought that was really interesting, and then on top of t…”
Andreessen: AlexNet and transformers were the true inflection points of modern AI
“I think the real story is it was the Alex net basically breakthrough in like 2013. That was the real knee in the curve. And then it was obviously the transformer breakthrough in 17.”
Manning: Transformer internal weights can act as joint representations for world models
“I'm not actually convinced that's right, because although the token production is this autoregressive process that's heading, you know, left to right, I guess don't have to be left or right, but anyway, in sequence of tokens, we could have right to left Arabic…”
Dehghani initially thought Google's Transformer architecture was random and would die
“And I was like, I don't know if I want to go with this team. It's just like, they're doing something random. Like who, like everybody's doing LST. I'm like, why should I go and work with like a group of people who are working on this like random architecture, …”
Unified Transformer architectures simplified the training of natively multimodal AI models
“Even if this is not, like, the only architecture that would be, like, in a multi-model, but it made it really simple to train these models, like, natively, because you have, like, a single architecture and you can have all the modalities in, during training.”
Transformers compute precise Bayesian posteriors down to 10^-3 bits accuracy
“We trained these models and we found that the transformer got the precise Bayesian posterior down to 10 to the power minus three bits accuracy. It was matching the distribution perfectly. So it is actually doing Bayesian in the mathematical sense, given a task…”
Transformers perform all Bayesian tasks, Mamba does most, and MLPs fail
“Transformer does everything. Mamba does most of it. LSTMs do only partially, and MLPs fail completely.”
Kann: Lead times for gas turbines, transformers, and switchgear are 3 to 7 years
“The five to seven years is the timeline to get new gas turbines, if you're ordering them, and it's close to the timeline to get new transformers and other, and switchgear and stuff like that. Like, we're all in the, like, three to five or maybe seven year time…”
Elder: Off-grid setups partially bypass generation bottlenecks with modular equipment
“The off-grid option in some ways maybe shortcuts the actual supply chain bottlenecks on the generation equipment side, at least to some extent relative to the other options. But I agree with you, there's still a bunch of other pieces of equipment, transformers…”
Kamath: State space models are not a big leap over transformers
“Our strong belief is like state space models are not like that big a leap over transformers.”
Baglino: Higher frequency linearly reduces transformer size and power electronics costs
“The best way to make power electronics Systems that involve isolation more affordable is go up in frequency, because to get isolation, you basically need to use a transformer of some type, and transformers become smaller as you go up in frequency. It's just a,…”
Baglino: Utility-scale solar plant transformers fail at about 1.4% per year
“Transformers are not really designed to run at their rated power, you know, as long as they do in these Desert power plants, you know, where they're, first of all, very hot, because they're sitting in the sun, and second of all, because they're running at name…”
Dean: Transformers delivered 10x to 100x compute efficiency over LSTMs
“Transformers similarly gave you a 10 X to a hundred X improvement in, you know compute cost to a given quality level versus say LSTMs at the time.”
Barrett: Transformers Are Not the End State of AI
“Transformers are not the end state of AI, right? AI is Embarrassingly stupid compared to many of the examples we are carrying around in our heads, both in terms of what it can achieve, but also how much power it uses to do it.”
Pineau: AI is far away from a breakthrough transformer moment for reasoning
“The transformer moment for reasoning and choosing action and being able to plan at different levels of granularity. We're still far away from doing that.”
Yi Tay: Today's Models Likely Couldn't Invent the Transformer from Pre-2015 Data
“Even today's models, they might not even be able to invent the transformer. Like, if you freeze the time at a certain time, and even you bring the time, I mean, the model is a transformer, so I just say there's no, assuming there's no leakage.”
Yi Tay: The Architecture That Achieves AGI Will Still Be a Transformer
“It will be a transformer, I think. Like people, it depends on what you call it, but I think unless the paradigm shifts completely, which is, I mean, as a scientist, you cannot like completely say no to like that, this would never happen. But my feeling is that…”
Yi Tay: Gradient descent learning paradigm is AI's bottleneck, not architecture
“It's not architecture itself. That's, that there's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the to…”
Yi Tay: The AI Industry Is Stuck in a Transformer Local Minimum
“So now we are like in this local minima of like transformers, everything, everything, right? Maybe it's not easy to like get totally out Of this, because also a lot of people's investment optimization have been done. So the things that play well needs to play …”
Nat Bullard: Transformer prices keep rising despite domestic manufacturing push
“If anything, they're still ticking slightly upward. And as much as there's energy and attention and noise being paid to, we need to build stuff here, we're not.”
Kann: Gas turbine manufacturing is more oligopolistic than electrical transformers
“Gas turbine manufacturing is even more capital intensive. It's currently more of an oligopoly than transformers are.”
Izmailov: Transformer architectures will prove highly suboptimal for certain computational tasks
“At least for some tasks, I'm pretty confident that the transformers will be highly suboptimal.”
DeepMind uses transformers to model spatial interactions in weather, not time
“We actually use transformers not to model the spatial, the interactions in weather over time, like the sequence of text, but in space.”
Schulhoff: All Transformer-Based Chatbots Are Vulnerable to Adversarial Attacks
“And because all I guess for the most part, all currently deployed chatbots are based on transformers or transformer adjacent technologies. They're all vulnerable to Prompt injection, jailbreaking, forms of adversarial attacks.”
Bourgeau: Google and DeepMind are actively researching post-Transformer architectures
“I believe so. There's groups doing research on the model architecture side, for sure, within Google and within DeepMind”
Rao: Transformers succeeded because they maximized GPU hardware efficiency
“Transformers are really they're a big innovation because they made the constructs of the GPU work extremely well.”
Kaiser: Reasoning models are the second major milestone after Transformers
“One point was, of course, the Transformers when it started, but the other point was reasoning models.”
Johnson: Transformers are natively models of sets, not sequences
“Transformers are actually not a model of sequences. A transformer is natively a model of sets.”
Guo: Reasoning is the biggest AI architecture shift since the transformer
“Reasoning is the biggest paradigm shift in AI architecture since the transformer.”
Rathi: Transformer shortages are placing housing projects on 18-month holds
“In my reporting, like I've come across many, many projects like housing projects that are completely on hold for 18 months because they're missing this one object that they need before they can Power the area.”
Kann: Upstart hardware companies will help bridge the transformer supply gap
“I think there's going to be a wave of upstart companies, whether they're doing traditional oil-filled transformers or solid-state stuff like Heron is doing. I think they're going to fill in a lot of the supply gap in the next few years”
Modern AI Revolution Is Predicated on Google's 2017 Transformer Invention
“The entire AI revolution that we are in right now is predicated by the invention of the transformer out of the Google brain team in 2017.”
2017 Began a 5-Year Period of Google Failing on Transformers
“It is fair to say that 2017 begins the five-year period of Google not sufficiently seizing the opportunity that they had created.”
OpenAI Introduced Pre-trained Transformers with GPT-1 in June 2018
“So in June of 2018, OpenAI releases a paper describing how they have taken the transformer and developed a new approach of pre-training them on very large amounts of general text on the internet, and then fine tuning that general pre-training to specific use c…”
Douglas: Transformers successfully model any domain given sufficient data and compute
“I don't think that's true. I think we haven't yet really found anything that transformers haven't been able to model provided sufficient data and sufficient compute.”
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer
With a context length of, you know, 256,000 or more, they're not using a true transformer.
What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Bachman: Compute-optimal models on internet text don't need long context
“In general, most internet text has mostly short-term structure. There's just not that much value in capturing long-term structure, and so compute optimal models on internet text actually don't have that long context, and so, of course, you're perfectly fine us…”
Frosst: Google failed to quickly commercialize or scale the Transformer
“It wasn't then commercialized very quickly within Google. It wasn't scaled up very quickly within Google. A lot of that work had to be done elsewhere and years later.”
Morcos: The Transformer is just one of many equivalently good architectures
“And one of my like more controversial viewpoints, I think, is that I think the transformer is a great advance to be sure, but I think it's one of a very large Set of equivalently good architectures that we could have found. And there are many, many ways we cou…”
Google Invented and Published the Revolutionary Transformer Architecture in June 2017
“Google invented the Transformer and published the paper in June of 2017.”
Sohmers: Transformer inference is memory-bound with a 1:1 FLOP-to-byte ratio
“And on the other side of this chart, you have the case of a transformer where when you're actually, you know, doing attention, Or really just any case where you're doing, you're fundamentally doing matrix vector multiplication rather than matrix matrix multipl…”