PyTorch, every mention
78 scenes, the whole family · ← back to PyTorch
tap a year for its mentions
every year anyone Andrej Karpathy 27Shawn Wang 26George Hotz 25Chris Lattner 11comfyanonymous (Comfy) 6Alessio Fanelli 6Evan Feinberg 5Thomas Sohmers 4Batuhan Taskaya 4Ankur Goyal 4
Verbatim, from the transcripts: the passages where PyTorch comes up
🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
- ▶ 6:52 Anima Anandkumar So instead of writing in, like, PyTorch, it's like a PyTorch-like abstraction, but you can, like, kind of, you know, write it in Lean, and so it can be fully formalized in Lean. 2 times in the scene
- ▶ 1:17:47 Anima Anandkumar It's part of the PyTorch ecosystem.
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 25:17 Shawn Wang Even, but I'm surprised by the race condition one because, uh, I thought PyTorch was a graph that, like, guarantees that you at least, you know, execute things in the right order.
The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
- ▶ 14:59 Akshat Bubna It's not universally true that everyone else can autoscale, and we've gone deeper into it on the tech side, but we've incorporated GPU snapshotting to the product so we can actually, uh, take the GPU state, like your Torch compiler model,…
🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
- ▶ 59:07 Evan Feinberg I will say that the last line of, of PyTorch I've written is much further back in history than the most recent line of PyTorch that Sergey has committed. 2 times in the scene
- ▶ 1:28:44 Evan Feinberg When we were, you know, coding and PyTorch .17 building, you know, the first, you know, when Sergey was scaling transformers on PyTorch, the first one to do that, and we were scaling graph neural nets in PyTorch, one of the first ones to… 3 times in the scene
AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
- ▶ 30:56 Mikhail Parakhin Uh, you know, like you can grab, uh, XGBoost and you can grab some, some PyTorch module and then grab some, you know, Grap and other tools and combine them.
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 11:13 Kyle Kranen And then, and then we built, you know, like when deep learning was getting big, we built, we built Torch and, and,
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- ▶ 1:36:59 Shawn Wang I, I don't know if this is something that affects your analysis at all, because I don't have any appreciation for the, or the, the sizes that we're talking about here, that JAX is helping TPUs win, or JAX is winning relatively to PyTorch,…
[NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
- ▶ 16:31 unnamed speaker Uh, maybe, uh, and, you know, most people are familiar with PyTor. 2 times in the scene
[State of Context Engineering] Agentic RAG, Context Rot, MCP, Subagents — Nina Lopatina, Contextual
- ▶ 18:14 unnamed speaker Um, so it's like a nice, like, sort of PyTorch-like
One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
- ▶ 1:24:02 Shawn Wang LF has many other funds and, and, and organizations, including data, data and AI foundation, as well as like dedicated like PyTorch and all the other ones. 3 times in the scene
World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
- ▶ 42:02 Pim de Witte Super rudimentary, um, PyTorch-based physics engine, which I would not recommend writing a physics engine in PyTorch for obvious reasons, but I wanted to be able to, um, because it's, it's differential, so you can, uh, you can generate,… 2 times in the scene
After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
- ▶ 12:14 Justin Johnson But the way we talk about neural networks is still as if they are a monolithic thing that could be coded like in one GPU in PyTorch.
- ▶ 19:21 Justin Johnson Um, this was all pre Pytorch.
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 5:49 Quentin Anthony I'm actually, uh, you can get pretty far by using the rock and backend, um, on a lot of higher level software, like torch, for example. 2 times in the scene
Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
- ▶ 19:28 Kyle Corbitt If you were to show that exact same equation, just, like, code, not, maybe not PyTorch code, because that you also have to, like, understand, but if you just, like, did the naive implementation in, like, Python, and, like, showed someone,…
Context Engineering for Agents - Lance Martin, LangChain
- ▶ 34:19 Shawn Wang Uh, a lot of AI engineers don't even need to use PyTorch because you can just prompt and do typical software engineering.
A Technical History of Generative Media
- ▶ 12:52 Batuhan Taskaya You know, you go from like 10 seconds with PyTorch. 4 times in the scene
⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
- ▶ 20:02 Thomas Sohmers The approaches that people have taken before have been just saying, well, we're going to build a PyTorch backend. 4 times in the scene
⚡️The Future of Notebooks - with Akshay Agrawal of Marimo
- ▶ 21:26 Akshay Agrawal So you try to import torch.
Information Theory for Language Models: Jack Morris
- ▶ 9:49 Jack Morris And it's not like they're learning how to do like multi-node distributed FSTP training, like with whatever deep speed, you have to learn that from the internet and from other people.
- ▶ 10:25 Shawn Wang For grad students who are looking to get up to speed on that, I will recommend the GPU mode discord, where basically the PyTorch team is hanging out in there, just waiting to help you.
The Shape of Compute (Chris Lattner of Modular)
- ▶ 11:49 Chris Lattner You can run arbitrary PyTorch models. 3 times in the scene
- ▶ 35:44 Chris Lattner And so things like PyTorch came on the scene and I think PyTorch gets all credit for democratizing model training, right? 3 times in the scene
- ▶ 41:06 Chris Lattner Well, again, you get back into hacking the internals of VLM and PyTorch isn't really designed for KV cache optimizations and all the modern transformer features and things like this. 2 times in the scene
- ▶ 49:56 Chris Lattner And so that is a huge moment that set the stage for PyTorch to be open source and for the research to be open and for all of these things, because they decided the value system was AI go faster. 2 times in the scene
- ▶ 1:01:39 Chris Lattner You've got PyTorch if you want to train a model, but nothing is set up to do this.
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
AI Engineering for Art - with comfyanonymous
- ▶ 2:45 comfyanonymous (Comfy) So basically October, 20, 22, just, uh, like I hadn't written a line of PyTorch before that.
- ▶ 33:13 comfyanonymous (Comfy) And yeah, the problem with PyTorch is it's, uh, it's high levels. 5 times in the scene
2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
- ▶ 39:36 Sarah Chieng Um, okay, so the Cerebris Graph Compiler, um, the CGC right here integrates with machine learning frameworks, such as, such as TensorFlow and PyTorch.
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 2:14 Shawn Wang You ran the PyTorch team at Meta for a number of years, and we previously had Sumith Chintala on, and I think we were just all very interested in, like, the, the history of Gen.E.I. 8 times in the scene
- ▶ 6:59 Shawn Wang When you and I chatted about like the origins of fireworks, it was originally envisioned more as a PyTorch platform, and then later became much more focused on generative AI. 9 times in the scene
- ▶ 9:50 Alessio Fanelli What was the transition from you start focus on PyTorch and like people want to understand the framework, get it live. 2 times in the scene
- ▶ 14:59 Lin Qiao Because we wrote PyTorch code, we know we basically have a special PyTorch build for that, uh, together with a custom kernel we wrote. 2 times in the scene
- ▶ 29:08 Shawn Wang I understand that's also why PyTorch won.
[Paper Club] Upcycling Large Language Models into Mixture of Experts
Production AI Engineering starts with Evals
- ▶ 1:23:17 Ankur Goyal In particular, in DSPY's case, code that looks a lot like PyTorch code. 4 times in the scene
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
- ▶ 2:03 Andrej Karpathy And then, um, you, you've all worked with PyTorch, of course, right? 5 times in the scene
- ▶ 4:06 Andrej Karpathy So, uh, let's think about, like, what is PyTorch offering you, really? 6 times in the scene
- ▶ 5:39 Andrej Karpathy So for example, layer norm here is like a PyTorch layer, and, uh, we'd like to basically port this over to C. 6 times in the scene
- ▶ 9:05 Andrej Karpathy And then we can verify that PyTorch code is identical to the C code, and everything is great, and we're just running in C. 2 times in the scene
- ▶ 18:14 Andrej Karpathy There's no need for Python, no need for PyTorch. 6 times in the scene
- ▶ 21:29 Andrej Karpathy Like, if PyTorch is, especially when push compiles, it's a bit like TCC for software, it's a compiler. 2 times in the scene
Answer.ai & AI Magic with Jeremy Howard
- ▶ 39:31 Jeremy Howard Um, and then, you know, FSTP is this very complicated library in PyTorch, which not particularly well documented.
- ▶ 44:37 Jeremy Howard Also, the PyTorch team have created this Torch AO project on quantization, um, and so there's a, yeah, big overlap now between, kind of, the fast AI and answer AI and CUDA mode communities of people, um,
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 1:10:55 unnamed speaker You can set the torch seed and the inference seed.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 40:14 unnamed speaker So if I, like, Sling, like, SlingPyTorch, okay, great.
- ▶ 49:44 Yi Tay How much I think that people inside Google don't care about what people think outside Google, uh, like, I kind of feel like, okay, we were a bit like, like, I think, I don't think we considered, uh, uh, uh, I mean, not like, forever not… 2 times in the scene
- ▶ 1:11:57 Yi Tay A little bit of like, if you, if you propose like some very complicated thing for like, for like, easy in PyTorch.
How to train a Million Context LLM — with Mark Huang of Gradient.ai
- ▶ 26:32 Mark Huang The other PyTorch implementations outside of ECContext, they, um, they just didn't really work. 3 times in the scene
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 24:12 unnamed speaker So, things like, um, the libraries that we're using, um, JAX, PyTorch, TensorFlow, amongst others, um, there's this idea of distributed training, which means that, um, can we use multiple GPUs to train our models so that we are able to…
Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
- ▶ 0:27 unnamed speaker I was actually listening to your interview on the Gradient podcast, so if people want to know more about, like, the history of Smith, uh, the history of PyTorch, uh, they can go to that podcast. 3 times in the scene
- ▶ 4:26 unnamed speaker Yeah, we kind of touched on PyTorch in a lot of episodes. 6 times in the scene
- ▶ 10:49 unnamed speaker PyTorch has been such a Switzerland versus just making Meta hardware go bird? 11 times in the scene
- ▶ 19:47 unnamed speaker Who are some of the other FAIR, PyTorch alumni that are building cool companies? 4 times in the scene
- ▶ 29:59 unnamed speaker Like, what's the most interesting usage of PyTorch that you're seeing, maybe, outside of this little bubble? 5 times in the scene
- ▶ 1:01:43 Soumith Chintala Um, they fund a huge amount of the PyTorch, uh, development.
- ▶ 1:29:21 Soumith Chintala Yeah, I mean, first, like, ah, the very interesting reason I invested in Osmo is because Alex Wilsko, the founder of Osmo, also was, like, a, um, before PyTorch, there was Torch, and Alex Wilsko actually worked on Torch.
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 40:34 Ben Firshman So they want to be able to use this stuff without having to, like, figure out all the internals of the models and, you know, like, touch PyTorch and whatever.
- ▶ 1:16:37 Ben Firshman You don't need to be like, you know, the metaphor here is that you don't need to be digging down into like, uh, this sort of PyTorch level if you don't want to in the same way as a software engineer in the nineties.
The AI-First Graphics Editor - with Suhail Doshi of Playground AI
- ▶ 1:03:34 Suhail Doshi We don't, yeah, we don't even, no, we don't, I think we use very basic tools, like, you know, Slurm for scheduling, and just normal PyTorch, PyTorch Lightning, that kind of thing.
The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
- ▶ 1:08:03 Shawn Wang Yeah, I met Lynn, so she was entirely the, uh, she was like, with Sumith, she was like the co-manager of PyTorch for five years.
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
- ▶ 11:15 Alessio Fanelli The goal of Google is, like, make TPUs go fast with TensorFlow, but then you also had a post about PyTorch kind of stealing the, the thunder, so to speak. 4 times in the scene
- ▶ 11:44 Dylan Patel TPUs through PyTorch XLA is amazing, but it's, it's not bad, right?
- ▶ 58:57 Dylan Patel Yeah,
The End of Finetuning — with Jeremy Howard of Fast.ai
- ▶ 28:59 unnamed speaker Uh, so just to reintroduce Fast.ai for, uh, people who may not have, uh, dived into it much, um, there is the, the courses that you do, there is the library that is, uh, that is, um, very, uh, well loved, and I kind of think it, think of…
- ▶ 1:19:46 Jeremy Howard Like, so there was actually an interesting blog post that came out just today from the PyTorch team, uh, where some of them have created this, like, uh, three-D matrix product visualization thing.
RWKV: Reinventing RNNs for the Transformer Era
- ▶ 1:39:42 Eugene Cheah Like he re-implemented the backprop and all that, and we're just going to use Torch for that. 2 times in the scene
FlashAttention-2: Making Transformers 800% faster AND exact
- ▶ 11:17 Tri Dao So the PyTorch folks have been working on this as well. 2 times in the scene
- ▶ 39:40 Tri Dao You know, for example, they can, um, you know, if you, if you can run PyTorch on, on this stuff, like, you know, lots of people will be, will be using it, but, uh, um, you know, supporting all the, all the operations in PyTorch will take,…
Ep 18: Petaflops to the People — with George Hotz of tinycorp
- ▶ 11:56 George Hotz If you go read PyTorch, uh, you know, PyTorch I think is actually pretty good code. 16 times in the scene
- ▶ 24:32 George Hotz I wasn't complaining about PyTorch not compiling. 9 times in the scene