GPT-3, every mention
96 scenes, the whole family · ← back to GPT-3
tap a year for its mentions
every year anyone Shawn Wang 16Ankur Goyal 7Varun Mohan 6Suhail Doshi 6Jerry Liu 5Yasser Elsaid 4Stanislas Polu 4Barak Lenz 4Will Bryk 3Marc Andreessen 3
Verbatim, from the transcripts: the passages where GPT-3 comes up
Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
- ▶ 4:16 Joon Sung Park So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year when we were about to get GPT-III to be available.
⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
- ▶ 1:24 Ahmad Awais And I think Greg Brokman and Sam Altman ended up giving me access to GPT-III early.
🔬How GPT‑5 derived new results in theoretical physics and quantum gravity — Alex Lupsasca, OpenAI
- ▶ 4:22 Alex Lupsasca And I, you know, I remember thinking, well, okay, GPT-III could write email. 2 times in the scene
⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
- ▶ 0:22 Yasser Elsaid So, DaVinci was the model.
- ▶ 1:43 Yasser Elsaid So while I was working on the other projects, I, of course, AI started, like, the very first days of AI starting to get into, you know, like, mainstream, I saw people talking about GPT-III, and it was, 3 times in the scene
🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
- ▶ 1:14:50 Brandon Anderson GPT-II, GPT-III, GPT-III, you know, GPT-I, two, and three, like, there was a clear progression there. 3 times in the scene
Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
- ▶ 5:18 Marc Andreessen Um, so the, there was like a year where like the only way for a normal person to use GPT-III was in an AI dungeon. 3 times in the scene
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- ▶ 8:40 Doug O'Laughlin Moore's Law was, uh, Moore's Law was ending, and everything would change, and so, coming in with that, like, thesis at the top level, just, like, made me want to attack every little assumption, and something that really changed as well, I,…
- ▶ 2:04:05 Shawn Wang But after the GPC-III essay. 2 times in the scene
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 36:56 Shawn Wang What's the rumored size of GPT three pro and to be clear, not confirmed for any official source, just rumors, but rumors do fly around rumors.
[State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
- ▶ 1:12 Jack Merullo I was working on grounding in language models, which is basically the idea that you need more than text data to represent meaning in the world, and it was basically right away as I started, like, at some point, you know, GPT-III had come…
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 24:30 Ashvin Nair And then I think, and then GPT-III happened.
⚡️ Building the AI Hardware Engineer with Matthias Wagner, Co-founder of Flux
- ▶ 4:11 unnamed speaker Did you just plug in GPT-III and like everything took off or was there a more complicated engineering story behind that?
Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
- ▶ 1:57 Barak Lenz We were, again, and also training our model, GPT-III came out, and then we were very heavy users, and we decided to train our own GPT-III. 3 times in the scene
- ▶ 6:31 Barak Lenz Before anyone was doing models, training a GPT three scale model, but the Lama architecture made it more approachable because they made a lot of the adjustment that they made the optimization work and the activations not explode.
Taste is your Moat (Dylan Field of Figma)
- ▶ 6:50 Dylan Field But I think GPT-III was probably the first time that I was like, wow, 2 times in the scene
Greg Brockman on OpenAI's Road to AGI
- ▶ 18:14 Greg Brockman A GPT-III or GPT-IV, not a GPT-V for sure, right?
- ▶ 19:37 unnamed speaker If I think about three, four, five as the major versions, I think three is very text-based, kind of like RLHF really getting started. 2 times in the scene
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 8:11 Varun Mohan You know, GPT-III is great, but the precision recall, you wouldn't trust someone's life with that, right? 3 times in the scene
- ▶ 27:58 Shawn Wang GPT three is 175.
- ▶ 46:59 Varun Mohan Here, which is that let's go with GPT-III, which is like a hundred and seventy billion parameters.
- ▶ 48:44 Varun Mohan But beyond that, I think Eleuther is a pretty special group, especially it's been now probably more than a year and a half since they released like GPTJ, which was like back when open source GPT-III Curie, which was comparable. 2 times in the scene
- ▶ 3:04:12 Scott Wu And, um, you know, in GPT three, for example, it's, if you were to go and ask GPT three to do something, you know, it could probably get through a few words or so, 2 times in the scene
AI is Eating Search
- ▶ 25:37 Robert McCloy It was just a wrapper around, ah, like, GPT-III.
Information Theory for Language Models: Jack Morris
- ▶ 3:06 Jack Morris Like around when I guess GPT three, one hundred and seventy five billion had been released, but not instruct GPT.
- ▶ 20:40 Shawn Wang You know, let's say GPT-III can store two Wikipedias.
- ▶ 23:02 Jack Morris Yeah, yeah, that can make sense, that can make sense, because I, I guess what you, you say from, you know, if you want to do apples to apples comparisons, GPT-III can store two Wikipedias, is that right? 2 times in the scene
- ▶ 43:05 Shawn Wang You need a GPT-III and GPT-IV in order to then get O-one.
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 27:58 Vibhu (Viboo) GPT-III hadn't really launched.
SF Compute: Commoditizing Compute
- ▶ 1:06:09 Michael Swix (Swyx) You were, you were so excited in the early GPT three days, three days.
Building Manus AI (first ever Manus Meetup)
- ▶ 18:15 unnamed speaker Because all you guys know, uh, GPT-III is far, uh, is far earlier. 2 times in the scene
Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
- ▶ 5:01 Misha Laskin So it's really that, I don't know if you've been able, you would have been able to do the same stuff with an earlier model, like with a GPT-III or a GPT-II, but GPT-IV had this kind of base of intelligence that you could actually go from…
The AI Architect: Bret Taylor
- ▶ 1:34:34 Bret Taylor And you just saw like every, and I think with these reasoning models, just how we're using sort of inference time compute and.
The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
- ▶ 2:27 Alessio Fanelli UX prototypes with GPT-Tree as well, and kind of like maybe how that is informed the way you build products.
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 6:33 William Beauchamp I think OpenAI had said, they hadn't released GPT-III yet, but they'd said GPT-III is so powerful, we can't release it to the world or something.
OpenAI o1 isn’t a chat model (and that’s the point)
- ▶ 2:24 Dan McAteer So, like, when ChatGPT first came out, right, I think it was GPT-III.
Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
- ▶ 1:17 Will Bryk So you might have to remember the times of summer, 21 and, uh, GBD three had come out. 2 times in the scene
- ▶ 20:09 Will Bryk Also, like, Google was built in a time where, like, in, you know, in 1988, where we didn't have LMs, we didn't have embeddings, and so they never thought to build those things, and so now they have this, like, gigantic system that is built…
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 39:55 Shawn Wang So let's say, 20, 2020, 20, three was let's scale big models territory because we had GPT three in 20, 20. 2 times in the scene
- ▶ 1:14:08 Shawn Wang And people really want GPT three base.
- ▶ 1:43:20 Alessio Fanelli Why Google failed to make GPT-III.
[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
- ▶ 36:46 Sarah Chieng So the paper actually includes a lot of research on this as well, um, looking at GPT-III, creating a sparse representation of GPT-III and showing that, you know,
- ▶ 40:15 Sarah Chieng And so activation tensors and models, um, like GPT-III have three dimensions, batch, sequence, and hidden feature.
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 34:50 Shawn Wang You know, like, I think that, that is the, the core insight of the GPTs, the GPC one, two, three, that was
Agents @ Work: Dust.tt — with Stanislas Polu
- ▶ 10:05 Stanislas Polu Most of the compute was, uh, going to a product called Nest, which was basically GPT-free. 3 times in the scene
- ▶ 16:53 Stanislas Polu At the time, the split was pretty surprising because they had been trying GPT-III.
Building the Silicon Brain - Drew Houston of Dropbox
- ▶ 3:21 Drew Houston I mean, actually, I started trying some of, like, this, what we would call, like, very small LLMs before, kind of, the GPT class models, and it was, like, super hard to get that, those things working, so, like, these 500 parameter models… 2 times in the scene
Production AI Engineering starts with Evals
- ▶ 20:10 Ankur Goyal And then I started playing with GPT-III, and that just totally blew my mind. 3 times in the scene
- ▶ 32:49 Ankur Goyal What is incredibly empowering about these, I would just maybe say that the quality that transformers bring to the table, and even BERT does this, but you know, GPT three and then four, like very emphatically do it, is that software…
- ▶ 1:34:47 Ankur Goyal I mean, it was ridiculous to GPT-III, too. 3 times in the scene
Building AGI in Real Time (OpenAI Dev Day 2024)
- ▶ 1:34:22 Sam Altman And also, when we made the first GPT-III, um, if you asked me for the techniques that would have worked for us to be able to
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- ▶ 7:08 Shunyu Yao So I think one day after, you know, I've seen all the GPT-III stuff, I just think, think about, you know, how, how can I solve the game?
- ▶ 28:31 Shawn Wang How language models, at least in the sort of GPC three era was that they're, they over optimized to some sets of tokens in sequence.
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
- ▶ 3:17 Sander Schulhoff And I started using GPT-III prompting it to do the translation, and that was, I think, my first intro to prompting, and I just started doing a bunch of reading about prompting, and I had an English class project where we had to write a…
- ▶ 17:24 Sander Schulhoff I think it might have worked on older ones, like GPT-III.
Building AGI with OpenAI's Structured Outputs API
- ▶ 8:29 Shawn Wang Not even everyone had access to the GPT-III model.
Is finetuning GPT4o worth it?
- ▶ 4:01 Alistair Pullen So we, we tried a small, um, some small projects in various different areas, but then Sam talked to me about GPT-III.
- ▶ 4:30 Alistair Pullen So I was like, okay, this GPT-III thing sounds interesting. 2 times in the scene
Answer.ai & AI Magic with Jeremy Howard
- ▶ 32:25 Swix (Shawn) One of his arguments was actually people kind of over-index on the decoder-only GPT-III type paradigm.
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 19:09 Alessio Fanelli They're kind of like speed running GPT one, GPT two, GPT three in open source.
- ▶ 1:10:58 unnamed speaker Because GPT-III took about a year to drop order magnitude, but now GPT-IV, it's really crazy.
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 10:40 Thomas Scialom And for that, they figured that model size is what matters, so GPT-free was way too big compared to the actual number of training tokens, because they did a mistake not adapting the scheduler.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 3:21 Yi Tay Like, WMT, like, machine translation, and like, complexity, and stuff like that, it's not really about, you know, there wasn't, like, I think, feel short learning, and feel short in context learning came only about, like, you know, when…
- ▶ 1:37:52 unnamed speaker I think OPT and, uh, GPT, uh, 3 times in the scene
This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
- ▶ 1:26:55 Joscha Bach And, ah, while arguably this agent paradigm of the chatbot is, ah, what made, ah, chat GPT so successful and moved it away from GPT-free to something that people started to use in their everyday work much more. 2 times in the scene
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
- ▶ 1:38 Jason Liu Yeah, I mean, this was, like, I mean, this was, like, GPT 2 times in the scene
- ▶ 5:05 unnamed speaker Uh, one more system that I'm interested in finding out more about is your similarity search system using Clip, uh, and GPT-T-E-Embedding in FICE, um, where you, you said 50, over fifty million dollars in annual revenue. 3 times in the scene
- ▶ 8:53 Jason Liu and so I kind of took, sort of, half of it as medical leave, the other half I became more of, like, a tech lead, just, like, making sure the systems were, like, lights were on, and then when I went to, uh, New York, I spent some time…
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
- ▶ 20:32 Jungwon Byun And then after GPT-III came out, I think by that time we, um, kind of realized that originally we were trying to help people convert their beliefs into probability distributions. 2 times in the scene
- ▶ 26:05 unnamed speaker Uh, yeah, I think, I think you were, you're about to start us on, like, GPT-III and how, like, that changed things for you. 4 times in the scene
- ▶ 37:26 Jungwon Byun So in 20, 21, we had this thing called composite tasks where you could use GPT three to brainstorm a bunch of research questions and then take each research question and decompose those further into sub questions.
- ▶ 44:43 Andreas Stuhlmüller Let's say the first model is Lama, and let's say the second model is GPG-C.
Why Google failed to make GPT-3 -- with David Luan of Adept
- ▶ 9:05 David Luan You know, every day we were scaling up GPT-III, I would wake up and just be stressed.
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 14:36 unnamed speaker Sampling or your autoregressive sampling of tokens to form your target sequence, which is what we've seen in T five and the decoder models, which I think all of us are familiar with things like GPT two, GP three, three, they are all there.
- ▶ 16:38 unnamed speaker So essentially, we'll just talk about the last point over here, where we are seeing that large language models, uh, in particular things like GPT-III, are able to perform your downstream task without specific fine-tuning. 3 times in the scene
- ▶ 28:03 unnamed speaker So that's how GPT-III will output in sentences.
Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
- ▶ 43:20 Shawn Wang So I could say like, because I, you know, the number of times I hit the GPT, GPT, GPT three, uh, API at the time, uh, was, was going to be subject to the rate limit.
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
- ▶ 25:52 Nathan Lambert Either GPT-II or GPT-III, I always get the exact views. 2 times in the scene
- ▶ 30:28 unnamed speaker Uh, pushed Llama One from, like, let's say a GPT-III-ish model to a GPT-III model in, in pure open source with not a lot of resources.
The AI-First Graphics Editor - with Suhail Doshi of Playground AI
- ▶ 11:46 Suhail Doshi Uh, address bar prediction with like GPT three and three, 3.5 and that kind of thing. 2 times in the scene
- ▶ 18:25 Suhail Doshi Like even GPT three, even when GPT three came out, it was exciting, but it was like, what are you going to use this for? 3 times in the scene
- ▶ 50:53 Suhail Doshi So I think it goes back to, you know, how we started the company, which was kind of looking at GPT three's playground that the reason why we're named playground is, is a homage to that actually.
The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
- ▶ 59:36 Shawn Wang Um, and actually, uh, there's a pretty good cadence from GPT-II, three, and four, uh, that you can, if you project out, um, so four, uh, four is, uh, based on George Hotz's, um, uh, concept of, like, uh, 20 petaflots being, like, a human,…
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
- ▶ 31:57 Shawn Wang Because we know three, but we don't know 3.5.
- ▶ 53:03 Dylan Patel Um, to build four, you need to be able to build a 3.5, build 3.5, you need to be able to have three, right?
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 8:43 Michael Royzen When it worked, it blew my mind, um, that instead of doing this few shot thing, like people were doing GPT-III at the time, which is all the rage, you could just ask a model a question, um, provide no extra context, and it would know what… 2 times in the scene
- ▶ 37:51 unnamed speaker Three to four. 3 times in the scene
Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue
- ▶ 3:08 Kanjun Qiu I started the archive with two friends, oh, with Josh, my co-founder, and a couple other folks in 2015, that's right, and GBD-III are housemates built, so.
- ▶ 7:39 Kanjun Qiu Not exactly this idea, but, uh, in late 2019, so I mentioned our housemates Tom Brown and Ben Mann, they're the first two authors on GPT-III. 2 times in the scene
- ▶ 1:01:39 Shawn Wang Like, this is, this is very special, it's, and a lot of people want to do that and fail, and you are one, like, you had the co-authors of GPT-III in your house.
RAG is a hack - with Jerry Liu of LlamaIndex
- ▶ 6:15 Jerry Liu Uh, this was just like exploring GPT-III, or it was October actually. 3 times in the scene
- ▶ 1:00:20 Jerry Liu Um, and I think the issue is if you also build your own models, like you're also just gonna have to keep up with like the rate of L and advances, like how, like basically the question is when GPT five and six and whatever, like anthropic…
- ▶ 1:09:52 Jerry Liu Because GPT-free had been out for a while, like, just the fact that, um, there was this engine that was capable, like, reasoning, and no one was really, like, tapping into it, um, and then the fact that, uh, you know, I used to work in…
RWKV: Reinventing RNNs for the Transformer Era
Ep 18: Petaflops to the People — with George Hotz of tinycorp
- ▶ 50:54 George Hotz No, it's a little bigger than GPT-III, and they did an eight-way mixture of experts.