GPT-2, every mention
49 scenes · ← back to GPT-2
tap a year for its mentions
every year anyone Shawn Wang 13Nathan Lambert 5Dylan Patel 4Ashvin Nair 4Andrej Karpathy 4Alessio Fanelli 4Andreas Stuhlmüller 3William Beauchamp 2Nicholas Carlini 2Jason Liu 2
Verbatim, from the transcripts: the passages where GPT-2 comes up
Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
- ▶ 4:29 Joon Sung Park So we already had GPT-II, and you could sense that there's this new class of models that was just becoming available in the market.
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
- ▶ 1:14:50 Brandon Anderson GPT-II, GPT-III, GPT-III, you know, GPT-I, two, and three, like, there was a clear progression there. 2 times in the scene
Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
- ▶ 5:06 Marc Andreessen And then, you know, and then open AI developed chat GPT or GPT two,
Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
- ▶ 3:17 Alessio Fanelli Have the thread models been updated a lot, or do you feel like you're still using the same thread models as GPT-II of like, you know, paperclip factory, blah, blah, blah, you know, but like, how much are you rising, you know, increasing…
- ▶ 9:58 Joel Becker You know, I think GPT-II can like sometimes do that task and, and, and sometimes not.
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 4:54 Ashvin Nair Yeah, like I would say that robotics is in kind of like the GPT-one to GPT-two area right now. 2 times in the scene
- ▶ 24:17 Ashvin Nair I think a lot of people didn't really think of GPT or GPT-II as something that was like super compelling probably. 2 times in the scene
Greg Brockman on OpenAI's Road to AGI
- ▶ 18:02 Greg Brockman The results to be felt like GPT one, maybe starting to be GPT two level, right?
⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
- ▶ 2:46 Stefano Ermon And so we published a paper at ICML last year that won the best paper award showing that for the first time, basically the discrete diffusion models were competitive on language generation without the regressive models up to the GPT-II
Information Theory for Language Models: Jack Morris
- ▶ 2:20 Jack Morris GPT two, GPT one from open AI were like interesting, but I think most people were into BERT at that time.
- ▶ 42:58 Shawn Wang If you gave the O-one harness on top of GPT-II, you would get nothing because GPT-II didn't know enough. 2 times in the scene
Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
- ▶ 9:35 Noam Brown I think it could have happened earlier, but if you try to do the reasoning paradigm on top of GPT-II, I don't think it would have gotten you almost anything.
- ▶ 32:28 Spooks (Swyx) We had David Luan on, who I think was VP Eng at the time of GPT-I and II.
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 35:21 Emmanuel Ameisen well you have all your curve detectors, but think about all of the concepts that like Claude or even GPT-II need to, to know, like just in terms of, it needs to know about like all of the different
The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
- ▶ 18:39 Nikunj Handa It's like the GPT-II of computer use or maybe GPT-I of computer use right now.
Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
- ▶ 5:01 Misha Laskin So it's really that, I don't know if you've been able, you would have been able to do the same stuff with an earlier model, like with a GPT-III or a GPT-II, but GPT-IV had this kind of base of intelligence that you could actually go from…
Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
- ▶ 8:58 Logan Kilpatrick I tweeted the other day that it's the, it feels like the GBT two era, the like order in which the timeframe in which model progress can actually happen.
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 6:41 William Beauchamp Was it GPT-II?
- ▶ 12:34 William Beauchamp Then it would prompt whatever, like I just used some external API for like BERT or GPT-II or like it was a very, very small thing.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 28:38 Shawn Wang Zero capabilities, and it's sudden emergence of GPT-IV. 2 times in the scene
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 34:50 Shawn Wang You know, like, I think that, that is the, the core insight of the GPTs, the GPC one, two, three, that was
Agents @ Work: Dust.tt — with Stanislas Polu
- ▶ 9:41 Alessio Fanelli How did the research evolve from, you know, the GPT-II and then getting closer to like…
Building AGI in Real Time (OpenAI Dev Day 2024)
- ▶ 1:48:56 Sam Altman Like, we are so early, this is like, you know, maybe it's the GPT-II scale moment, but
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- ▶ 2:46 Alessio Fanelli And that was GPT-II, was that at the time?
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
- ▶ 8:41 Andrej Karpathy We download a starter pack, which is really just a GPT-II weights in a single binary file. 2 times in the scene
- ▶ 17:52 Andrej Karpathy And where that leads us to is that we can actually turn QPD too, and we can actually reproduce it after all of that work.
- ▶ 22:41 Andrej Karpathy You just prompted them to write GPT-C.
Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
- ▶ 24:53 Nicholas Carlini I still would have said the same thing with like, after I had seen GPT two, I had written a couple of papers studying GPT two very carefully. 2 times in the scene
Is finetuning GPT4o worth it?
- ▶ 4:15 Alistair Pullen I'd actually heard of GPT-II. 2 times in the scene
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 19:09 Alessio Fanelli They're kind of like speed running GPT one, GPT two, GPT three in open source.
LLM Asia Paper Club Survey Round
- ▶ 40:01 unnamed speaker Um, on all sorts of models, uh, GPT-II Small is a favorite, just because it's very well understood by interpretability researchers, and you can see all sorts of features identified in GPT-II Small, um, and it's actually super easy to, to… 3 times in the scene
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
- ▶ 13:07 Andreas Stuhlmüller And then the initial, I think the first models that kind of make sense were TPT-II and TNLG and like the, the early, early, um, 3 times in the scene
- ▶ 26:58 Jungwon Byun It felt less like a level, GPT-III over GPT-II was like qualitative level shift.
Why Google failed to make GPT-3 -- with David Luan of Adept
- ▶ 7:10 Shawn Wang Um, but you mentioned GPT-II. 3 times in the scene
- ▶ 12:38 Shawn Wang Is there another, you know, sort of GPT-II story that, like, you know, you love to get out there, um, that I think is, you think it's like underappreciated for, like, the amount of work that people put into it? 4 times in the scene
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 14:36 unnamed speaker Sampling or your autoregressive sampling of tokens to form your target sequence, which is what we've seen in T five and the decoder models, which I think all of us are familiar with things like GPT two, GP three, three, they are all there.
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
- ▶ 25:52 Nathan Lambert Either GPT-II or GPT-III, I always get the exact views. 2 times in the scene
- ▶ 1:09:56 Nathan Lambert Talking about GPT-II still, which is such a kind of like odd model to focus on. 3 times in the scene
The AI-First Graphics Editor - with Suhail Doshi of Playground AI
- ▶ 18:21 Suhail Doshi It feels a little like graphics is in like this GPT two moment, right?
The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
- ▶ 59:36 Shawn Wang Um, and actually, uh, there's a pretty good cadence from GPT-II, three, and four, uh, that you can, if you project out, um, so four, uh, four is, uh, based on George Hotz's, um, uh, concept of, like, uh, 20 petaflots being, like, a human,…
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
- ▶ 10:28 Dylan Patel Like we are growing, you know, GPT-two to four, two to four is like twenty-twenty to twenty-twenty-two, right?
- ▶ 37:18 Dylan Patel If you look at the total flops, right, you know, parameters times tokens times six, right, it's like, it's like a tiny, tiny fraction of GPT-II, which came out just a few months later, which was like, okay, 3 times in the scene
The End of Finetuning — with Jeremy Howard of Fast.ai
- ▶ 1:19:59 Jeremy Howard Yeah, and they actually showed some nice, nice examples of, like, a GPT-II attention layer, and, like,
RWKV: Reinventing RNNs for the Transformer Era
- ▶ 47:01 unnamed speaker So, GPT-II was actually open source.
- ▶ 52:59 Eugene Cheah So, so, so the history, the history behind it, right, and is that, um, I, I think, I think, like, a few years back, once with GBT II, Transformers started to pick up Steam, and I guess the whole world is starting to think, let's just…