GPT-2, every mention

49 scenes · ← back to GPT-2

tap a year for its mentions
0020840152023202420252026episodesmentions
08152023202420252026episodes it came up in
001.37.52.5152023202420252026episodesmentions per episode

every year anyone Shawn Wang 13Nathan Lambert 5Dylan Patel 4Ashvin Nair 4Andrej Karpathy 4Alessio Fanelli 4Andreas Stuhlmüller 3William Beauchamp 2Nicholas Carlini 2Jason Liu 2

Verbatim, from the transcripts: the passages where GPT-2 comes up

loading…

Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI Aug 21, 2026 · 1 mention

  • ▶ 4:29 Joon Sung Park So we already had GPT-II, and you could sense that there's this new class of models that was just becoming available in the market.

The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI Jul 22, 2026 · 2 mentions

  • ▶ 1:32:47 Eiso Kant What I do want to call out is that people have been calling for the fear of misuse of these models since GPT-II. 2 times in the scene

🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu) Jul 21, 2026 · 1 mention

  • ▶ 1:15:53 Bo Wang I remember vividly we have GPT-II, we update the lectures, and then different language models, how the multi-thread GPU communication is used in training, large-scale neural networks, etc.

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik Apr 20, 2026 · 2 mentions

  • ▶ 1:14:50 Brandon Anderson GPT-II, GPT-III, GPT-III, you know, GPT-I, two, and three, like, there was a clear progression there. 2 times in the scene

Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different" Apr 3, 2026 · 1 mention

Measuring Exponential Trends Rising (in AI) — Joel Becker, METR Feb 27, 2026 · 2 mentions

  • ▶ 3:17 Alessio Fanelli Have the thread models been updated a lot, or do you feel like you're still using the same thread models as GPT-II of like, you know, paperclip factory, blah, blah, blah, you know, but like, how much are you rising, you know, increasing…
  • ▶ 9:58 Joel Becker You know, I think GPT-II can like sometimes do that task and, and, and sometimes not.

[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor Dec 30, 2025 · 4 mentions

  • ▶ 4:54 Ashvin Nair Yeah, like I would say that robotics is in kind of like the GPT-one to GPT-two area right now. 2 times in the scene
  • ▶ 24:17 Ashvin Nair I think a lot of people didn't really think of GPT or GPT-II as something that was like super compelling probably. 2 times in the scene

Greg Brockman on OpenAI's Road to AGI Aug 15, 2025 · 1 mention

⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs Aug 4, 2025 · 1 mention

  • ▶ 2:46 Stefano Ermon And so we published a paper at ICML last year that won the best paper award showing that for the first time, basically the discrete diffusion models were competitive on language generation without the regressive models up to the GPT-II

Information Theory for Language Models: Jack Morris Jul 2, 2025 · 3 mentions

  • ▶ 2:20 Jack Morris GPT two, GPT one from open AI were like interesting, but I think most people were into BERT at that time.
  • ▶ 42:58 Shawn Wang If you gave the O-one harness on top of GPT-II, you would get nothing because GPT-II didn't know enough. 2 times in the scene

Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI Jun 19, 2025 · 2 mentions

  • ▶ 9:35 Noam Brown I think it could have happened earlier, but if you try to do the reasoning paradigm on top of GPT-II, I don't think it would have gotten you almost anything.
  • ▶ 32:28 Spooks (Swyx) We had David Luan on, who I think was VP Eng at the time of GPT-I and II.

The Utility of Interpretability — Emmanuel Amiesen Jun 6, 2025 · 1 mention

  • ▶ 35:21 Emmanuel Ameisen well you have all your curve detectors, but think about all of the concepts that like Claude or even GPT-II need to, to know, like just in terms of, it needs to know about like all of the different

The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!! Mar 11, 2025 · 1 mention

Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin] Mar 7, 2025 · 1 mention

  • ▶ 5:01 Misha Laskin So it's really that, I don't know if you've been able, you would have been able to do the same stuff with an earlier model, like with a GPT-III or a GPT-II, but GPT-IV had this kind of base of intelligence that you could actually go from…

Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era Feb 28, 2025 · 1 mention

  • ▶ 8:58 Logan Kilpatrick I tweeted the other day that it's the, it feels like the GBT two era, the like order in which the timeframe in which model progress can actually happen.

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research Jan 26, 2025 · 2 mentions

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 2 mentions

  • ▶ 28:38 Shawn Wang Zero capabilities, and it's sudden emergence of GPT-IV. 2 times in the scene

Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI Nov 25, 2024 · 1 mention

  • ▶ 34:50 Shawn Wang You know, like, I think that, that is the, the core insight of the GPTs, the GPC one, two, three, that was

Agents @ Work: Dust.tt — with Stanislas Polu Nov 11, 2024 · 1 mention

  • ▶ 9:41 Alessio Fanelli How did the research evolve from, you know, the GPT-II and then getting closer to like…

Building AGI in Real Time (OpenAI Dev Day 2024) Oct 4, 2024 · 1 mention

  • ▶ 1:48:56 Sam Altman Like, we are so early, this is like, you know, maybe it's the GPT-II scale moment, but

Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph Sep 27, 2024 · 1 mention

llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE Sep 21, 2024 · 4 mentions

Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind Aug 28, 2024 · 2 mentions

  • ▶ 24:53 Nicholas Carlini I still would have said the same thing with like, after I had seen GPT two, I had written a couple of papers studying GPT two very carefully. 2 times in the scene

Is finetuning GPT4o worth it? Aug 22, 2024 · 2 mentions

The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap) Aug 2, 2024 · 1 mention

LLM Asia Paper Club Survey Round May 22, 2024 · 3 mentions

  • ▶ 40:01 unnamed speaker Um, on all sorts of models, uh, GPT-II Small is a favorite, just because it's very well understood by interpretability researchers, and you can see all sorts of features identified in GPT-II Small, um, and it's actually super easy to, to… 3 times in the scene

High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor Apr 24, 2024 · 2 mentions

  • ▶ 1:58 Jason Liu When GPT-II came out, I fine-tuned my own GPT-II to write, like, rap lyrics, and I was like, 2 times in the scene

Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit Apr 11, 2024 · 4 mentions

  • ▶ 13:07 Andreas Stuhlmüller And then the initial, I think the first models that kind of make sense were TPT-II and TNLG and like the, the early, early, um, 3 times in the scene
  • ▶ 26:58 Jungwon Byun It felt less like a level, GPT-III over GPT-II was like qualitative level shift.

Why Google failed to make GPT-3 -- with David Luan of Adept Mar 27, 2024 · 7 mentions

  • ▶ 7:10 Shawn Wang Um, but you mentioned GPT-II. 3 times in the scene
  • ▶ 12:38 Shawn Wang Is there another, you know, sort of GPT-II story that, like, you know, you love to get out there, um, that I think is, you think it's like underappreciated for, like, the amount of work that people put into it? 4 times in the scene

A Comprehensive Overview of Large Language Models - Latent Space Paper Club Mar 15, 2024 · 1 mention

  • ▶ 14:36 unnamed speaker Sampling or your autoregressive sampling of tokens to form your target sequence, which is what we've seen in T five and the decoder models, which I think all of us are familiar with things like GPT two, GP three, three, they are all there.

The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert Jan 11, 2024 · 5 mentions

The AI-First Graphics Editor - with Suhail Doshi of Playground AI Jan 2, 2024 · 1 mention

The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph Dec 17, 2023 · 1 mention

  • ▶ 59:36 Shawn Wang Um, and actually, uh, there's a pretty good cadence from GPT-II, three, and four, uh, that you can, if you project out, um, so four, uh, four is, uh, based on George Hotz's, um, uh, concept of, like, uh, 20 petaflots being, like, a human,…

The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis Dec 5, 2023 · 4 mentions

  • ▶ 10:28 Dylan Patel Like we are growing, you know, GPT-two to four, two to four is like twenty-twenty to twenty-twenty-two, right?
  • ▶ 37:18 Dylan Patel If you look at the total flops, right, you know, parameters times tokens times six, right, it's like, it's like a tiny, tiny fraction of GPT-II, which came out just a few months later, which was like, okay, 3 times in the scene

The End of Finetuning — with Jeremy Howard of Fast.ai Oct 20, 2023 · 1 mention

  • ▶ 1:19:59 Jeremy Howard Yeah, and they actually showed some nice, nice examples of, like, a GPT-II attention layer, and, like,

RWKV: Reinventing RNNs for the Transformer Era Aug 31, 2023 · 2 mentions

  • ▶ 47:01 unnamed speaker So, GPT-II was actually open source.
  • ▶ 52:59 Eugene Cheah So, so, so the history, the history behind it, right, and is that, um, I, I think, I think, like, a few years back, once with GBT II, Transformers started to pick up Steam, and I guess the whole world is starting to think, let's just…

FlashAttention-2: Making Transformers 800% faster AND exact Aug 3, 2023 · 1 mention

  • ▶ 24:21 Tri Dao Like, that's, that, that was the bet that they, they, they made, you know, after, I think, GPT-II, so they saw that, oh, scaling,
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.