GPT 2

product on 28 shows · 22 statements across 20 episodes · said 282 times in 174 episodes since 2019

Latent Space 73 No Priors 28 the MAD Podcast 28 TBPN 27 the a16z Podcast 26 Big Technology 24 20VC 14 the Y Combinator Startup Podcast 9 Acquired 7 A Product Market Fit Show 7 Sourcery 7 Lenny's Podcast 4 the Neon Show 4 My First Million 3 All-In 3 In Depth 2 BG2 Pod 2 the Knowledge Project 2 the Official SaaStr Podcast 2 the Startup Ideas Podcast 2 Mixergy 1 Innovators & Investors 1 Cheeky Pint 1 We Live to Build 1 American Optimist 1 WTF is with Nikhil Kamath 1 Top Founders 1 How I Built This 1

Mentions by year, every show

tap a year for its mentions
0075401508020192020202120222023202420252026episodesmentions
0408020192020202120222023202420252026episodes it came up in
0014028020192020202120222023202420252026episodesmentions per episode

Latent Space 73the MAD Podcast 28No Priors 28TBPN 27the a16z Podcast 26Big Technology 2420VC 14the Y Combinator Startup Podcast 920 more shows

2026 45 mentions in 32 episodes 1 per episode
2025 101 mentions in 64 episodes 2 per episode
2024 91 mentions in 51 episodes 2 per episode
2023 37 mentions in 22 episodes 2 per episode
2021 1 mention in 1 episode
2020 1 mention in 1 episode
2019 6 mentions in 3 episodes 2 per episode

every mention on every show, scene by scene, with the transcript →

22 statements about GPT 2, every show

NEON SHOW Assertion Contradicted
Koratana: GitHub Copilot originally began building on a GPT-2 Codex variant
“In 2019, 20, 20 time GPT-II had just been trained. And that was, I think, one of the first language models that came together and was coherent over longer horizons. And there was a variant of GPT-II that that was called Codex. And this is kind of the same name…”
Animesh Koratana May 20, 2026 ▶ 2:19 Great Founders Aren't Writing Code. They're Building Its Immune System. | Animesh, PlayerZero
MAD Assertion Not checkable as stated
DeepSeek's architecture is still built on a GPT-2 scaffold
“And you can actually, in fact, Take a GPT one or two model and with a few, I mean, few lines of code almost, you can transform it into the latest let's say deep seek version, 3.2 architecture. It's not like a big leap. It's still the same as scaffold.”
Sebastian Raschka Jan 29, 2026 ▶ 2:54 State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka
Nair: AI robotics is currently in its 'GPT-1 to GPT-2' era
“Yeah, like I would say that robotics is in kind of like the GPT-one to GPT-two area right now.”
Ashvin Nair Dec 30, 2025 ▶ 4:54 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Altman: Current AI memory is in the 'GPT-2 era'
“We're in like the, you know, the GPT-II era of memory”
Sam Altman Dec 18, 2025 ▶ 15:32 Sam Altman: How OpenAI Wins, ChatGPT’s Future, AI Buildout Logic, IPO in 2026?
a16z Assertion Not checkable as stated
Hsu: Current AI biological models are at GPT-1 or GPT-2 capability levels
“I find it helpful to frame these in terms of like GPT one, two, three, four, five capabilities, right? And I think most people would agree we're somewhere between GPT one and two, right?”
Patrick Hsu Sep 15, 2025 ▶ 15:28 Faster Science, Better Drugs
LATENT SPACE Assertion Partly supported
Brockman: Arc Institute trained 40B DNA model on 13T base pairs
“I'd say that maybe the neural net we produced, you know, it's a 40 B neural net trained on, you know, like 13 trillion base pairs or something like that. The results to be felt like GPT one, maybe starting to be GPT two level, right? It's like accessible or, a…”
Greg Brockman Aug 15, 2025 ▶ 17:53 Greg Brockman on OpenAI's Road to AGI
BIG TECHNOLOGY Assertion Not checkable as stated
Amodei: GPT-2 and GPT-3 were built to test RLHF at scale
“Actually, the original reason for building GPT-II and GPT-III, it was an outgrowth of the kind of AI alignment work that we were doing, right? Where myself and Paul Cristiano and some of the Anthropic co-founders had invented this technique called RL from huma…”
Dario Amodei Jul 30, 2025 ▶ 50:57 Anthropic CEO Dario Amodei: AI's Potential, OpenAI Rivalry, GenAI Business, Doomerism
Noam Brown: Reasoning paradigms would have failed on GPT-2
“If you try to do the reasoning paradigm on top of GPT-II, I don't think it would have gotten you almost anything.”
Noam Brown Jun 19, 2025 ▶ 9:35 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
SOURCERY Assertion Supported
Scale AI provided data labeling for OpenAI's GPT-2
“We actually helped OpenAI with GPT too.”
Leigh Marie Braswell May 2, 2025 ▶ 43:58 Windsurf: The Making of a Billion-Dollar AI Company | Leigh Marie Braswell, Kleiner Perkins · Sourcery with Molly O'Shea
20VC Assertion Supported
Nayak: Scale AI Partnered With OpenAI on Early GPT-2 RLHF
“Scale partnered with open AI very, very early on before chat GPT came out. This was like very early on RLHF when they were trying to tune models to summarize better based off of Reddit passages. And this is on GPT two.”
Aatish Nayak Apr 11, 2025 ▶ 8:22 20Product: How Scale AI and Harvey Build Product | Why PMs Are Wrong: They are not the CEOs of the Product | How to do Pre and Post Mortems Effectively and How to Nail PRDs | The Future of Product Management in a World of AI with Aatish Nayak
MAD Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Francois Chollet Apr 3, 2025 ▶ 11:44 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
NO PRIORS Prediction Not checkable as stated
Elad Gil compares OpenAI's o1 model to the GPT-2 stage
“Yeah, it's kind of like GPT-II for that approach, and the idea is that that could scale further over time, and it's kind of the initial early proof point that you feel like something really real is happening. And then it's just going to be subject to the same …”
Elad Gil Oct 17, 2024 ▶ 16:18 No Priors Ep. 86 | With Sarah Guo & Elad Gil
LATENT SPACE Assertion Supported
Karpathy: llm.c trains GPT-2 on one H100 node in 24 hours for $600
“You can train it on a single node of H-one-hundreds in about 24 hours, and that costs roughly 600 dollars.”
Andrej Karpathy Sep 21, 2024 ▶ 18:02 llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
LATENT SPACE Assertion Supported
Karpathy: llm.c was 20% faster and used 30% less memory than PyTorch
“At the time of that post, we were using, in LL and that's in 30% less memory, and we were 20% faster in training, just the truth.”
Andrej Karpathy Sep 21, 2024 ▶ 19:10 llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
a16z Assertion Supported
Andreessen: Andrej Karpathy's open-source GPT-2 costs under $100 to train
“And Andre Karpathy, who was a co-founder of OpenAI, who has left, actually now has an open source version of GPT-II where you can train your own GPT-II from scratch for less than 100 dollars in compute costs.”
Marc Andreessen Jul 16, 2024 ▶ 47:46 Trump Vs. Biden: Tech Policy
a16z Prediction Open · timeframe Jul 2029
Andreessen: GPT-2 scale AI models will run on smartphones within five years
“I think within five years, you're going to have models of this size on your phone.”
Marc Andreessen Jul 16, 2024 ▶ 48:16 Trump Vs. Biden: Tech Policy
20VC Assertion Contradicted
Luan: Scaling GPT-2 without architectural changes unlocked three-digit arithmetic
“When we were training GPT-II, we trained GPT-II in various different sizes. And at the smallest size, the model was just, like, unable to do three-digit arithmetic. But as the models got bigger and bigger and bigger, we didn't change anything else. We just had…”
David Luan Jun 24, 2024 ▶ 17:20 David Luan: Why Nvidia Will Enter the Model Space & Models Will Enter the Chip Space | E1169 · 20VC with Harry Stebbings
a16z Assertion Not checkable as stated
AI App Day-90 Retention Increased Consistently With GPT Model Upgrades
“There's an example of an AI-native companionship product where they actually did this test where they used GPT-one, GPT-two, GPT-three, and they actually tracked the cohort of user retention using each of the models. And it's very clear as the model quality im…”
Bryan Kim May 31, 2024 ▶ 5:05 7 Ways to Boost Retention (Both Pre- and Post-AI)
LATENT SPACE Assertion Not checkable as stated
Stuhlmüller: Elicit continues to use T5-based models
“We do also use, like, T-Five-based models, even, even now but started, yeah, started with GPT-II.”
Andreas Stuhlmüller Apr 11, 2024 ▶ 13:18 Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
Doshi: Image Generative AI Is Stuck in a 'GPT-2 Moment'
“So I think that we continue to feel like graphics and these foundation models for anything really related to pixels, but also definitely images continues to be very under invested. It feels a little like graphics is in like this GPT two moment, right? Like eve…”
Suhail Doshi Jan 2, 2024 ▶ 18:07 The AI-First Graphics Editor - with Suhail Doshi of Playground AI
NO PRIORS Assertion Supported
Henry: Square Used GPT-2 in Square Messages Product
“Even kind of prior to this kind of latest big shift in the AI landscape, you know, we were using GPT two as part of us, the square messages product being you know, a virtual assistant to help customers, you know, answer responses to customer inquiries and vari…”
Alyssa Henry Dec 14, 2023 ▶ 5:19 No Priors Ep. 44 | With Former Square CEO Alyssa Henry
ACQUIRED Assertion Supported
GPT parameter counts grew from 120 million in GPT-1 to 1.7 trillion
“GPT-I had roughly a hundred and twenty million parameters that it was trained on. GPT-II had 1.5 billion. GPT-III had a hundred and seventy five billion, and GPT-IV, OpenAI hasn't announced, but it's rumored that it has about 1.7 trillion parameters that it wa…”
David Rosenthal Sep 6, 2023 ▶ 41:11 Nvidia Part III: The Dawn of the AI Era (2022-2023) (Audio) · Acquired

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.