GPT-4, every mention
192 scenes, the whole family · ← back to GPT-4
tap a year for its mentions
every year anyone Shawn Wang 45Nathan Lambert 22Michael Royzen 15Alessio Fanelli 14Thomas Scialom 10Dylan Patel 10Alistair Pullen 9Jerry Liu 6Sander Schulhoff 5Pratik Bhavsar 5
Verbatim, from the transcripts: the passages where GPT-4 comes up
⏭️ Forward Deployed: Voice AI on what works in 2026
- ▶ 19:33 unnamed speaker Are you guys using the latest ones, or are you using, you know, like GPT-IV, because it's, like, the right balance between speed and intelligence?
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 1:09:37 Shawn Wang There's still people out there using four.
Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
- ▶ 2:55 Simon Last I mean, like it was even right when we got access to like GPG four in late, one of the first ideas we had is like,
⚡️ The best engineers don't write the most code. They delete the most code. — Stay Sassy
- ▶ 54:59 unnamed speaker Like, we're no longer impressed by GPT-IV, which was so impressive when it just, just came out, right?
Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
- ▶ 11:59 Ryan Lopopolo And this I think is like the fundamental difference between reasoning models and the four ones and four O's of the past where these models could not think.
Measuring Exponential Trends Rising (in AI) — Joel Becker, METR
- ▶ 2:04 Joel Becker Um, but there's this, this wonderful report on our website, GPT-V, um, report, an analogous one for GPT-V as well, trying to make this more sort of structured case, that it doesn't pose these really large-scale risks, you know, eventually… 3 times in the scene
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- ▶ 53:07 Doug O'Laughlin GPT, 3.5 or four for me, where there's that first time where you're like, okay,
Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
- ▶ 31:40 Sarah Wang You know, if you fast, you know, if you, if you rewind time even before that, GPT-IV was number one for nine months, 10 months.
⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
- ▶ 13:17 Pratyush Maini So coming back to this phenomenon, uh, and I was, like, curious, like, okay, now we know that this suddenly, uh, happens, or, like, this infection point happens in GPT-IV.
- ▶ 22:54 Pratyush Maini For instance, it could be the GPT-IV model and you're trying to query that model to generate a textbook or an essay or a paragraph.
🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
- ▶ 9:13 Andrew White And so I was a red teamer for GPT-IV 4 times in the scene
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 8:36 Micah Hill-Smith That in the extreme and like you get crazy cases, like back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four.
- ▶ 38:47 Shawn Wang It's not as good for the content creator rumor mill where I can say, oh, GPT four is this small circle.
- ▶ 58:55 Micah Hill-Smith The, like, one fact on that is that you can get intelligence at the level of GPT-IV for over a hundred times cheaper than GPT-IV was at launch right now. 3 times in the scene
[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
- ▶ 0:26 unnamed speaker Yeah, yeah, and you were on with us for GPT-IV.
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 17:18 unnamed speaker The idea was, you, you started with codecs, someone else was doing instroft GPT, then we launched GPT, four, four O, I guess O-one.
- ▶ 26:35 unnamed speaker I think the question is, there used to be more of this, and now I know this less, which is, well, the stuff we've released, we're like, you know, internally, we're like six months ahead of, like, you say, part of the reason why people, uh,…
⚡️ Building the AI Hardware Engineer with Matthias Wagner, Co-founder of Flux
- ▶ 4:25 Matthias Wagner We shipped that I think a month or two months before GPT-IV became publicly available.
⚡ Inside GitHub’s AI Revolution: Jared Palmer Reveals Agent HQ & The Future of Coding Agents
- ▶ 8:25 unnamed speaker And it was the four, GVT four era? 3 times in the scene
- ▶ 9:21 Jared Palmer And from, you know, GPT-FOR, GPT-FOR, the big boy, we never really got GPT-FOR turbo working.
The Agents Economy Backbone - with Emily Glassberg Sands, Head of Data & AI at Stripe
- ▶ 1:03:42 Emily Glassberg Sands With, like, GPT-IV, like, it didn't work, and then you swap in GPT-V, and all of a sudden, it does.
Breaking AI to Fix It: Ian Webster's Journey from Discord's Clyde to Promptfoo's $18M Series A
- ▶ 4:17 Ian Webster And then we quickly found, especially because this was like before GPT four was, was economical.
Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
- ▶ 3:24 Kyle Corbitt We were looking at different ideas, and one thing we realized was we actually started the company immediately after the GPT-IV launch, and what we saw as the opportunity in the market at the time, which has changed since then, was GPT-IV… 4 times in the scene
Long Live Context Engineering - with Jeff Huber of Chroma
- ▶ 17:57 Shawn Wang And then GPT four one and Gemini flash are, are, uh, degrade a lot quicker in terms of the context length.
Greg Brockman on OpenAI's Road to AGI
- ▶ 1:32 Greg Brockman Well, I'd say that after we trained GPT-IV, we had a model that you could talk to, and I remember doing the very first, we did the post-training, we actually did a instruction following post-train on it, so it was really just a data set… 2 times in the scene
- ▶ 16:00 unnamed speaker We're in the, you know, multiple low double digit to high single digit range for, you know, GPT four, 4.5 and five, but we, you know, we're not confirming that, but like, you know, we're, we're, we're scaling there.
- ▶ 18:14 Greg Brockman A GPT-III or GPT-IV, not a GPT-V for sure, right?
- ▶ 19:37 unnamed speaker If I think about three, four, five as the major versions, I think three is very text-based, kind of like RLHF really getting started. 2 times in the scene
- ▶ 25:48 Greg Brockman So I definitely have my library of prompts that I've built up since, you know, the GPT-IV days. 2 times in the scene
- ▶ 45:44 unnamed speaker Uh, it's since the day you launched GPT-IV,
The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
- ▶ 28:11 Stephanie Palazzolo You know, the past year we've seen a lot of progress, especially in reasoning models, but I think, like, under all that has been the fact that, like, GPT-IV has been kind of the, like, leading GPT model for, for a very long time, and…
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 1:05:14 Nathan Lambert We're serving GPT-IV.
- ▶ 1:16:23 Nathan Lambert A lot of it is just more resources, but it's like, like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 28:00 Shawn Wang GPT four is going to be 100 trillion billion.
- ▶ 3:04:24 Scott Wu Um, and GPT 3.5 was better, and GPT four was better, right?
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
- ▶ 7:21 Pratik Bhavsar And some shaking of the leaderboard also later on, but when we released, we found that Gemini models are performing really good, and they were extremely cost-efficient at the time compared to the second and second model like GPT-IV at the…
- ▶ 23:41 Pratik Bhavsar It looks way, we have seen situation previously, where we were using probably GPT-IV turbo or something, or even GPT-IV.V long back, I'm talking like two years back. 3 times in the scene
- ▶ 23:41 Pratik Bhavsar It looks way, we have seen situation previously, where we were using probably GPT-IV turbo or something, or even GPT-IV.V long back, I'm talking like two years back.
Personalized AI Language Education — with Andrew Hsu, Speak
- ▶ 1:00:14 Andrew Hsu In 2023, when we first launched our AI role plays using GPT-IV, back then people were way more concerned about safety, right?
- ▶ 1:02:15 Andrew Hsu People, if you recall, expected the world to kind of explode when GPT-IV came out.
⚡️Anthropic vs Cognition on Multi-Agents: A Breakdown with Dylan Davis
- ▶ 24:27 Shawn Wang We are in one of these situations where, you know, in a world where the cost of intelligence for a given set of intelligence, let's say GPT-IV, let's say O-one, whatever, it is literally falling a hundred X over the course of one year.
Information Theory for Language Models: Jack Morris
- ▶ 34:03 Jack Morris A lot of people have this shared idea, like, you know, Claude and GBT-IV probably do a lot of very similar internal computation because both of them are trained on trillions of tokens of human written text, even if they have different
- ▶ 43:05 Shawn Wang You need a GPT-III and GPT-IV in order to then get O-one. 2 times in the scene
Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
- ▶ 1:07:38 Spooks (Swyx) So let's say that the pre-training scaling paradigm took about five years from like discovery of GPT to scaling it up to GPT-IV.
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 1:42:50 Vibhu (Viboo) Recently, you know, as OpenAI has Sunset GPT-IV, a lot of people are like, oh, can we put out the weights?
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
- ▶ 21:57 Will Brown People, the sorts of people who I think also really like GBT-IV.
DeepWiki: The GitHub Encyclopedia
- ▶ 29:13 unnamed speaker They used to use GPT-IV
⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
- ▶ 16:38 Jack Hopkins Whereas GPT four would use defensive programming, use self assertions. 2 times in the scene
- ▶ 26:38 Jack Hopkins Um, and then you have, uh, your Gemini's and your GPT-IVs and all the way down here, you have GPT-IV and 2 times in the scene
Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
- ▶ 28:33 Charles Packer I think if you brought, like, GPT-IV back in time, 10 years, and, like, you didn't really explain how it worked, and then you just kind of showed it to somebody, they'd be kind of surprised if, about, like, all the inherent limitations,…
SF Compute: Commoditizing Compute
- ▶ 19:46 Michael Swix (Swyx) So let's say GPT-IV and O-ONE both had total training costs of like a five hundred million dollars is the rough estimate.
- ▶ 33:19 Evan Conrad Um, but the, if you imagine like what an open AI needs, um, for, um, like GPT-IV, it's like tremendously big, but that's because it's a consumer product that has almost all the inference demand.
The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
- ▶ 21:15 Romain Huet Last year when we introduced vision capabilities, you know, vision capabilities were in like a vision preview model based off of GPT-IV and then vision capabilities now are like obviously built into GPT-IV. 2 times in the scene
Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
- ▶ 4:34 Misha Laskin It was really that, so Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were… 2 times in the scene
smol agents are all you need
- ▶ 14:19 Aymeric (Emmerich) So, GPT-IV performed well with small agents already, and when I tried O-one, I was
The AI Architect: Bret Taylor
- ▶ 1:32:55 Bret Taylor And I'll say actually performance is an ambiguous word, basically the latency of these reasoning models more reasonable, because if you think about say GPT-IV, which was, I think a huge step change in intelligence, it was quite slow and… 2 times in the scene
Why every AI Engineer needs an AI Gateway (ft Portkey.ai CEO)
- ▶ 2:25 unnamed speaker I think before maybe there wasn't as much value of like routing between four or like four or many, you kind of knew what to use, but now the latency of the reasoning model is much higher.
The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
- ▶ 15:19 Karina Nguyen So you actually need to, like, go back to, like, I don't know, like, GPT-E for model card and, like, read the appendix just to, like, make sure that, like,
Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- ▶ 14:11 Shawn Lewis And, and applying a one is, is very different than, than using GPT-IV. 3 times in the scene
OpenAI o1 isn’t a chat model (and that’s the point)
- ▶ 12:08 Ben Hillock Even with LLMs in general, I think that, uh, you know, even since like, you know, GP four, GP four five, like there hasn't been an instruction manual.
- ▶ 16:18 Ben Hillock Like I think in a lot of ways that there's not a huge Delta right now, like essentially if you're using like GPT four or something to like do those sorts of insights, it's like, 2 times in the scene
Beating Google at Search with Neural PageRank and $5M of H200s — with Will Bryk of Exa.ai
- ▶ 12:51 Will Bryk You could imagine, for example, running GPT-IV over the entire web.
- ▶ 19:00 Will Bryk I mean, you could also have GPT-IV level systems calling search, but it's just because of the cost of inference, it's just better to have a very efficient search tool and a very efficient LLM, and they're built for different things.
- ▶ 48:39 Will Bryk But then you're actually completely missing the meaning of the document, whereas an LLM like GBP-IV is really good at labeling.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 11:04 Shawn Wang GPT, 3.5 and GT four being 95%.
- ▶ 28:38 Shawn Wang Zero capabilities, and it's sudden emergence of GPT-IV.
- ▶ 40:04 Shawn Wang Uh, we were like, okay, we go from one 75 B to one point A B, uh, one point T and that was GPT three to GPT four. 2 times in the scene
- ▶ 1:16:48 Shawn Wang Um, so for the same amount of ELO, let's say GPT four, 23, uh, would, would be about sort of 1175 in ELO.
- ▶ 1:27:47 Shawn Wang Like this is where the messaging of Omni model really started kicking in, you know, previously four and four turbo were all text.
- ▶ 1:27:47 Shawn Wang Like this is where the messaging of Omni model really started kicking in, you know, previously four and four turbo were all text.
0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
- ▶ 54:50 Itamar Friedman But at some point, OpenAI closed their GPT-EV or whatever, and, uh, that was part of my answer. 3 times in the scene
[Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
- ▶ 42:44 Eugene Yan Yeah, the poll, poll essentially, the poll paper is actually talking about, um, just combining multiple, using multiple weaker LLMs, you can get, uh, as, as good as, uh, an LLM judge as, um, GPT-IV. 3 times in the scene
[Paper Club] BERT: Bidirectional Encoder Representations from Transformers
- ▶ 0:37 unnamed speaker I wonder if OpenPipe does this, that only does, mirrors a structured output GPT-IVO call and just mirrors it until it has enough data for BERT and then just switches you to BERT.
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 32:43 Alessio Fanelli Because I think everybody agrees with open source models eventually will catch up, and I think with four, then with Lama to three, one, four, five B, we close the gap, and then all one just reopened the gap so much, and it's unclear.
Agents @ Work: Lindy.ai (with live demo!)
- ▶ 10:48 Florent Crivello I use GPT for turbo.
Agents @ Work: Dust.tt — with Stanislas Polu
- ▶ 13:59 Stanislas Polu With the, the, the success of GPT-IV and stuff shows that it was the right approach.
- ▶ 19:16 Stanislas Polu I had seen GPT-IV internally at the time. 3 times in the scene
- ▶ 41:09 Shawn Wang You can just say, oh, I'm going to run it on GPT-Fort, GPT-Fort Turbo, or... 4 times in the scene
- ▶ 41:09 Shawn Wang You can just say, oh, I'm going to run it on GPT-Fort, GPT-Fort Turbo, or... 6 times in the scene
In the Arena: How LMSys changed LLM Benchmarking Forever
- ▶ 3:37 Wei-Lin Chiang At that time, it was, like, GPT-FOR was just announced. 2 times in the scene
- ▶ 15:20 Anastasios Angelopoulos And when the coefficient is large, it tells you that GPT four or whatever, very large coefficient.
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- ▶ 7:40 Jesse Hu They, like, facilitate the community, they review all the submissions, they update the leaderboard, uh, on, on, like, some sort of, like, biologically basis, and you can see here there's been a ton of submissions where they started out…
- ▶ 11:58 Jesse Hu Or, you know, what they call two percent for GPT-IV, which is not really legit, right?
Building the Silicon Brain - Drew Houston of Dropbox
- ▶ 3:50 Drew Houston ChatGPT launch being the starting gun for all, for the whole AI era of computing and then having API access to three and then early access to GPT-IV.
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 1:29 Vibhu Sapra So think GPT-IV, GPT-IV, Gemini, whatnot, and
- ▶ 26:05 Vibhu Sapra Write question answers on this OCR. 2 times in the scene
- ▶ 37:36 Vibhu Sapra I think when you look a little deeper, they, they mentioned, like, you know, the one B is on par with GPT-IV on most benchmarks, but
Production AI Engineering starts with Evals
- ▶ 32:49 Ankur Goyal What is incredibly empowering about these, I would just maybe say that the quality that transformers bring to the table, and even BERT does this, but you know, GPT three and then four, like very emphatically do it, is that software…
- ▶ 1:33:26 Ankur Goyal In fact, no one knows whether this is true or whatever, but GPT-IV was famously a mixture of experts. 2 times in the scene
Building AGI in Real Time (OpenAI Dev Day 2024)
- ▶ 6:49 NotebookLM Host 1 What can O-one do that, say, GPT-IV can't? 3 times in the scene
- ▶ 18:30 Ilan Biggio Like I've done back in the day, like from GPT-IV to GPT-IV. 2 times in the scene
- ▶ 41:46 Romain Huet But can GPT-IV do some of that? 2 times in the scene
- ▶ 49:59 Simon Willison Using, like, uh, GPT-IV wrote me the JavaScript, and I got that live just in time, and
- ▶ 56:47 Simon Willison I didn't think GPT-IV Vision could do bounding boxes at all.
- ▶ 1:11:53 Alistair Pullen This process continues, like even going from, you know, when we first started going from 3.5 to four, we saw this happen.
- ▶ 1:12:00 Alistair Pullen Um, and then from four turbo to four O and then from four O to O one, we've seen the performance get better every time.
- ▶ 1:25:53 Sam Altman If you think about what happened from last decade to this one in terms of model capabilities, and you're like, I mean, if you go look at like, if you go from like, oh, one on a hard problem back to like four turbo that we launched a little…
- ▶ 1:49:02 Sam Altman like, we know how to get to GPT-IV, and we have the fundamental stuff in place now to get to GPT-IV. 2 times in the scene
- ▶ 1:49:20 Sam Altman Plan for it to feel like way more of a year of improvement than from, ah, four turbo to one.
[Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
- ▶ 7:48 Eugene Yan So essentially, you're just getting GPT-IV to generate some kind of criteria, and this kind of criteria can be in the form of code, 2 times in the scene
page 1 of 2 · 100 scenes per page · newest episode first next →