GPT-4o, every mention
87 scenes, the whole family · ← back to GPT-4o
tap a year for its mentions
every year anyone Shawn Wang 32Alessio Fanelli 15Michelle Pokrass 14Vibhu Sapra 8Karina Nguyen 5Noam Brown 4Thomas Paul Mann 3Raaz Dwivedi 3Pratyush Maini 3Ankur Goyal 3
Verbatim, from the transcripts: the passages where GPT-4o comes up
⏭️ Forward Deployed: Voice AI on what works in 2026
- ▶ 35:58 unnamed speaker Over GPT-Foto is 4.1, you know, the popular real-time models.
Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
- ▶ 7:28 unnamed speaker And I think this, this branch of the LLM tech tree of real time was kind of started with the four O launch, which like I think a lot of people indexed on.
Podcast Crossover: AIE, AGI, frontier lab strategy with @matthew_berman and @swyxtv
- ▶ 4:28 Shawn Wang The current workloads on existing models, like, four-oh is still being used, right, by some, by some folks, because they, they just don't move, like, once the thing works, it works, like, don't, don't touch it.
Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
- ▶ 11:59 Ryan Lopopolo And this I think is like the fundamental difference between reasoning models and the four ones and four O's of the past where these models could not think.
The Stove Guy: Sam D'Amico Shows New AI Cooking Features on America's Most Powerful Stove at Impulse
- ▶ 33:30 Sam D'Amico and, you know, as long as I don't get one-shotted by GPT, or, you know, other stuff, I avoided the whole four-oh.
Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
- ▶ 34:00 unnamed speaker Gone away from the original four old vision of the Omni model.
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- ▶ 1:00:22 Shawn Wang Where, so like, you can see the, the nice, uh, increase from four O to Opus four one.
⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
- ▶ 9:55 Pratyush Maini And then the chat GPT photo model, uh, update that was 2 times in the scene
- ▶ 16:43 Pratyush Maini O-one reasoning traces would have been available to the pre-training team in December, or the mid-training team, I should say, in December, and then from December until, ah, what's like four months from there, end of April, the…
[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 17:18 unnamed speaker The idea was, you, you started with codecs, someone else was doing instroft GPT, then we launched GPT, four, four O, I guess O-one. 2 times in the scene
Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
The Agents Economy Backbone - with Emily Glassberg Sands, Head of Data & AI at Stripe
- ▶ 1:03:47 Emily Glassberg Sands Or we use, like, uh, GPT-IV-O, and it was, like, kind of a little bit expensive to justify the humans that it was replacing for a particular risk-related task, but then next thing we know, like, O-three mini is out, and it's, like, you…
⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
- ▶ 37:28 Shawn Wang Uh, I think that the, the other thing, you know, uh, especially you mentioned the GPT-V router is the difference because like MOEs have routers.
Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
- ▶ 30:09 Barak Lenz It's a GPT-IV-O is already an AI system.
⚡️Traversal: Causal ML and Reinforcement Learning
- ▶ 30:44 Raaz Dwivedi You are trying to argue across so many symptoms and such a complex architecture of what the root cause is and a simple, you know, photo and these, the flagship models would not work. 3 times in the scene
Amp: The Emperor Has No Clothes
- ▶ 49:32 Alessio Fanelli Yeah, I wrote this article for the GPT-V release about, um, models self-improving for coding. 2 times in the scene
A Technical History of Generative Media
- ▶ 29:56 unnamed speaker Like, uh, obviously Gem, I, I honestly, I still think Gemini is underrated because they were first and then, but then obviously, uh, opening, I did the four or image gen and that was a huge thing.
- ▶ 48:10 Batuhan Taskaya You can't like, even the editing models, you know, GPT image one, or like flux context, whatever coin image, if you put your face or like, if you put like multiple people, whatever, you can't get the quality.
Greg Brockman on OpenAI's Road to AGI
- ▶ 42:43 Greg Brockman Four.
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
- ▶ 1:09:04 Shawn Wang My pushback is on this is just, if you're doing image editing, four O should do it, do all of it. 2 times in the scene
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 1:59:21 Shawn Wang It's Claude, it's Oro. 2 times in the scene
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
- ▶ 4:45 Pratik Bhavsar Uh, powered by GPT-FORO, and it has multiple judges so that we get higher performance, and that's how we kind of get the score for each sample, and we aggregate for a data set, and then we do average over the data set, and that's how we…
- ▶ 6:45 Alessio Fanelli For me, it was Mistral Small being the best open source model above Foro and DeepSeq VIII.
- ▶ 32:53 Pratik Bhavsar Like let's say I generate by 3.7 and then check on 4.1
Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
- ▶ 3:28 Noam Brown So at this point, like, you know, the truth is, you know, GPT-IV-O and like O-III, these models are like passing the Turing test.
- ▶ 11:04 Spooks (Swyx) So let's say we have, now we have the four O like natively omni model type of thing, then that also makes O three really good at GeoGuessr.
- ▶ 18:23 Noam Brown Like, before the reasoning models emerged, there was, like, all of this work that went into engineering these, like, agentic systems that, like, made a lot of calls to GPT-IV-O or, like, these non-reasoning models to get reasoning behavior.
- ▶ 36:06 Noam Brown Obviously, a lot more people use GPT-IV-O and just, like, the default on chat GPT and that kind of stuff. 2 times in the scene
Quadratic: The AI Spreadsheet
- ▶ 17:19 Shawn Wang So there's, you know, a lot of people when ImageGen came out for GPT-IVO, they started gibbifying everything, but you can also gibbify charts.
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 44:52 unnamed speaker GPT-FORO update, but we don't know that for sure, because they have interp teams, they just
Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
- ▶ 23:10 Alessio Fanelli Yeah, I'm curious because you have for a mini and you have for all, I'm curious, like, if you think that makes a big difference or not as much. 2 times in the scene
GPT 4.1: The New OpenAI Workhorse
- ▶ 4:22 Michelle Pokrass Basically, the way we got here is that GPT-IV. is like a pretty big improvement over the four O line, and we really wanted to signify that. 5 times in the scene
- ▶ 23:43 Michelle Pokrass Um, and we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot.
- ▶ 32:15 Michelle Pokrass Um, so you can see, like, 4.1 mini is actually quite significantly better than four o mini, um, and not that far away from the old four o.
- ▶ 42:45 Shawn Wang Oh, uh, but like not, not a ton, but like cheaper. 2 times in the scene
Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
- ▶ 30:28 unnamed speaker Um, everything else from that was pretty bog standard, Ruby on Rails application with SQLite, GPT-FORO, structured output. 3 times in the scene
Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
- ▶ 17:56 Sujay Jayakar You know, oh, three does do better than four.
The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
- ▶ 7:52 Alessio Fanelli I think the, when I first said web search, I thought you were going to just expose a API that then return kind of like a nice list of thing, but the way its name is like GPT for O search preview. 3 times in the scene
- ▶ 8:15 Alessio Fanelli GPT four O is 30% accuracy.
- ▶ 20:40 Shawn Wang And you seem all like fine tunes of Foro.
Raycast: Your AI Automation Assistant
- ▶ 8:15 Thomas Paul Mann and so we looked into all the various models we had, um, and then we picked, at the moment, it's gbd-for-o and gbd-for-o-mini, 3 times in the scene
The AI Architect: Bret Taylor
- ▶ 1:33:13 Bret Taylor Once you got to four O and four O mini, you know, it, it opened the door to a lot of different applications, both for cost and latency. 2 times in the scene
Why every AI Engineer needs an AI Gateway (ft Portkey.ai CEO)
- ▶ 19:16 Rohit Agarwal So you could always say, okay, if it comes to node A, always hit GPT-FORO, and node B is DeepSeek R-ONE,
The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
- ▶ 18:04 Alessio Fanelli And especially now, I know that with Foro for Canvas, you've done RL after on the model.
- ▶ 30:26 Karina Nguyen So we actually, what we did was actually retrain the entire full O plus our canvas stuff. 5 times in the scene
The Unreasonable Effectiveness of Reasoning Distillation: using DeepSeek R1 to beat OpenAI o1
- ▶ 19:35 Shawn Wang And then we also, uh, on this podcast, we interviewed Cosign that actually fine tune four O to, on three bench, sorry, uh, to, to, to, to achieve data on three bench.
OpenAI o1 isn’t a chat model (and that’s the point)
- ▶ 6:32 Dan McAteer I'm typically using LLMs, mostly for, for coding use cases, and, like, using Sonnet, 3.5, and GPT four out.
- ▶ 19:14 Shawn Wang And what I do is actually I ask, um, four O's perfectly fine at this.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 14:11 Shawn Wang LAMA-FORO-FIB does well compared to Gemini and GPT-E-FORO.
- ▶ 37:07 Shawn Wang That, that Cosign, which we talked about, we talked to him on Dev Day, can just fine-tune for O to beat O-one.
- ▶ 1:17:43 Shawn Wang Um, and I think what you're starting to see now, uh, in July is the emergence of four O mini and deep sea V two as outliers to the July frontier where July frontier used to be maintained by four O Lama four five. 3 times in the scene
- ▶ 1:25:27 Shawn Wang And obviously now we know that Sonnet is, is kind of the workhorse, um, just like four O is the workhorse of, of OpenAI. 3 times in the scene
Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
- ▶ 14:53 Graham Neubig This is old, um, and we need to update this, basically, but, um, we evaluated Claude, GPT-FORO, O-ONE-MINI, um,…
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 54:58 unnamed speaker This is the year that vision language models became mainstream, with every model from GPT-Forty to One, to Claude Three, to Gemini One, and Two, to Llama 3.2, to Mistral's Pix-Trol, to AI-Two's Pixmo, going multimodal.
Windsurf: The Enterprise AI IDE
- ▶ 22:14 unnamed speaker It's Claude, it's Foro.
[Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
- ▶ 46:06 Shreya Shankar But we use GPT four.
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
- ▶ 37:53 Shawn Wang So much so that they, they made Claude's on it, and for, oh, just, like, they, they made the previous state of the art look bad.
Agents @ Work: Lindy.ai (with live demo!)
- ▶ 32:01 unnamed speaker Was it a Foreo? 3 times in the scene
Agents @ Work: Dust.tt — with Stanislas Polu
- ▶ 41:58 Stanislas Polu It's funny how the comparison between GPT for all and GPT for turbo is still up in the air on function calling. 2 times in the scene
In the Arena: How LMSys changed LLM Benchmarking Forever
- ▶ 33:03 Shawn Wang So it's, for example, GPT-FORO August has fallen from 1290 to 1260 over the past few months. 2 times in the scene
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 7:53 Vibhu Sapra Oh, then there's a large one, which is mobile. 2 times in the scene
- ▶ 30:52 Vibhu Sapra but yeah, I mean, across the board, the seven B, the one B, the 72 B very on par with four O four V 1.5 thoughts on it and whatnot. 3 times in the scene
- ▶ 45:19 Vibhu Sapra And the, wait, so our seven B and QEN seven B based models comfortably sit between GPT four V and four O on both. 3 times in the scene
Production AI Engineering starts with Evals
- ▶ 1:32:05 Ankur Goyal So I, I think that sort of shook out of me, that, that temptation as an engineer that you have to say, oh, you know, GPT-IV-O is good at this, 3 times in the scene
Building AGI in Real Time (OpenAI Dev Day 2024)
- ▶ 0:06 unnamed speaker Until dev daylights, code ignites Real-time voice streams reach new heights O-one and GPT-FOR-O in flight Fine-tune the future,
- ▶ 26:38 Olivier Godement The vision here is, at the moment, like, most developers, like, use, like, a one-size-fits-all model, like, the off-the-shelf, like, GP-for-O, essentially.
- ▶ 1:03:32 Simon Willison It'll be interesting to see if GPT-FORO can handle that or not.
- ▶ 1:07:00 unnamed speaker Special shout-out to listeners like Jesse from Morph Labs when he came on to talk about how he created synthetic datasets to fine-tune the largest lauras that had ever been created for GPT-for-O to post the highest-ever scores on Sweebench… 5 times in the scene
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
- ▶ 44:15 Shawn Wang So, GPC four O is currently five dollars per million input tokens. 2 times in the scene
[Paper Club] 🍓 On Reasoning: Q-STaR and Friends!
- ▶ 34:25 unnamed speaker Uh, the backspace token paper is not called the backspace token paper, but anyway, there was, there was kind of one, uh, observation in the wild on four O where, um,
Building AGI with OpenAI's Structured Outputs API
- ▶ 30:14 Shawn Wang Able to apply the structured output system on backdated models, like, uh, for May, as well as mini, as well as August. 4 times in the scene
- ▶ 41:11 Michelle Pokrass And if Foro works well for you, that's great. 7 times in the scene
Is finetuning GPT4o worth it?
- ▶ 12:58 Alistair Pullen Like for a fine tuning came out either.
- ▶ 52:44 Alistair Pullen To my knowledge, at least right now, state of the art also, which makes sense, but also GPT four, oh, gets, I believe, 33%, which is like, I double check that, but the August one, the new one. 2 times in the scene
Answer.ai & AI Magic with Jeremy Howard
- ▶ 1:02:13 Jeremy Howard I created Claudette to be as Claud-friendly as possible, and then after I did that, um, and then with Claud- with, particularly with GPT-IV-O coming out, I kind of thought, okay, now let's create something that's as,
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 42:56 unnamed speaker And by the way, also OpenAI did it with GPC four O
- ▶ 44:13 Alessio Fanelli Well, the first news is that, uh, Four O Voice is still not out, even though the, the demo was great. 3 times in the scene
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 13:01 Thomas Scialom There's now GPT-IV-Zero, of course, and we're closed, but we're not there yet.
- ▶ 37:51 Thomas Scialom But, by far, compared to the version originally released, uh, even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.
- ▶ 54:52 Alessio Fanelli GVD four was a hundred K four O is 200 K. 2 times in the scene
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 1:29:50 unnamed speaker Uh, so this is a little bit of commentary on GPT-IV-O and Chameleon. 2 times in the scene
How to train a Million Context LLM — with Mark Huang of Gradient.ai
- ▶ 58:39 Shawn Wang Um, some people were complaining about GPT-IV-O that, um, 2 times in the scene