GPT-4

also referred to as: gpt-iv

includes GPT 4 Turbo, GPT 4 With Vision, GPT 4 Vision

43 statements across 32 episodes · 17 bullish · 11 bearish · 29 people on the record · first statement Jun 20, 2023 by George Hotz · said 323 times in 87 episodes since 2023 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (45), Nathan Lambert (22), Michael Royzen (15), Alessio Fanelli (14), Thomas Scialom (10), Dylan Patel (10), Alistair Pullen (9), Jerry Liu (6)

tap a year for its mentions
0010020200402023202420252026episodesmentions
020402023202420252026episodes it came up in
004208402023202420252026episodesmentions per episode
2026 21 mentions in 11 episodes 2 per episode
2025 76 mentions in 33 episodes 2 per episode
2024 161 mentions in 34 episodes 5 per episode
2023 65 mentions in 9 episodes 7 per episode

every mention, scene by scene, with the transcript →

Everything said about GPT-4, oldest first

Jun 20, 2023 neutral
Assertion Not publicly verifiable
Hotz: GPT-4 is an 8-way mixture model with 220B parameters per head
“GPT-IV is two hundred twenty billion in each head, and then it's an eight-way mixture model.”
George Hotz Jun 20, 2023 ▶ 49:48 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Oct 12, 2023 negative
Insight
Liu: Avoid pre-GPT-4 models for tasks requiring complex reasoning
“Like, I'm one of the first to say, like, you know, you shouldn't use anything pre-GPT-IV for anything that requires, like, complex reasoning because it's just going to gonna be unreliable. Okay, disregarding stuff like fine-tuning.”
Jerry Liu Oct 12, 2023 ▶ 9:28 RAG is a hack - with Jerry Liu of LlamaIndex
Nov 3, 2023 positive
Assertion Not checkable as stated
Royzen: Users switch to Phind when ChatGPT-4 fails on code
“What really shocks us is that a lot of the people who do that they're coming from ChatGPT. So they tried it in ChatGPT with ChatGPT-IV. It didn't work. Maybe it required like some multi-step reasoning. Maybe it required to like, Some internet context or someth…”
Michael Royzen Nov 3, 2023 ▶ 27:07 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Nov 3, 2023 bearish
Prediction Not checkable as stated
Royzen: The leap from GPT-4 to GPT-5 will be smaller
“I think that GPT-IV, my hypothesis is that the jump from four to 4.5, or four to five, will be smaller than the jump from Three to four.”
Michael Royzen Nov 3, 2023 ▶ 37:38 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Nov 3, 2023 positive
Assertion Not checkable as stated
Royzen: Training on code unlocked general spatial and temporal reasoning
“We've seen emerging capabilities in the find model, whereby training it on high quality code, it can actually, like, reason better. It went from not being able to solve like, World problems where like riddles where like with like temporal and like low, like pl…”
Michael Royzen Nov 3, 2023 ▶ 43:51 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Nov 3, 2023 negative
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Dec 5, 2023
Assertion Supported
Patel: SemiAnalysis reported GPT-4 mixture of experts architecture in January
“Just being clear, I talked about mixture of experts in January, it's just people didn't really notice it.”
Dylan Patel Dec 5, 2023 ▶ 1:08 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 neutral
Opinion
Patel: Frontier AI model training costs are effectively irrelevant
“In my opinion, I think that's a little bit spicy, but yeah, it's like training costs are irrelevant, right? Like GPT-IV, right? Like 20,000 A-one hundreds. That's like, I know it sounds like a lot of money.”
Dylan Patel Dec 5, 2023 ▶ 4:34 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 bullish
Prediction Not checkable as stated
Patel: Roughly 10 companies have compute to beat GPT-4 within six months
“There's, like, 10 companies that have enough compute in one single data center to be able to beat GPT-IV, right? Like, straight up, like, if not today, within the next six months, right?”
Dylan Patel Dec 5, 2023 ▶ 33:51 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Jan 11, 2024 neutral
Opinion
Lambert: Frontier Labs Lack Visibility into Cross-Model RLHF Sensitivity
“I think big labs are so over-indexed, are indexed on their own base models, so they don't know, like, what's swapping between CloudBase or GPT-IV-Base, how that would change any notion of preference or what you do with RLHF.”
Nathan Lambert Jan 11, 2024 ▶ 1:30:15 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Assertion Supported
Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans
“Essentially, people also think that synthetic data is, like, GPT-IV is more accurate than humans at labeling preferences, so if you look at these diagrams, like, humans are about 60 to 70% agreement, or, like, that's what the models get to, and if humans are a…”
Nathan Lambert Jan 11, 2024 ▶ 48:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Nathan Lambert Jan 11, 2024 ▶ 1:26:04 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Feb 7, 2024 negative
Assertion Partly supported
Retool survey found GPT-4V NPS was around 14 vs 45 for GPT-4
“GPT-IV.V. MPS, I want to say it was, like, 14 or something, like, it was, like, not high, actually. But the GPT-IV MPS thing was, like, 45 or something like that”
David Hsu Feb 7, 2024 ▶ 51:30 The State of AI in production — with David Hsu of Retool
Apr 6, 2024
Assertion Not checkable as stated
Murphy: Azure OpenAI significantly beats OpenAI's hosted API 400-600ms latency.
“And then GPD, 3.5 turbo or four, you probably get, you know, 400, maybe 600 milliseconds of latency in their hosted API. And if you go into Azure and you use their services, you can get that down a lot lower.”
Damien Murphy Apr 6, 2024 ▶ 10:17 Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Apr 11, 2024
Opinion
Byun: GPT-3 was a qualitative shift, while GPT-4 was an extension
“I think GPT-III was a big change because it kind of said, oh, now is the time to build to you that we can use AI to build these tools. And then GPT-IV was maybe a little bit more of an extension of GPT-III. It felt less like a level, GPT-III over GPT-II was li…”
Jungwon Byun Apr 11, 2024 ▶ 26:47 Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
Apr 24, 2024 negative
Opinion
Liu: GPT-4 and Claude 3 Opus still write poor quality essays
“Or those are two sort of systems that I wish you before or Opus was actually good enough to just write me an essay, but most of the essays are still pretty bad.”
Jason Liu Apr 24, 2024 ▶ 49:50 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Apr 27, 2024 positive
Opinion
Claude 3 can bypass its assistant persona to expose the underlying simulator
“Instead of having this entity, like GPT-IV, that's an assistant that just pops up in your face that you have to kind of, like, punch your way through and continue to have to deal with as a headache, instead, there's ways to kindly coax Claude into having the a…”
Karan Malhotra Apr 27, 2024 ▶ 8:37 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Jul 5, 2024 negative
Insight
Yi Tay: Distilled open-source model variants disappeared after failing to climb LMSYS
“When people realize that, like, this, like, turning on the GPT-IV tab and running some DPO is not going to give them the reward signal that they want anymore, right? Then all these variants gone, right? You know, there was this era where there's, wow, there's …”
Yi Tay Jul 5, 2024 ▶ 2:01:00 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Jul 23, 2024 positive
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Aug 28, 2024 negative
Insight
Carlini: GPT-4 would exist identically without adversarial machine learning research
“Nothing about GPT-IV would be at all different if the field of, like the entire field of Everson machine learning disappeared. Like everything to do with Everson examples, like all of the, like for the most part, like GPT-IV would exist identically.”
Nicholas Carlini Aug 28, 2024 ▶ 58:56 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Sep 20, 2024 neutral
Assertion Not checkable as stated
Schulhoff: GPT-4 Fails to Output Reasoning on 1 in 100 to 1,000 Prompts
“I remember I did a lot of experiments with GPT-IV, and especially when you look at it at scale, so I'll run thousands of prompts against it through the API, and I'll see, you know, every one in a hundred, every one in a thousand outputs no reasoning whatsoever…”
Sander Schulhoff Sep 20, 2024 ▶ 33:03 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Sep 27, 2024 negative
Opinion
Yao: Text adventure games remain very hard even for GPT-4
“Like those texting are just too hard. I think today it's still very hard. Like if you used to be before to solve it, it's still very hard.”
Shunyu Yao Sep 27, 2024 ▶ 8:29 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Oct 4, 2024 bullish
Assertion Supported
Swyx: Distilling GPT-4 to Mini cut costs 15x with 2% hit
“Yeah, I sat in the distillation session just now, and they showed how they distilled from four to four mini, and it was like only like a two percent hit in the performance, and 15 X cheaper.”
Shawn Wang Oct 4, 2024 ▶ 48:14 Building AGI in Real Time (OpenAI Dev Day 2024)
Oct 19, 2024 neutral
Assertion Supported
Hu: SWE-bench scores jumped from 3% on GPT-4 RAG to 43%
“If we just use GPT-IV plus RAG, what do we get? It's, like, a measly three percent. And then up to the most recent submissions where they get up to 43%.”
Jesse Hu Oct 19, 2024 ▶ 7:40 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Oct 19, 2024 positive
Insight
Jesse Hu: SWE-bench gains over baseline GPT-4 come entirely from agent scaffolding
“The diff between that and something like Devin is all in like, sort of like the agent scaffold or the agent code, right? So that's, what's really exciting about this stuff. It shows off what you can do just from prompting and just from adding tools.”
Jesse Hu Oct 19, 2024 ▶ 12:09 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Nov 11, 2024
Assertion Partly supported
Polu: GPT-4 was ready internally at OpenAI months before September 2022
“I had seen GPT-IV internally at the time. It was September, 20, 22. So it was pre-chat GPT, but GPT-IV was ready since, I mean, I'd been ready for a few months internally.”
Stanislas Polu Nov 11, 2024 ▶ 19:16 Agents @ Work: Dust.tt — with Stanislas Polu
Nov 29, 2024 bullish
Assertion Partly supported
Eugene Yan: Ensembled LLMs outperform standalone GPT-4 for evaluation tasks
“So in this paper here by Kohir, what they did was they have a reference model, and this reference model is GPT-IV. And then essentially what they did was the ensemble command R, Haiku, and GPT-IV. And I can't remember what the I think the ensemble was just maj…”
Eugene Yan Nov 29, 2024 ▶ 43:37 [Paper Club] DocETL: Agentic Query Rewriting + Eval for Complex Document Processing w Shreya Shankar
Jan 1, 2025 bearish
Prediction Not checkable as stated
Swyx: AI models hit a 2T parameter wall, won't reach 10T
“As far as everyone is concerned, Claude, you know, Opus 3.5 is not coming out. GPT 4.5 is not coming out. And Gemini two, like we don't have pro whatever we've hit that wall, whatever that wall is. Maybe I'll call it like the two trillion parameter wall. Like …”
Shawn Wang Jan 1, 2025 ▶ 40:13 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jan 28, 2025 neutral
Insight
OpenAI o1 struggles with multi-step agentic tasks compared to GPT-4
“O-One is very good at programming, but it's kind of, the agent part was the harder part to get it to do here. I think it's like less trained To take the next step in like an agentic task, whereas GPT-IV for like the last two years has been really, you know, pr…”
Shawn Lewis Jan 28, 2025 ▶ 16:09 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Feb 1, 2025 negative
Insight
Nguyen: AI model card benchmark numbers are never apples-to-apples across labs
“None of the numbers are, like, apples to apples. So you actually need to, like, go back to, like, I don't know, like, GPT-E for model card and, like, read the appendix just to, like, make sure that, like, The settings were the same as you're running the settin…”
Karina Nguyen Feb 1, 2025 ▶ 15:11 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Mar 7, 2025 positive
Disclosure
Gemini 1 Proved GPT-4-Level Models Can Bootstrap Reinforcement Learning
“Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were basically at GPT-IV level or above, were capable en…”
Misha Laskin Mar 7, 2025 ▶ 4:34 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Apr 27, 2025 neutral
Assertion Supported
Claude 3.5 wrote fire-and-forget code while GPT-4 used defensive programming
“Claude, for instance, the Sonnet 3.5 was very much fire and forget. It would write code in a kind of Pythonic way, just like, let it fail. Don't be careful about it. Whereas GPT four would use defensive programming, use self assertions.”
Jack Hopkins Apr 27, 2025 ▶ 16:27 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Jul 5, 2025 bullish
Assertion Partly supported
Swix: AI Inference Costs for Fixed Intelligence Fall 100x Annually
“The cost of intelligence for a given set of intelligence, let's say GPT-IV, let's say O-one, whatever, it is literally falling a hundred X over the course of one year.”
Shawn Wang Jul 5, 2025 ▶ 24:27 ⚡️Anthropic vs Cognition on Multi-Agents: A Breakdown with Dylan Davis
Jul 11, 2025 neutral
Opinion
Hsu: Real-World Inertia Leaves AI's Daily Impact Outside Bay Area Near Zero
“And I think if you, like, go to another state outside of the Bay Area, probably even in California, outside of the Bay Area, and then you ask somebody how much their life has materially changed, it's, like, pretty close to zero. Real-world inertia is enormous.…”
Andrew Hsu Jul 11, 2025 ▶ 1:02:23 Personalized AI Language Education — with Andrew Hsu, Speak
Jul 11, 2025 neutral
Assertion Not checkable as stated
Hsu: Early Speak GPT-4 role-play users submitted questionable custom scenarios
“In 2023, when we first launched our AI role plays using GPT-IV, back then people were way more concerned about safety, right? And obviously the models now are much better at like refusals and line sharper between what's appropriate and not. But we did see a lo…”
Andrew Hsu Jul 11, 2025 ▶ 1:00:14 Personalized AI Language Education — with Andrew Hsu, Speak
Jul 31, 2025 positive
Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Nathan Lambert Jul 31, 2025 ▶ 1:16:23 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Jul 31, 2025 bullish
Insight
Lambert: Custom personality fine-tuning is open source AI's winning turf
“If open models are to win, part of it could be just, like, everybody can have exactly the model they want. We're serving GPT-IV. It's kind of its thing. You can prompt it, but if fine-tuning is more effective than prompting, everybody can have the model. That …”
Nathan Lambert Jul 31, 2025 ▶ 1:05:05 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Aug 15, 2025
Assertion Not checkable as stated
Brockman: GPT-4 handled multi-turn chat without being trained on it
“We actually did a instruction following post-train on it, so it was really just a data set that was, here's a query, here's what the model completion should be, and I remember that we were like, well, what happens if you just follow up with another query? And …”
Greg Brockman Aug 15, 2025 ▶ 1:32 Greg Brockman on OpenAI's Road to AGI
Nov 22, 2025 positive
Assertion Contradicted
Wagner claims Flux shipped embedded AI chat before GPT-4 released
“So I think I'm going to claim here, I think we were the first engineering tool or design tool that had an AI chat in it. We shipped that I think a month or two months before GPT-IV became publicly available.”
Matthias Wagner Nov 22, 2025 ▶ 4:18 ⚡️ Building the AI Hardware Engineer with Matthias Wagner, Co-founder of Flux
Jan 9, 2026 negative
Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-Smith Jan 9, 2026 ▶ 8:36 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 positive
Assertion Supported
GPT-4-level intelligence is now over 100 times cheaper than at launch
“The, like, one fact on that is that you can get intelligence at the level of GPT-IV for over a hundred times cheaper than GPT-IV was at launch right now.”
Micah Hill-Smith Jan 9, 2026 ▶ 58:55 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 28, 2026 neutral
Disclosure
Andrew White Began Red-Teaming GPT-4 in August 2022
“And so I was a red teamer for GPT-IV and I was using it like nine months with me for release with August. So GPT-IV came out in March and I was using it in August.”
Andrew White Jan 28, 2026 ▶ 9:13 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Apr 7, 2026 positive
Insight
Lopopolo: Reasoning models eliminate need for rigid state-machine scaffolding
“And this I think is like the fundamental difference between reasoning models and the four ones and four O's of the past where these models could not think. So you kind of had to put them in boxes with a predefined set of state transitions. Whereas here we have…”
Ryan Lopopolo Apr 7, 2026 ▶ 11:59 Extreme Harness Engineering: 1M LOC, 1B toks/day, 0% human code or review — Ryan Lopopolo, OpenAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.