AI progression has shifted from post-GPT-4 to a coding agent era
“There's sort of a post GPT-IV era, and then there's a more recent, like, post wide adoption of coding agents era, and then probably soon there's going to be, you know, additional eras, and things are going quite a bit faster, and development is going, you know…”
Altman: I was wrong to expect immediate software disruption after GPT-4
“When we got to GBT four, which was back in 20, 23, I think that very quickly after that, there was going to be much more disruption in software business being up for grabs right, right away than turned out to be. And the thing that I think I was wrong about a …”
Altman: OpenAI gained conviction in reasoning from GPT-4, not 3.5
“I would say we got real conviction with GPT-IV. Not even 3.5.”
Brockman: GPT-4 met previous AGI criteria despite clearly not being AGI
“It clearly wasn't an AGI. It was lacking something, but just if you'd describe your criteria for AGI two months prior, it probably would have been compatible with what GPT-IV was”
Lopopolo: Reasoning models eliminate need for rigid state-machine scaffolding
“And this I think is like the fundamental difference between reasoning models and the four ones and four O's of the past where these models could not think. So you kind of had to put them in boxes with a predefined set of state transitions. Whereas here we have…”
Turley: Giving flagship AI model away for free boosted revenue and retention
“So we just gave it away for free and that ended up being totally revenue positive and retention positive because it just provided access to the tech.”
Turley: Raw GPT-4 Did Not Impress Anyone Before Post-Training
“A few weeks or so after I joined OpenAI GPT-IV had finished training, and I remember trying it out, and it actually It didn't impress me at all nor anyone else that week, because it kind of didn't work, and it's because we hadn't figured out how to post train …”
Simon Last: GPT-4 Proto-Interface Sparked Notion's AI Pivot
“It wasn't until I played with GPT-IV that it became really, really real. So, you know, we, when we got access to it was sort of like a proto-ChatGPT-like interface. And my co-founder Ivan and I both got access, and it was just immediately clear, like, I would …”
Gil: GPT-4-level token pricing dropped 150x in 21 months
“We looked at the cost of a GPT-IV level or equivalent model. We looked at that a year or two ago and basically in 21 months, it went from like 37 bucks for a million tokens to 25 cents. And so, you know, pricing dropped by a 150 X in 21 months.”
Acharya: Token cost for GPT-4 has dropped 100x since release
“The cost of actually a token on GPT-IV has, you know, gone down a hundred X since the model was released.”
Andrew White Began Red-Teaming GPT-4 in August 2022
“And so I was a red teamer for GPT-IV and I was using it like nine months with me for release with August. So GPT-IV came out in March and I was using it in August.”
Hoffman: Gates demanded GPT pass AP Biology to prove its utility
“We'd gotten the demand from him to say, man, challenge to say, hey, I'll believe you have general purpose, you know, and that's not AGI, but it's like a good general purpose functional technology. If without special training, you can pass the AP bio exam at a …”
Weinberg: Consumer AI is plateauing because reasoning is already sufficient
“So I think that we're seeing a plateau in performance for consumer use cases. And the reason why I think this is like a misnomer or something that people actually shouldn't pay attention to is we don't need them to be better for consumer use cases. Like, I fee…”
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
GPT-4-level intelligence is now over 100 times cheaper than at launch
“The, like, one fact on that is that you can get intelligence at the level of GPT-IV for over a hundred times cheaper than GPT-IV was at launch right now.”
Gil: GPT-4 equivalent token costs dropped over 100x in 2024
“In 20, 24, the cost of GPT-IV equivalent models, if you look at a million tokens, it came down over a hundred X. You know, so many of my team did this analysis to show that.”
Tan: GPT-4 wiped out YC startups' RL fine-tuning lead over GPT-3.5
“You know, we've also had YC companies where, ah, they had something that beat OpenAI, ah, you know, GPT-III.V and they were doing fine-tuning with RL, but then, ah, yeah, GPT-IV and then, ah, GPT-IV came out and, ah, you know, basically blew their fine-tuning …”
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Clark: Thrive received a GPT-4 demo in November 2022
“We actually got a demo of GBT-IV in November of twenty-twenty-two before it was publicly released in March, and it was one of the most astounding pieces of technology we'd ever seen.”
Wagner claims Flux shipped embedded AI chat before GPT-4 released
“So I think I'm going to claim here, I think we were the first engineering tool or design tool that had an AI chat in it. We shipped that I think a month or two months before GPT-IV became publicly available.”
Nadella: Bill Gates was skeptical of Microsoft's initial $1B OpenAI investment
“Obviously you go to the board and say, hey, I have an idea of taking a billion dollars and giving it to this crazy structure, which we don't even kind of understand. What is it? It's a nonprofit, blah, blah, blah. And saying, go for it. There was a debate. Bil…”
Heller: Casetext received early access to GPT-4 in summer 2022
“Because we were so focused on large language models and were researching deeply in this space, we got really early access to GPT-IV. Like summer, 20, 22.”
Masad: GPT-5 regressed in human tone compared to GPT-4
“My feeling is that you know, GPT-Five got good at verifiable domains. It didn't feel that much better at anything else. The more human angle of it felt like it regressed”
Masad: GPT-5 shows no reasoning progress on open-ended controversial topics
“Go you know, dig up GPT-IV or other models and go to GPT-V. You're not gonna find that much difference of, okay, let's reason together. Let's try to figure out what was the origins of COVID. Because it's still an unanswered question, you know? And I don't see …”
Tan: Reliable LLM output requires decomposing tasks into small contextual steps
“At that moment if you chopped it down to a bite-sized chunk, like you gave it some amount of context, That a human being, given the same context and the same prompt, would answer in a certain way. He found that he could, ah, you know, given inputs and outputs,…”
Tworek: OpenAI team was initially underwhelmed by pre-trained GPT-4
“When we trained GPT-IV, we were pretty underwhelmed internally, and then there was a lot of moments, oh, we trained this small, we spent a lot of money on it, and it's kind of like, you know, pretty dumb, at least, like, you know, we have GPT-IV, GPT-III alrea…”
Labenz: GPT-4 to GPT-5 capability leap matches GPT-3 to GPT-4
“And if you look back to GPT three, you know, there's a huge leap. I would contend that the leap is similar from GPT four to five.”
Labenz: Pure reasoning AI models achieved IMO gold without external tools
“Well, I mean, a big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models with no access to tools from multiple companies. And, you know, that is night and day compared to what GPT-IV could do with math, right?”
Labenz: Per-token model costs fell 95% from GPT-4 to GPT-5
“It's like 90 it's like a 95% discount from GPT-IV to GPT-V.”
Altman: OpenAI knows how to build a GPT-4 equivalent for video
“It was really not until GPT-IV where these text models started providing real value for people, and we know how to go make the GPT-IV equivalent of video models, and we will do that, and then a lot of these things that are currently annoying, like doors, or, y…”
Patel: GPT-4 Turbo was less than half the size of GPT-4
“Four to four turbo was, like, the model was less than half the size. And four turbo to four O, like four O's cost is way lower than four. And they just kept shrinking the cost.”
Wu: GPT-4 already existed internally at OpenAI by September 2022
“The first one was right when I joined the company in September, 20, 22. We, it was pre-TiGPT. But at the time, GPT-IV already existed internally.”
Brockman: GPT-4 handled multi-turn chat without being trained on it
“We actually did a instruction following post-train on it, so it was really just a data set that was, here's a query, here's what the model completion should be, and I remember that we were like, well, what happens if you just follow up with another query? And …”
Altman: GPT-5 will integrate all reasoning models into a single interface
“It's an integrated model, so you don't have to, like, pick in our model switcher and know if you should use GPT-IV or O-III or O-IV or any of the complicated things. It's just one thing that, that works”
Andreessen: Google could have built a GPT-4-level ChatGPT by 2019
“Google developed the transformer in 2017. And then they basically let it sit on the shelf, right? Cause it was a research project. They didn't productize it. They were very worried about, you know, from people I've talked to, they were very worried about the, …”
Turley: Pre-ChatGPT, OpenAI shipped models like hardware instead of software
“Treating the model as a product was not a thing before ChatGPT, because we would ship it more like hardware, where, you know, there'd be a release like GPT-III, and then we would start working on GPT-IV, and these weird giant Big spend R&D projects that would …”
Andreessen: Google Could Have Built a GPT-4 Level Chatbot by 2019
“I talked to somebody senior who was there at the time and I asked them, you know, when could you have had chat GPT with GPT four level output if you had just gotten, you know, gone, gone flat out starting in 2017. And they said by 2019.”
Kim: The leap from GPT-4 to GPT-5 is OpenAI's most impressive yet
“Maybe I'm biased, recency biased, but I think to jump to four to five is most impressive for me, because I guess with 3.5 when we first released it, the most common use case for me then also was still just for coding. And, but now, like, Even though four was b…”
Lightcap: GPT-5 bakes in tool use and longer-horizon reasoning
“So using tools, for example, is something that really thinks really important for overall intelligence, GPT two and three couldn't really do that as well. GPT-IV could do it in a more nascent way. And now GPT-V, you get that baked in with the benefit of these …”
Lambert: Custom personality fine-tuning is open source AI's winning turf
“If open models are to win, part of it could be just, like, everybody can have exactly the model they want. We're serving GPT-IV. It's kind of its thing. You can prompt it, but if fine-tuning is more effective than prompting, everybody can have the model. That …”
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Gerstner: AI models shift from internet compression to tool-using reasoning engines
“So you go back to GPT-IV, you know, and you're compressing, you know, the entire internet. But now we really don't need to do it because we've trained them to use tools like the internet, right? They're true reasoning engines.”
Hsu: Early Speak GPT-4 role-play users submitted questionable custom scenarios
“In 2023, when we first launched our AI role plays using GPT-IV, back then people were way more concerned about safety, right? And obviously the models now are much better at like refusals and line sharper between what's appropriate and not. But we did see a lo…”
Hsu: Real-World Inertia Leaves AI's Daily Impact Outside Bay Area Near Zero
“And I think if you, like, go to another state outside of the Bay Area, probably even in California, outside of the Bay Area, and then you ask somebody how much their life has materially changed, it's, like, pretty close to zero. Real-world inertia is enormous.…”
Swix: AI Inference Costs for Fixed Intelligence Fall 100x Annually
“The cost of intelligence for a given set of intelligence, let's say GPT-IV, let's say O-one, whatever, it is literally falling a hundred X over the course of one year.”
Schulhoff: Explicit chain-of-thought prompting is still needed for GPT-4 and GPT-4o
“Actually for those models, I'd say no need, but if you're using GPT-IV, GPT-IV-O, then it's still worth it.”
Patel: GPT-4-Level Training Costs Have Dropped 10x to 100x
“If you look at what it costs to train GBT for originally, I think it was like 20,008, 100 over the course of a hundred days. So I think it costs on the order of like half a million to a hundred million dollars, somewhere in that range. And I think you could tr…”
Douglas: Generalist models will obsolete specialized fine-tuned models
“I really do think that similar to how we saw with large pre-trained models before with small fine-tuned models made it like, had gains over the sort of GPT-II era, but then were obsoleted by GPT-IV being generally good at everything. I think, to be honest, you…”
Duolingo was an launch partner for OpenAI's GPT-4 release
“We were one of the launch partners of OpenAI when they first launched GPT-IV”
Marcus: Models like o1 are not systematically better than GPT-4
“Models like O-I are not systematically better than GPT-IV. There's, they're better in certain use cases. Especially ones where you can create data in advance.”
Tan: Casetext founder Jake Heller first commercialized GPT-4 in legal
“Garry Tan: He was one of the first people to get access to GPT-IV, and we think of him at YC as the first man on the moon, and that he was the first to successfully commercialize GPT-IV in the legal space.”
Claude 3.5 wrote fire-and-forget code while GPT-4 used defensive programming
“Claude, for instance, the Sonnet 3.5 was very much fire and forget. It would write code in a kind of Pythonic way, just like, let it fail. Don't be careful about it. Whereas GPT four would use defensive programming, use self assertions.”
Will Brown: Morgan Stanley deployed integrations on day GPT-4 launched
“Like, the day GPT-IV launched, Morgan Stanley had integrations, because we had been working on it, and these were, like, we had press releases for these, like, we are ready to go”
Srinivas: Most ChatGPT users do not know o1 or GPT-4 differences
“In fact, most people using tracks between the world don't even know there's a model called O one or O three and don't even know what the difference is for GPT four.”
Srinivas: Launching Perplexity around GPT-4 would have been too late
“You want to build at the right moment where, like, kind of, like, how we built perplexity around the GPT 3.5 time, not GPT four time. It would have been too late then.”
Blomfield: GPT-4 Underperformed on Coding by Misimplementing and Asking Too Many Questions
“I tried GPT-IV just a couple of days ago, and honestly, I wasn't yet as impressed. It just came back with me with too many questions and actually got the implementation wrong too many times.”
Patel: Model inference costs dropped 60x from GPT-4 to DeepSeek-V3
“And likewise, when we look at from GPT-IV to DeepSeq VIII it's fallen roughly 600 X in cost. Right. So we're not quite at that 1200 X, but it has fallen 600 X in cost from 60 dollars to less than you know, to about a dollar. Right. Or to less than a dollar. So…”
Patel: Frontier AI cluster costs have scaled from $100M to $10B
“For GPT-IV, it was a few hundred million dollars and it's one building full of GPUs, too. GPT-IV 4.5 and the reasoning models, like, oh, one, oh, three were done in a, in three buildings on the same site, and, you know, billions of dollars to, hey, these next …”
Suleyman: GPT-4 and GPT-4o Efficiency Will Improve 100x
“So I expect that to happen for GPT-IV, GPT-IV-O and all of the other models down the road.”
Chollet: GPT-4 lacks fluid intelligence, but OpenAI's o3 model has it
“GPT-IV does not have fluid intelligence, for instance, but O-III does.”