why aren't all 37 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Open · timeframe May 2028
Mostaque: AI Will Eliminate All Outsourced Programming Jobs in India
“India, all of the outsourcing jobs in programming will go, because GPT-IV can go level three Google programmer exam and pass it.”
Opinion
Narayanan: GPT-4 passing bar and medical exams meant nothing for actual practice
“So when GPT-IV came out and OpenAI claimed that it passed the bar exam and the medical licensing exam people were very excited slash scared about what this means for doctors and lawyers, and the answer turned out to be approximately nothing, right? Because it'…”
Assertion Not checkable as stated
Narayanan: No AI model has meaningfully surpassed GPT-4 in 18 months
“And what we've seen in the nearly year and a half since GPT-IV came out is that we haven't really had models That have surpassed it in a meaningful way.”
Prediction Not checkable as stated
Narayanan: Skeptical GPT-5 will yield a leap comparable to GPT-4
“Are we going to see a GPT-V that's as big a leap over GPT-V as GPT-V was over GPT-V? I'm frankly skeptical.”
Opinion
Weinberg: Consumer AI is plateauing because reasoning is already sufficient
“So I think that we're seeing a plateau in performance for consumer use cases. And the reason why I think this is like a misnomer or something that people actually shouldn't pay attention to is we don't need them to be better for consumer use cases. Like, I fee…”
Assertion Partly supported
Sivulka: GPT-4 beat BloombergGPT at every financial task
“So then GPT-IV was released, I think, like a few weeks later. I don't know exactly any of the right timeline, but it just destroyed Bloomberg GPT at every single finance task.”
Prediction Held up
Altman: Competing AI labs will successfully replicate OpenAI's o1 model
“After, after a research lab does something, even if you don't know exactly how they did it, it's, I won't say easy, but it's doable to go off and copy it, and you can see this in the replications of GPT-IV, and I'm sure you'll see this in replications of O-one…”
Opinion
Kolter: Open-sourcing models at GPT-4 capability poses low catastrophic risk
“If you look at the current best models that there are right now, so things like GPT-IV, Claude, 3.5 Gemini, things like this, I would not currently be all that nervous about having an open source model that was as capable as these in terms of the catastrophic …”
Assertion Partly supported
Gomez: 13-billion-parameter models now outperform original 1.7-trillion GPT-4
“GPT-IV, if it's true what they say, and it's 1.7 trillion parameters, this big MOE, we have models that are better than that model that are like, thirteen billion parameters.”
Assertion Partly supported
Wang: JP Morgan's internal data is 150 petabytes versus GPT-4's sub-petabyte dataset
“JP Morgan's proprietary internal data set is a 150 petabytes. The GPT-IV was trained on an internet data set that was less than one petabyte.”
Assertion Partly supported
Srinivas: GPT-4 convincingly beats BloombergGPT on finance benchmarks
“Bloomberg spent a lot of money training Bloomberg GPT... And that model is, is beaten convincingly by, like, a GPT-IV on all the finance benchmarks.”
Insight
Lightcap: Many enterprises mistakenly assume GPT-4 is the peak of AI
“A lot of companies think it's static, so a lot of companies think GPT-IV is the best the models will ever get. That's understandable. Every technology they've ever had to adopt has been relatively static.”
Prediction Didn’t hold up
Socher: An open-source GPT-4 equivalent will arrive by end of 2023
“I predicted that we'll have a GPT-IV equivalent model Before the end of the year, that's open source.”
Prediction Not checkable as stated
Douwe Kiela: GPT-4 Will Disrupt Data Annotators Like Mechanical Turk
“One of the use cases I've been seeing now for GPT-IV is actually that people are using it to generate data and then they're training on that data with cheaper models. So GPT-IV might end up disrupting, not knowledge workers necessarily, but it might just disru…”
Prediction Not checkable as stated
Douwe Kiela: AI Model Parameter Sizes Will Stop Growing Rapidly
“I think Sam Altman had this interesting quote where he was saying that he thought models would stop growing in size. GPT-IV kind of hit this ceiling. I think that's probably right, but not really because size doesn't matter. It's just that data size matters ev…”
Prediction Didn’t hold up
Socher: Open-source GPT-4 equivalent model will launch before end of 2023
“I predicted that we'll have a GBD four equivalent model before the end of the year. That's open source. Of course, GBD four keeps getting better and better. So my prediction was for the version we had like a few months ago,”
Assertion Not checkable as stated
Douwe Kiela: GPT-4's coding skills may be inflated by dataset contamination
“Data contamination where a bunch of these language models are trained on the things that they are being evaluated on. So GPT-IV looks like it's an amazing coder, but it might also just be trained on the data that it's evaluated on, which means that it's not ac…”
Prediction Not checkable as stated
Douwe Kiela: GPT-4 will disrupt Mechanical Turk before knowledge workers
“And so, so GPT-IV might end up disrupting, not like knowledge workers necessarily, but it might just disrupt like mechanical Turk and is just a, an annotator on steroids.”
Disclosure
Emad Mostaque Uses GPT-4 as a Personal Therapist
“I already use GPT-IV as a therapist and things like that.”
Assertion Supported
Acharya: Token cost for GPT-4 has dropped 100x since release
“The cost of actually a token on GPT-IV has, you know, gone down a hundred X since the model was released.”
Disclosure
Altman: OpenAI faced severe, unknown technical issues early in GPT-4 development
“Well, when we started working on GPT-IV, there were some issues that caused us a lot of consternation that we really didn't know how to solve. We figured it out, but there was definitely a time period where we just didn't know how we were gonna Do that model.”
Assertion Supported
Mollick: Unprompted GPT-4 math tutoring led to lower test scores
“The first randomized control trial we have, I have some of my colleagues at Wharton was giving GPT-IV people for math tutoring in Turkey. Now, they didn't do a huge amount of, like, you know, it was an assigned class, and they used the system, but it turns out…”
Prediction Not checkable as stated
Mollick: Open-Source GPT-4 Level AI Models Will Become Ubiquitous
“We now have an open source GPT four capable model and it's going to be everywhere.”
Assertion Not checkable as stated
Mollick: Only 5-10% of professionals have used advanced AI models
“Five to 10% of people in any room, whether, by the way, Silicon Valley, actual people, right, who aren't at a lab, whether that's at a large bank, whether that's at a conference of innovation professionals, maybe five to 10% have used those models, and maybe t…”
Prediction Not checkable as stated
Levie: Users Will Wire Into Specific AI Models Rather Than Abstracted Selection
“I don't think you're going to have complete commoditization of sort of the personality of the models for the sort of style and the response to the point where then, you know, if you're a user, you know, any given response could come from Gemini, Versus GPT-IV …”
Opinion
Kleinerman: Shallow GPT-4 wrappers built in weeks are not real companies
“There are some very shallow wrappers on top of a GPT-IV. I don't place much value on them. I usually ask, hey, how long did it take you to build this? Oftentimes it's a week or two. I don't think there's a company there.”
Opinion
Kleinerman: Deep GPT-4 wrappers with domain expertise build real value
“I do think that there are some very deep wrappers on top of GPT-IV that apply domain specific legal or other domain that I think you end up with a true way to bring GPT-IV to a given market or industry. I think those are value.”
Assertion Supported
Hankes: OpenAI Delayed GPT-4 Release by Five Months for Safety Testing
“On GPT four, we saw the demo in the fall. They didn't just release the product then they took, I think four or five months To test and learn about safety in the edges of the model and then ultimately released it to the world.”
Disclosure
Hankes: Sam Altman Pitched OpenAI Round With Closed GPT-4 Demos
“The way we kicked off the round was he did almost like a closed demo with lots of investors on a couple of calls of the technology they're working on and ultimately in GPT-IV.”
Disclosure
Narayanan: GPT-4's 18-month training period created an illusion of rapid progress
“So I think, like a lot of people, I was fooled by how quickly after GPT-III.V, GPT-IV came out. It was just, you know, three months or so, but it had been in training for 18 months.”
Assertion Not checkable as stated
Alex Wang: GPT-4 was trained on nearly all internet data
“GPT-IV was a model basically trained on nearly all of the internet and using a huge amount of computational capabilities.”
Insight
Frontier AI Model Training Requires Unwritten Art, Not Just Compute
“So it isn't just, you know put in data, right, you know, put in compute, press button, go away for six months, come back, and go, ha ha, I've got GBD-IV. There's a whole bunch of art and science, and what I mean by art is like, well, you train it this way firs…”
Assertion Not checkable as stated
Srinivas: GPT-4 tier models are not yet commoditized
“I think GPT four quality models are not yet commoditized. There's only probably one or two alternatives for the people today, like Claude Opus or some people, Gemini, let's say. If it's just like two or three alternatives, it's not, I wouldn't call it a commod…”
Insight
Söderström: Product designers must understand GPT-4 as deeply as user needs
“Something that designers need to get good at in this world. They need to understand GPT-IV as well as they understand the user.”
Assertion Supported
Rauch: Llama is nowhere near as capable as GPT-4
“Right now, Llama is not as good as GPT-IV. Not even close.”
Assertion Supported
Lebrun: LIMA fine-tuned on 1,000 examples beats GPT-3, rivals GPT-4
“Three weeks ago, there was a paper about Lima. So, so this Lima paper shows that with only 1000 question and answer examples, so very, very small data sets they get something for, use for fine tuning, so the second stage, they get something that performs bette…”
Assertion Supported
Duolingo was an launch partner for OpenAI's GPT-4 release
“We were one of the launch partners of OpenAI when they first launched GPT-IV”