why aren't all 13 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Sivulka: GPT-4 beat BloombergGPT at every financial task
“So then GPT-IV was released, I think, like a few weeks later. I don't know exactly any of the right timeline, but it just destroyed Bloomberg GPT at every single finance task.”
Prediction Held up
Altman: Competing AI labs will successfully replicate OpenAI's o1 model
“After, after a research lab does something, even if you don't know exactly how they did it, it's, I won't say easy, but it's doable to go off and copy it, and you can see this in the replications of GPT-IV, and I'm sure you'll see this in replications of O-one…”
Assertion Partly supported
Gomez: 13-billion-parameter models now outperform original 1.7-trillion GPT-4
“GPT-IV, if it's true what they say, and it's 1.7 trillion parameters, this big MOE, we have models that are better than that model that are like, thirteen billion parameters.”
Assertion Partly supported
Wang: JP Morgan's internal data is 150 petabytes versus GPT-4's sub-petabyte dataset
“JP Morgan's proprietary internal data set is a 150 petabytes. The GPT-IV was trained on an internet data set that was less than one petabyte.”
Assertion Partly supported
Srinivas: GPT-4 convincingly beats BloombergGPT on finance benchmarks
“Bloomberg spent a lot of money training Bloomberg GPT... And that model is, is beaten convincingly by, like, a GPT-IV on all the finance benchmarks.”
Prediction Didn’t hold up
Socher: An open-source GPT-4 equivalent will arrive by end of 2023
“I predicted that we'll have a GPT-IV equivalent model Before the end of the year, that's open source.”
Prediction Didn’t hold up
Socher: Open-source GPT-4 equivalent model will launch before end of 2023
“I predicted that we'll have a GBD four equivalent model before the end of the year. That's open source. Of course, GBD four keeps getting better and better. So my prediction was for the version we had like a few months ago,”
Assertion Supported
Acharya: Token cost for GPT-4 has dropped 100x since release
“The cost of actually a token on GPT-IV has, you know, gone down a hundred X since the model was released.”
Assertion Supported
Mollick: Unprompted GPT-4 math tutoring led to lower test scores
“The first randomized control trial we have, I have some of my colleagues at Wharton was giving GPT-IV people for math tutoring in Turkey. Now, they didn't do a huge amount of, like, you know, it was an assigned class, and they used the system, but it turns out…”
Assertion Supported
Hankes: OpenAI Delayed GPT-4 Release by Five Months for Safety Testing
“On GPT four, we saw the demo in the fall. They didn't just release the product then they took, I think four or five months To test and learn about safety in the edges of the model and then ultimately released it to the world.”
Assertion Supported
Rauch: Llama is nowhere near as capable as GPT-4
“Right now, Llama is not as good as GPT-IV. Not even close.”
Assertion Supported
Lebrun: LIMA fine-tuned on 1,000 examples beats GPT-3, rivals GPT-4
“Three weeks ago, there was a paper about Lima. So, so this Lima paper shows that with only 1000 question and answer examples, so very, very small data sets they get something for, use for fine tuning, so the second stage, they get something that performs bette…”
Assertion Supported
Duolingo was an launch partner for OpenAI's GPT-4 release
“We were one of the launch partners of OpenAI when they first launched GPT-IV”