why aren't all 12 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Gil: GPT-4 equivalent token costs dropped over 100x in 2024
“In 20, 24, the cost of GPT-IV equivalent models, if you look at a million tokens, it came down over a hundred X. You know, so many of my team did this analysis to show that.”
Assertion Supported
Gil: GPT-4 equivalent inference costs dropped 180x over 18 months
“Somebody on my team kind of worked out that in the last 18 months we saw a 180 X decrease in cost per token for equivalent level models. 180 X, not 180%, 180 times.”
Assertion Supported
Gil: GPT-4 Equivalent Token Costs Dropped 240x in 18 Months
“Over the last 18 months or so, the cost of a million tokens going into a GPT-IV equivalent model has basically dropped 240 X.”
Assertion Partly supported
Liberty: RAG Over Internet Data Reduces LLM Hallucinations by 50%
“And you could see that if you augment all of them with RAG on, even on the internet, which is data that they were trained on, you can reduce hallucinations significantly up to 50% sometimes.”
Assertion Supported
ChatGPT launched as a 10-month-old model with RLHF as a practice test
“ChatGPT was a 10 month old model with a little bit of RLHF on top of it. And, you know, like by, you know, admission, like You know, not a beautiful user interface. It was just sort of a way to get something out there because you know, you needed some practice…”
Assertion Supported
Gil: GPT-4-level token pricing dropped 150x in 21 months
“We looked at the cost of a GPT-IV level or equivalent model. We looked at that a year or two ago and basically in 21 months, it went from like 37 bucks for a million tokens to 25 cents. And so, you know, pricing dropped by a 150 X in 21 months.”
Assertion Supported
Alexandr Wang: JPMorgan has 150 petabytes of data; GPT-4 used under one.
“JP Morgan's proprietary data set is a 150 petabytes of data. GPT-IV is trained on less than one petabyte. Of data.”
Prediction Held up
Guo: 2024 will end with a handful of GPT-4-level models
“I think it's very it's very likely at this point that you end this year with a handful of GPT-IV level models, and that some of those are open source, right?”
Assertion Supported
Gil: Mistral reached near GPT-4 capability within one year of founding
“They went from basically starting the company to almost GPT-IV level in less than a year.”
Assertion Supported
Khan: OpenAI approached Khan Academy before finishing GPT-4's first training run
“They hadn't even finished the first training run of GPT-IV, but they said, you know, we think it's going to be done in about two weeks, and we think This is going to be the model that really wakes up people to the power of generative AI. We want two reasons wh…”
Assertion Supported
Sankar: OpenAI's GPT-4 is currently unavailable in classified government environments
“We live in a world where we can't count on GPT-IV everywhere. Like we don't have that on classified environments, right?”
Assertion Supported
Guo: No open-source model matches GPT-3.5, GPT-4, or Claude quality
“So there's nothing out there today in open source that is like GPT four, three, five or anthropic cloud quality, right? So there is a, there's one player out in front and that's open AI”