why aren't all 25 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Chase: Open-source models still lag behind Claude 3 and GPT-4
“Like there's, I think we see increasingly interest in open source, but the reasoning abilities are still just like lagging behind Cloud three or GPT four. And I think like for a lot of the applications that it kind of, it probably depends on the types of appli…”
Assertion Not checkable as stated
GPT-4 Class Enterprise Models Now Cost $10M to $20M to Train
“What we've seen is today you can build a model that's as good as GPT-IV in all the things that enterprises might care about. For ten million dollars, twenty million dollars, like just orders of magnitude less than what was spent to develop that model. And so i…”
Insight
Alexandr Wang: Data abundance is the fundamental bottleneck for post-GPT-4 models.
“The key to the scaling of these large language models and the, you know, these language models in general is the ability to scale data. And I think that one of the fundamental bottlenecks to, you know, what's, what's in the way of us getting from GPT-IV to GPT…”
Insight
Gil: Existing models yield 10x to 100x gains without base model scaling
“In order to 10 X or even a hundred X use cases and usages for AI outside of that, there's things that could just be done on existing models today. So you don't need to wait for GPT seven or whatever. You could start with GPT four or GPT 3.5 and add these thing…”
Prediction Not checkable as stated
Rauch: AI could operate in a purer logic layer before compiling to code
“So I definitely believe that AIs could operate in a more pure layer. Of logic that then gets converted and mapped back to whatever problem at hand that you have.”
Assertion Supported
Gil: GPT-4 equivalent token costs dropped over 100x in 2024
“In 20, 24, the cost of GPT-IV equivalent models, if you look at a million tokens, it came down over a hundred X. You know, so many of my team did this analysis to show that.”
Assertion Supported
Gil: GPT-4 equivalent inference costs dropped 180x over 18 months
“Somebody on my team kind of worked out that in the last 18 months we saw a 180 X decrease in cost per token for equivalent level models. 180 X, not 180%, 180 times.”
Assertion Supported
Gil: GPT-4 Equivalent Token Costs Dropped 240x in 18 Months
“Over the last 18 months or so, the cost of a million tokens going into a GPT-IV equivalent model has basically dropped 240 X.”
Assertion Not checkable as stated
Knoop: Frontier AI labs have stopped publishing technical details
“Frontier AI research is also basically like completely stopped publishing. You know, the GPD four paper had zero technical details. The Gemini paper had zero technical details on the longer context stuff.”
Opinion
Alexandr Wang: GPT-4 was too early a model to sustain application hype.
“GPT-IV, I think, as a model, was a little early of a technology for us to have this entire hype wave around, and I think we, you know, the community very quickly discovered all the limitations of GPT-IV... It was probably a few generations too early of a model…”
Assertion Partly supported
Liberty: RAG Over Internet Data Reduces LLM Hallucinations by 50%
“And you could see that if you augment all of them with RAG on, even on the internet, which is data that they were trained on, you can reduce hallucinations significantly up to 50% sometimes.”
Disclosure
Zhao: Notion bet the company on AI after seeing GPT-4
“And after that, we sort of just bet the company on it.”
Assertion Supported
ChatGPT launched as a 10-month-old model with RLHF as a practice test
“ChatGPT was a 10 month old model with a little bit of RLHF on top of it. And, you know, like by, you know, admission, like You know, not a beautiful user interface. It was just sort of a way to get something out there because you know, you needed some practice…”
Assertion Supported
Gil: GPT-4-level token pricing dropped 150x in 21 months
“We looked at the cost of a GPT-IV level or equivalent model. We looked at that a year or two ago and basically in 21 months, it went from like 37 bucks for a million tokens to 25 cents. And so, you know, pricing dropped by a 150 X in 21 months.”
Insight
Weinberg: Intuition for AI progress comes from hammering models through failures
“It seems like there was this large gap between
I log into ChatGPT or I just use GPT-IV from an API or whatever it is, and I try something a couple times versus I'm going to sit there and just hammer on this until it works, right?
and I think if you do that …”
Prediction Not checkable as stated
Davis: AI systems will make millions of heterogeneous model calls per question
“I think that what we'll see people doing is kind of composing, this sounds funny, but massive networks, Where maybe each stage in the network will basically be maybe some best of A, best of K component with many, many calls to different language models, you kn…”
Assertion Supported
Alexandr Wang: JPMorgan has 150 petabytes of data; GPT-4 used under one.
“JP Morgan's proprietary data set is a 150 petabytes of data. GPT-IV is trained on less than one petabyte. Of data.”
Prediction Held up
Guo: 2024 will end with a handful of GPT-4-level models
“I think it's very it's very likely at this point that you end this year with a handful of GPT-IV level models, and that some of those are open source, right?”
Assertion Supported
Gil: Mistral reached near GPT-4 capability within one year of founding
“They went from basically starting the company to almost GPT-IV level in less than a year.”
Assertion Not checkable as stated
Hoffman: Pre-release GPT-4 gave useless MBA-style advice on AI investing
“When I first got access to GBD-IV you know, months before it was publicly accessed, because I was on the OpenAI board, I sat down and said how can I, Reid Hoffman, make money through investing in artificial intelligence? Because I just wanted to try it. And it…”
Assertion Supported
Khan: OpenAI approached Khan Academy before finishing GPT-4's first training run
“They hadn't even finished the first training run of GPT-IV, but they said, you know, we think it's going to be done in about two weeks, and we think This is going to be the model that really wakes up people to the power of generative AI. We want two reasons wh…”
Assertion Supported
Sankar: OpenAI's GPT-4 is currently unavailable in classified government environments
“We live in a world where we can't count on GPT-IV everywhere. Like we don't have that on classified environments, right?”
Assertion Supported
Guo: No open-source model matches GPT-3.5, GPT-4, or Claude quality
“So there's nothing out there today in open source that is like GPT four, three, five or anthropic cloud quality, right? So there is a, there's one player out in front and that's open AI”
Disclosure
Instacart deployed internal assistant 'Ava' for engineers to use GPT-4
“We've also deployed an internal assistant called Ava so that all of our engineers can leverage the GPT-IV models in a lot of what they're doing, but do that in an assisted way that helps them figure out how to integrate that.”
Disclosure
Simon Last: GPT-4 Proto-Interface Sparked Notion's AI Pivot
“It wasn't until I played with GPT-IV that it became really, really real. So, you know, we, when we got access to it was sort of like a proto-ChatGPT-like interface. And my co-founder Ivan and I both got access, and it was just immediately clear, like, I would …”