why aren't all 7 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Prediction Held up
Ratner: Private data models will exceed closed models in specialized tasks
“Closed source models, like a GPT-IV, five, six, seven, whatever comes, are going to be very hard to match in terms of generalist capability for, say, consumer use cases that are reflected in the web data they're trained on and the flywheels that get powered by…”
Assertion Supported
Azhar: Mistral matched GPT-4 quality far more computationally efficiently than US firms
“Mistral, which is this Parisian company has been doing some, you know, remarkable things, had done a couple, two things that I thought were really interesting. One was that they were able to get close to GPT-IV quality much more computationally efficiently tha…”
Assertion Supported
Shah: Hippocratic AI outperformed GPT-4 on 105 of 114 healthcare exams
“And then we took it, and then we had GPT-IV take it, and we had all the other language models take it, and we beat them all. And we beat them on a 105 of a 114 for GPT-IV, for example.”
Assertion Supported
Nvidia released an AI model that outperformed GPT-4
“Just a couple of weeks ago, Nvidia announced that In addition to having the best chips, they just released a model that was actually better than GPT-IV.”
Assertion Supported
GPT-4 price per token dropped roughly 90% in one year
“The, I think the price per token of GPT-IV dropped something like 90%. Over the last year”
Assertion Supported
Traynor: GPT-4 crossed the hallucination threshold required for customer support bots
“So Finn's built on GPT-IV, by the way, we've, we tried, we wanted to build on a three, 3.5, but it didn't, it's still, I remember back when we used to talk about hallucinations, like four was the sort of the perceptual change for us in terms of trust and relia…”