why aren't all 20 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Disclosure
Tworek: OpenAI team was initially underwhelmed by pre-trained GPT-4
“When we trained GPT-IV, we were pretty underwhelmed internally, and then there was a lot of moments, oh, we trained this small, we spent a lot of money on it, and it's kind of like, you know, pretty dumb, at least, like, you know, we have GPT-IV, GPT-III alrea…”
Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Opinion
Srinivas: Building a web search index is harder than competing with GPT-4
“We believe that building your own index is even harder than trying to compete with GPT-IV because it's not a problem that's just solved with money.”
Assertion Not checkable as stated
Chollet: GPT-4 lacks fluid intelligence, but OpenAI's o3 model has it
“GPT-IV does not have fluid intelligence, for instance, but O-III does.”
Prediction Not checkable as stated
UiPath's specialized models will outperform GPT-4 on enterprise document tasks
“But in the end, we will beat, you know, any performance of GPT-IV or whatever, and they would be extremely reliable.”
Prediction Not checkable as stated
Future AI systems will rely on small, dedicated models for routine tasks
“I don't think so, because I think nature is pretty smart. And if you look at how our Brain operates. We have a very good cognitive engine, but we are not using that cognitive engine to do trivial tasks. It's way too expensive. If you want to learn to swim or t…”
Disclosure
DeepScribe sees more impact fine-tuning in-house LLMs than GPT-4
“So with GPT-IV, to be honest, we haven't gotten the impact we'd like in terms of fine-tuning. Where we've seen the most impact is with our own in-house LLM.”
Assertion Contradicted
Lamini's AMD support unlocks 20,000 GPUs, enough to train GPT-4
“So that unlocks, what that means is that unlocks about 20,000 GPUs readily available today for enterprises to be able to use, and to get a sense of what that means you can train GBD-IV.”
Prediction Held up
Ratner: Private data models will exceed closed models in specialized tasks
“Closed source models, like a GPT-IV, five, six, seven, whatever comes, are going to be very hard to match in terms of generalist capability for, say, consumer use cases that are reflected in the web data they're trained on and the flywheels that get powered by…”
Assertion Supported
Azhar: Mistral matched GPT-4 quality far more computationally efficiently than US firms
“Mistral, which is this Parisian company has been doing some, you know, remarkable things, had done a couple, two things that I thought were really interesting. One was that they were able to get close to GPT-IV quality much more computationally efficiently tha…”
Assertion Not checkable as stated
Enterprises prototype with GPT-4 but shift to cheaper self-hosted models for production
“We see quite a bit of usage patterns where people would start GPT-IV for design and then decide to move to Essentially something cheaper, like nixtral self-hosted or nixtral self-hosted when moving to production.”
Assertion Not checkable as stated
Sapoznik: GPT-4 real-time call center suggestions cost $50-$100 per conversation
“In an average conversation, there's thousands of pings to this model. You can say, well, couldn't GPT-IV, for example, make those predictions? It actually could, and it'd probably do a very decent job out of the box. The problem is your conversation will cost …”
Insight
Polu: Enterprise AI needs frontier models rather than complex query routing
“Most of the tasks are pretty general, right? Most of the tasks are pretty like a human would do. And so you just want the best models. And as it happens today, the best models are before enclosed. So that's what you want.”
Assertion Not checkable as stated
Fine-tuned Llama 2 achieves performance comparable to GPT-3.5 and GPT-4
“Straight out of the bat, if you just use Lama tool directly, I don't think you could get like comparable performance, you know, with GPD, 3.5 or four today. But like with fine tuning, if you make that investment in curating your data set in running that fine t…”
Assertion Supported
Shah: Hippocratic AI outperformed GPT-4 on 105 of 114 healthcare exams
“And then we took it, and then we had GPT-IV take it, and we had all the other language models take it, and we beat them all. And we beat them on a 105 of a 114 for GPT-IV, for example.”
Assertion Supported
Nvidia released an AI model that outperformed GPT-4
“Just a couple of weeks ago, Nvidia announced that In addition to having the best chips, they just released a model that was actually better than GPT-IV.”
Insight
AI progression has shifted from post-GPT-4 to a coding agent era
“There's sort of a post GPT-IV era, and then there's a more recent, like, post wide adoption of coding agents era, and then probably soon there's going to be, you know, additional eras, and things are going quite a bit faster, and development is going, you know…”
Assertion Supported
GPT-4 price per token dropped roughly 90% in one year
“The, I think the price per token of GPT-IV dropped something like 90%. Over the last year”
Assertion Supported
Traynor: GPT-4 crossed the hallucination threshold required for customer support bots
“So Finn's built on GPT-IV, by the way, we've, we tried, we wanted to build on a three, 3.5, but it didn't, it's still, I remember back when we used to talk about hallucinations, like four was the sort of the perceptual change for us in terms of trust and relia…”