why aren't all 7 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Ratner: $200 of ChatGPT calls can clone closed models into open ones
“If you take a couple hundred bucks of API calls to, say, ChatGPT, and you graph that onto a model that is substantially smaller, say a seven billion parameter model like LLAMA, or now increasingly fully open for commercial use ones, like Red Pajama is one that…”
Prediction Held up
Ratner: Private data models will exceed closed models in specialized tasks
“Closed source models, like a GPT-IV, five, six, seven, whatever comes, are going to be very hard to match in terms of generalist capability for, say, consumer use cases that are reflected in the web data they're trained on and the flywheels that get powered by…”
Assertion Partly supported
Ratner: Distillation creates specialized models 100,000x smaller with higher accuracy
“It shows that you can take these massive kind of generalist foundation models, And not only can you tune them using that kind of programmatic labeling to be more accurate on a given task that you need to be really accurate on, but you can then distill down ver…”
Prediction Not checkable as stated
Ratner: Base models will open-source while production uses specialized models
“There are going to be these base models. They're probably going to be increasingly open source. And then the reality of what actually ships in production is going to be a whole kind of family tree of smaller specialized models that are tuned and adapted via da…”
Assertion Supported
Ratner: DataComp benchmark beat OpenAI models purely through data curation
“Just by cleaning, curating, sampling, filtering the data, we get a new state of the art score at compute parity, beating open AI models and others.”
Prediction Not checkable as stated
Ratner: Foundation models will require customization on enterprise-specific data
“Most of these foundation models, like a GPT-IV, five, six, seven, are going to need to be customized on your specific data and knowledge and workloads.”
Assertion Supported
Ratner: BloombergGPT achieves only low 60s accuracy on specialized financial tasks
“If you actually open the first page of the Bloomberg GPT paper, the FinServe tasks that it does better on by using private financial data are still getting low sixties in terms of accuracy.”