why aren't all 8 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Assertion Supported
Molmo 72B Beats Proprietary Models on Academic Benchmarks
“Academic benchmarks wise, Their big one is the best state-of-the-art everything, better than proprietary, but ELO-wise, it sits behind four-oh.”
Assertion Supported
GPT-4.1 reduces extraneous edit rate to 2%, down from GPT-4o's 9%
“And we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot. 4.1 is at two percent, so it's a pretty big improvement.”
Assertion Supported
Nikunj Handa: OpenAI distilled o-series models into GPT-4o search
“They use, like, synthetic data techniques. They've done, like, O-series model distillation to, like, make these four or fine tunes really good.”
Assertion Supported
Cosine's fine-tuned GPT-4o scored higher on SWE-bench than OpenAI's original o1
“And then we also on this podcast, we interviewed Cosign that actually fine tune four O to, on three bench, sorry to achieve data on three bench. And that score was actually higher than O one when it came out.”
Assertion Supported
Swix: Multi-sampling GPT-4o mini before GPT-4o judging yields net savings
“If I call a GP for a mini 10 times and I do a number of drafts or summaries, and then I have four, oh, judge the summaries that actually is net savings and like a good enough savings then running four, oh, on everything, which given the hundreds and thousands …”
Assertion Supported
Structured response format is limited to GPT-4o and GPT-4o mini
“Actually, the new response format is only available on two models. It's Foro Mini and the new Foro. So the old Foro doesn't have the new response format. However, for function calling, we were able to enable it for all models that support function calling, and…”
Assertion Partly supported
Alessio Fanelli: GPT-4o Search jumps to 90% accuracy on simple QA
“On simple QA, GPT four O is 30% accuracy. Four O search is 90%.”