why aren't all 19 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Crivello: OpenAI's GPT-4o Is Overhyped and Poor for Agents
“I think four O is overhyped. Frankly, we don't use four O. I don't think it's good for agentic behavior.”
Insight
Lopopolo: Reasoning models eliminate need for rigid state-machine scaffolding
“And this I think is like the fundamental difference between reasoning models and the four ones and four O's of the past where these models could not think. So you kind of had to put them in boxes with a predefined set of state transitions. Whereas here we have…”
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Assertion Not checkable as stated
Sands: Switching to o3-mini saved Stripe $3M yearly on risk task
“Or we use, like GPT-IV-O, and it was, like, kind of a little bit expensive to justify the humans that it was replacing for a particular risk-related task, but then next thing we know, like, O-three mini is out, and it's, like, you know, three million dollars a…”
Opinion
OpenAI models remain the industry best for chained function calling
“And so we find like the open AI models function calling wise for our use case. So for the best ones and like, yeah, basically verifying the others. We had a beta group and testing different models. So basically from all providers and yeah, find basically the b…”
Opinion
Polu: GPT-4 Turbo performs better than GPT-4o on function calling
“I personally don't have proof, but I know many people, and I'm probably part of them, to think that GPT-IV Turbo is still better than GPT-IV on function calling.”
Assertion Supported
Molmo 72B Beats Proprietary Models on Academic Benchmarks
“Academic benchmarks wise, Their big one is the best state-of-the-art everything, better than proprietary, but ELO-wise, it sits behind four-oh.”
Assertion Supported
GPT-4.1 reduces extraneous edit rate to 2%, down from GPT-4o's 9%
“And we found that from four O, which got nine percent, which is pretty crazy, nine percent of the time making an extraneous edit is a lot. 4.1 is at two percent, so it's a pretty big improvement.”
Assertion Not checkable as stated
OpenAI's o3 outperforms GPT-4o on Convex evals by a small margin
“You know, oh, three does do better than four. Oh, I mean, we use brain trust for tracking all this quantitatively, but I can't remember off the top of my head, but it's not like a slam dunk.”
Assertion Supported
Nikunj Handa: OpenAI distilled o-series models into GPT-4o search
“They use, like, synthetic data techniques. They've done, like, O-series model distillation to, like, make these four or fine tunes really good.”
Disclosure
Raycast built its Ray 1 models by fine-tuning GPT-4o and mini
“And so we looked into all the various models we had and then we picked, at the moment, it's gbd-for-o and gbd-for-o-mini, Which we basically did a fine tune to really optimize for our use case, and then basically shipping that in the app as Ray one and Ray one…”
Disclosure
Karina Nguyen: OpenAI Retrained GPT-4o to Handle Canvas Edge Cases
“The only way to like fix some of the edge cases is actually through post training. So we actually, what we did was actually retrain the entire full O plus our canvas stuff.”
Opinion
Nguyen: Original ChatGPT Canvas Beta Model Was More Creative Than GPT-4o Canvas
“I would say like the original better model that we released this canvas was actually much more creative than even right now when I use like for, oh, this canvas”
Assertion Supported
Cosine's fine-tuned GPT-4o scored higher on SWE-bench than OpenAI's original o1
“And then we also on this podcast, we interviewed Cosign that actually fine tune four O to, on three bench, sorry to achieve data on three bench. And that score was actually higher than O one when it came out.”
Assertion Supported
Swix: Multi-sampling GPT-4o mini before GPT-4o judging yields net savings
“If I call a GP for a mini 10 times and I do a number of drafts or summaries, and then I have four, oh, judge the summaries that actually is net savings and like a good enough savings then running four, oh, on everything, which given the hundreds and thousands …”
Assertion Supported
Structured response format is limited to GPT-4o and GPT-4o mini
“Actually, the new response format is only available on two models. It's Foro Mini and the new Foro. So the old Foro doesn't have the new response format. However, for function calling, we were able to enable it for all models that support function calling, and…”
Assertion Not checkable as stated
GPT-4.1 Mini significantly outperforms 4o Mini, nearing original GPT-4o performance
“4.1 mini is actually quite significantly better than four o mini and not that far away from the old four o.”
Assertion Partly supported
Alessio Fanelli: GPT-4o Search jumps to 90% accuracy on simple QA
“On simple QA, GPT four O is 30% accuracy. Four O search is 90%.”
What-if
Shreya Shankar: GPT-4o Mini reduces DocETL optimization cost by 90%
“The reason it was a hundred dollars, if I ran the optimizer with GPT-Foro mini as the LLMs, it would be 10 dollars. But we use GPT four. Oh, just because I think we did this at a time where many hadn't come out yet.”