why aren't all 7 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Assertion Supported
Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans
“Essentially, people also think that synthetic data is, like, GPT-IV is more accurate than humans at labeling preferences, so if you look at these diagrams, like, humans are about 60 to 70% agreement, or, like, that's what the models get to, and if humans are a…”
Assertion Supported
GPT-4-level intelligence is now over 100 times cheaper than at launch
“The, like, one fact on that is that you can get intelligence at the level of GPT-IV for over a hundred times cheaper than GPT-IV was at launch right now.”
Assertion Supported
Claude 3.5 wrote fire-and-forget code while GPT-4 used defensive programming
“Claude, for instance, the Sonnet 3.5 was very much fire and forget. It would write code in a kind of Pythonic way, just like, let it fail. Don't be careful about it. Whereas GPT four would use defensive programming, use self assertions.”
Assertion Supported
Swyx: Distilling GPT-4 to Mini cut costs 15x with 2% hit
“Yeah, I sat in the distillation session just now, and they showed how they distilled from four to four mini, and it was like only like a two percent hit in the performance, and 15 X cheaper.”
Assertion Supported
Patel: SemiAnalysis reported GPT-4 mixture of experts architecture in January
“Just being clear, I talked about mixture of experts in January, it's just people didn't really notice it.”
Assertion Supported
Hu: SWE-bench scores jumped from 3% on GPT-4 RAG to 43%
“If we just use GPT-IV plus RAG, what do we get? It's, like, a measly three percent. And then up to the most recent submissions where they get up to 43%.”