why aren't all 9 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Held up
Marcus: A death will be tied to an LLM within a year
“The prediction that I made Is basically that there will be a death tied to a large language model in the next year.”
Assertion Supported
Marcus: OpenAI's Project Orion failed and became GPT-4.5
“So OpenAI tried to build GPT-V and they had a thing called Project Orion and it actually failed. And eventually got released as GPT four and a half. So what they thought was going to be GPT five just didn't meet expectations.”
Assertion Supported
Marcus: Models like o1 are not systematically better than GPT-4
“Models like O-I are not systematically better than GPT-IV. There's, they're better in certain use cases. Especially ones where you can create data in advance.”
Assertion Supported
Marcus: Microsoft launched Bing Chat despite negative India test feedback
“And then the other really disturbing thing is that apparently they tested in India and got, you know, customer service requests saying it's not ready for prime time. And it was, you know, still put out.”
Assertion Supported
Marcus: OpenAI's o3 hallucinates more than preceding models
“I'll give you just one more example is O three apparently hallucinates more than the models that came before it.”
Assertion Supported
Marcus: Microsoft study suggests chatbot usage impairs human critical thinking
“Well, Microsoft did a study, in fact, suggesting that critical thinking was getting worse as a function of them.”
Assertion Supported
Marcus: Google's LaMDA Has No Sensors Perceiving the Physical World
“Lambda actually has fewer sensors than my watch. My watch has a lot of sensors and lamb doesn't really have anything sensing the real world, except for its linguistic input.”
Assertion Supported
Marcus: Vals AI benchmark shows LLM accuracy under 10% on financial charts
“Where they looked at things like, can you pull out a chart based on a series of financial statements, SEC statements from a bunch of companies and these systems all claimed to do it, but accuracy was under 10%. And overall on this new benchmark, accuracy was a…”
Assertion Supported
Marcus: An amateur Go player beat KataGo 14 out of 15 games
“There was another mind-blowing study this week that showed that one of the best Go programs, KataGo, could be fooled by some silly little strategy that would be obvious to a human player. But, you know, somebody was able to follow this strategy and, like, an a…”