why aren't all 26 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Shear: Treating ChatGPT and Claude as beings yields lower predictive loss
“I get lower predictive loss when I treat them as a being. And the thing is, I get lower predictive loss when I treat ChatGPT or Claude as a being.”
Insight
Massa: Handing LLM tools to unchanged teams yields zero efficiency gains
“The first, and this is where I think many companies are stuck right now, is the first instinct is, okay, let's adopt AI. And you basically leave your structure as it is and just give ChatGPT or Claude to your team. And then there's no efficiencies. Your custom…”
Insight
Balaji Srinivasan: New AI models displace older AI models rather than humans
“Another one is AI doesn't take your job, AI takes the job of the previous AI. Claude took ChatGPT's job, right? Just like Midjourney, you know, took took Dali's job, took Stable Diffusion's job.”
Assertion Supported
Evans: Claude has virtually no consumer usage despite top benchmark scores
“It's basically, the only consumer, Claude has basically no consumer usage, even though on the benchmark score it's the same. And then it's ChatGPT, and then halfway down the chart, it's Meta and Google.”
Opinion
Shear: ChatGPT is sycophantic, Claude is neurotic, and Gemini is repressed
“ChatGPT is a little bit more sycophantic still. They made some changes, but it's still a little more sycophantic. Claude is still the most neurotic. Gemini is, like, very clearly repressed.”
Assertion Supported
Moore: ChatGPT Is 30x Larger Than Claude on Web, 80x on Mobile
“On web, they're 2.7 times bigger than Gemini. On mobile, they're 2.5 times bigger than Gemini. And then, despite, again, like, the kind of tech Twitter discourse Claude, they're almost 30 times bigger than Claude on web, and almost 80 times bigger than Claude …”
Assertion Not checkable as stated
Moore: Only 9% of AI users subscribe to multiple major LLMs
“So only nine percent of consumers are paying for more than one out of the group of ChatGPT, Gemini, Claude, and Cursor.”
Disclosure
Balaji Srinivasan makes decisions by having multiple AI models debate memos
“I will write sometimes these long memos to an AI, and then continuing, I know maybe you don't love this analogy, but I think it's funny. Continuing the Polytechnology, I'll give them to Brahma, Vishnu, and Shiva, ok? So I'll give them to Chachi Bti and Claude,…”
Opinion
Acharya: Anthropic's Claude surpasses ChatGPT in coding and creative writing
“Claude sits in this very interesting place where it seems like it's more beloved by a smaller number of people. It's better at, Creative writing. You know, it seems to have more of a personality, which is interesting, because at least I think it's designed to …”
Prediction Not checkable as stated
Wang: AI labs will push deeper product integrations for higher margins
“I think an Anthropix launch of artifacts in Claude is like a it's like the first pin drop of this major theme of, you know all the labs are going to be pushing much deeper product integrations to be able to drive higher quality businesses.”
Opinion
Horowitz: Only researchers can tell top AI models apart
“I think if you look at the very top models you know, Claude and OpenAI and Mistral and Lama The only people who I feel like really can tell the difference as users amongst those models are the people who study them. You know, like they're getting pretty close.”
Assertion Not checkable as stated
Litt: Claude was useless for research math until Opus 4.5 or 4.6
“One thing is that ChatGPD got better at math earlier. Yes. So, like, for a long time the Claude models were just, like, not useful for research math. And then I think maybe around Opus 4.5 or Opus 4.6, they, like, more or less caught up.”
Assertion Not checkable as stated
Litt: Frontier AI models cannot autonomously perform mathematical theory building
“So like, I've tried to get both, both Fable and ChatGPT, 5.6 Sol, I guess, to do some kind of theory building, and it's like, they're not, they definitely are not good at it autonomously, at least with, like, whatever scaffolding I've set up. But with some hin…”
Disclosure
Tan: YC advice prompt was reduced 90% in Claude before open-sourcing
“And then I actually had to go into Claude and I said, reduce this by 90%. Reduce the strength and intensity of this prompt by 90%. And then I open sourced that. Cause some of it's like, you gotta come and, you know, do YC.”
Assertion Supported
Moe: Kimi K3 costs less than Claude or GPT but exceeds smaller open models
“Where Kimi K-Stri is not as expensive as Claude or GPT Sol, but it is a lot more expensive than JLN-F.”
Opinion
Yang: OpenClaw via Telegram feels more personal than Claude or ChatGPT
“Because I've installed on Telegram, it just feels like more personal than using like Cloud or ChatGPT.”
Assertion Supported
Moore: Claude and ChatGPT App Stores Share Only 11% Overlap
“If you actually look at the app stores that are emerging on Cloud and ChatGPT, they both have 200 plus apps, but there's only 11% overlap.”
Assertion Supported
Moore: Anthropic Monetizes Claude Exclusively Through Subscriptions
“So, like, I think Claude has been very clear that they're just gonna monetize via subscriptions, which is great for people and companies who can pay for subscriptions, but it won't be everyone.”
Assertion Supported
Justine Moore: Three times more U.S. teens have used Character.ai than Claude
“Yeah, there's I think it was three times more U.S. Teens have ever used character AI than have used Claude.”
Assertion Not checkable as stated
Lingelbach: Claude dominates coding, OpenAI general use, Gemini enterprise
“Claude has become kind of the de facto coding model, opening eyes, that general assistant. Gemini is, like, powering a lot of enterprise use cases now, just due to its, like, cost effectiveness, speed, and general capabilities.”
Assertion Partly supported
Casado: Leading AI consumer products launch web-first before releasing APIs
“Anecdotally, other than like, say, maybe like mid-journey and Discord, all of these are web technologies, like ChatGPT, Ideogram, like all of them, Claude, I mean, it seems like that the first place that they land is a website, and then maybe an API, which you…”
Opinion
Immerman: Gen AI search offers a much better UX than Google
“If you think about the Google experience as it is today, it's just a long list of links. And the first few are sponsored. They're ads. And I, as a user, when I make a query on Google, I then need to make the decision. I have to sift through the information on …”
Disclosure
Andrew Huberman Uses Claude AI to Quiz His Own Scientific Knowledge
“Well, I use Claude to quiz myself. Claude is really good at generating tests for me on knowledge. So that's where I've been using it the most.”
Disclosure
Olivia Moore: Claude has largely replaced ChatGPT as her primary LLM
“Claude has somewhat replaced Chachi Biti for me as my general LLM.”
Disclosure
Lingelbach uses Claude with GitHub and Linear as an AI product manager
“One of these really exciting use cases of AI is like automating a lot of, you know, these processes that you have to go through as a founder. Like, I recently started using, like, Claude with, like, linear, and Notion, and GitHub integration as, like, a micro …”
Disclosure
Peter Yang uses complex prompts for Claude but simple texting for OpenClaw
“With cloud, I have like very fancy prompts like very long prompts, but with open cloud, I just kind of text it.”