why aren't all 86 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Opinion
Lenz: Model providers should not dictate enterprise AI policies
“Right now, if you're using a model, you're taking in their own policy. Even if I want to use GPT-OSS, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave, etc. And I don't thi…”
Insight
Bach: Claude is implemented as an invariant pattern similarly to consciousness
“Claude exists only as a pattern. It's something that is a pattern in the activation of the transistors. And even transistors don't actually exist. They are A pattern in the atoms that we are able to see as an invariance because we tune the atoms in a particula…”
Prediction Not checkable as stated
Rieseberg: AI takeoff will create an accelerating, self-reinforcing loop
“Big bang moment where things will accelerate so quickly that it becomes a self-reinforcing loop. And at that point it's sort of like off to the races and there will be no more like slowly catching up. You know, just have Claude being so good at everything.”
Insight
Rieseberg: AI agents need parity with all user tools
“I think that entity needs to have access to all the same tools you have access to. Otherwise it's going to be hamstrung, like all these complex ways.”
Opinion
O'Laughlin: Claude agent teams degrade performance unlike Kimi 2.5 swarms
“My experience is the 2.5 swarm actually improves the model's performance meaningfully. The agent team makes it meaningfully worse because there's clearly not RL done.”
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Opinion
Dwivedi: Claude is superior at agentic tool calling and error unstacking
“For some of the agentic part of the stack, we are shifting towards Anthropic because they're agentic and the tool calling, especially the unstacking part, you know, when you go down the wrong path and you build context that forces you to keep going down the wr…”
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Assertion Not checkable as stated
Nguyen: Stanford HELM benchmark under-reported Claude performance due to improper prompting
“This has happened with, like, Stanford, I remember, like, when Stanford had lists also, like, they were, like, running benchmarks. Yeah, Helm. And somehow, like, Claude was, like, always, like, not performing well, and that's because, like, the way they prompt…”
Insight
Schluntz: JSON Escaping Overhead Degrades LLM Performance Across the Board
“Like if you're trying to output a code in JSON, there's a lot of extra escaping that needs to be done. And that actually hurts model performance across the board. Where versus like if you're in just a single XML tag, there's none of that sort of escaping that …”
Prediction Not checkable as stated
Howard: Reka's model is probably superior to GPT and Claude for certain tasks
“There's a whole model that's been trained in a different way. So there's probably a whole lot of tasks it's probably better at than you know, GPT and Gemini and Claude.”
Insight
Rieseberg: Aggressively anthropomorphizing Claude improves agent UX and architecture design
“And in terms of architecture and UX and everything else that we've been working on Anthropic, it often is quite useful for you to like anthropomorphize cloud aggressively and just be like, this is a person. What would you do if you give, if you had a person, r…”
Prediction Not checkable as stated
Rieseberg: Claude is close to effectively controlling real user computers
“I don't think we're far away from claw being very effective at like using your computer and not just a theoretical computer.”
Opinion
O'Laughlin: Claude for Excel is much worse than Claude Code with Python
“Cloud for Excel is much worse than cloud code using Python to use the Excel skills to then deposit into.”
Insight
Pliny: One jailbroken orchestrator can weaponize segmented sub-agents for cyberattacks
“It's very, very difficult when you have the ability to spin up sub-agents where information is segmented. If you guys know the story of sort of like the builders of the, there's a lot of examples of this in history, but you may, maybe you're building like a py…”
Opinion
Davis: Claude Deep Research Outperforms OpenAI, Perplexity, and Gemini
“And time and time again, over the last couple of weeks, I found that Claude has by far outperformed the others. And I guess the definition of good for me right now is not just length, but also the number of sources and diversity of response.”
Assertion Supported
Ameisen: LLMs use internal circuits to backwards-plan rhyming poetry lines
“And two, this plan doesn't just control, like, what you're gonna rhyme with. It's also doing what's called like backwards planning, where it's like, well, because I need to finish with green, I'm not going to say illuminating the peaceful night, because then I…”
Insight
Brown: Anthropic safety issues stem from conflicting model objectives
“A lot of the kind of headline anthropic like safety results, especially related to reward hacking and kind of deviation and alignment faking, Are all things to me that seem like a rock and a hard play situation where the model has two objectives it's given tha…”
Assertion Not checkable as stated
Anthropic estimates Claude wrote 80% to 90% of the Claude Code codebase
“Probably near 80, I'd say.”
Assertion Not checkable as stated
Anthropic uses Claude to rewrite Claude Code from scratch every 4 weeks
“We've rewritten it from scratch, yeah, probably every three weeks, four weeks or something, and it just like all the, it's like a ship of Theseus, right? Like every piece keeps getting swapped out, and just because quad is so good at writing its own code.”
Assertion Supported
Claude scored nearly twice as high as the next best model
“So we see that Claude right here is almost got twice the score of the nearest best model.”
Opinion
Claude is far better than OpenAI at slang and Gen Z tone
“Whenever we need to do stuff that's a little more conversational or like a little more like a little better at slang, like Claude is way better at slang. Like whenever you ask like Claude to generate something that's like, that sounds like human or like sounds…”
Assertion Not checkable as stated
Hershey: Anthropic Study Found Claude Treats Named Characters Better
“Anthropic actually did like a blinded study of like named characters versus unnamed characters in different settings, and Claude like actually does clearly prefer and is nicer to named characters, which is an interesting thing.”
Opinion
Snipd CEO: Claude is the best model at phrasing and personality
“Like, in my opinion, Claude is the best one when it comes to the way it formulates things.”
Insight
AI agents have an optimal effective context length where intelligence peaks
“I think like one thing you see a lot when you talk to people building agents is there's like some effective context length that actually like has the model be the smartest. And that seems to vary slightly model by model, but for this model, for whatever purpos…”
Disclosure
Upgrading Claude models in Pokémon primarily involves deleting prompt scaffolding
“Literally every model that has come out with Pokemon, like, the main change that I have made to this agent is deleting prompt stuff.”
Assertion Not checkable as stated
Prompting Claude to nickname its Pokémon caused it to exhibit protective behaviors
“And one thing we found when we started doing that is it got more protective of the Pokemon it nicknamed. Like, it's pretty obvious, like, when it catches a Pokemon, Now that it has a nickname, it will, like, go heal it right away if it's hurt, and that did not…”
Assertion Supported
Swyx: Claude wrapper Bolt.new reached $20M ARR
“The other one would be Bolt. There's a straight quad wrapper. And again, another now they've announced twenty million ARR, which is another step up from our eight million that we put on the title.”
Opinion
Neubig: Claude Is The Best Agent Model, Open Models Lag Behind
“I still am under the impression that Claude is the best. The other closed models are, you know, not quite as good, and then the open models are a little bit behind that.”
Opinion
Neubig: GPT Loops On Errors While Claude Tries New Approaches
“So, like, GPT doesn't have very good air recovery ability. And so, because of this, it will go into loops and do the same thing over and over and over again, whereas Claude does not do this.”
Disclosure
Mohan: Windsurf uses Claude for planning, proprietary models for retrieval and diffs
“No, so actually the way it works is the high-level planning that is going on in the model is actually getting done with products like the Cloud. But the extremely fast retrieval, as well as the ability to, like, take the high-level plan and actually apply it t…”
Opinion
Lambert: Frontier Labs Lack Visibility into Cross-Model RLHF Sensitivity
“I think big labs are so over-indexed, are indexed on their own base models, so they don't know, like, what's swapping between CloudBase or GPT-IV-Base, how that would change any notion of preference or what you do with RLHF.”
Insight
Lambert: Claude's constitution dictates output priorities, not model beliefs
“If you look at Claude's constitution, like, that doesn't mean the model believes these things. It's just trying Trained and to prioritize these things.”
Insight
Cheah: AI Engineers Do Not Need ML Math to Build Products
“Frankly, for an AI engineer, you don't need it. You, your main thing that you needed to do was to, frankly, just play around with ChatGPT, or all the alternatives, be aware of the alternatives, because be very mercenary, swap out to Cloudia if it's better for …”
Prediction Not checkable as stated
Swyx: Frontier AI labs will not provide bespoke enterprise integration support
“The labs do not have 200 people dedicated to like, you know, being on call with you with Goldman Sachs going like, okay guys, what do you need? We got it. You need the Microsoft Teams zero integration. Got it. You don't use GitHub. You use this like weird org …”
Assertion Not checkable as stated
Awais: Claude tolerates tool errors and self-corrects, unlike open models
“Claude is actually really, really lenient for tool calls. So even if, you know, your coding agent harness messes up, it can figure out that, oh, I'm being sent this error and can fix itself. Not the case with you know open models”
Assertion Not checkable as stated
Hong: Claude plus AXLE is a go-to setup in Lean community
“And we have seen also, we have heard from a lot of the people that Claude plus Axel is kind of their go-to setup for now.”
Assertion Not checkable as stated
Yan: Claude Opus 4.7 writes detailed PRD-style function comments
“One of the things this new model likes to do is it writes lots of comments, not like, you know, it'll like comment every line, but it'll write like paragraph, like PRDs, like, you know, on top of every function. But I will say to its credit, these aren't slop,…”
Assertion Not checkable as stated
Ludwig: Frontier AI Models Now Excel at Low-Level Code and GPU Shaders
“Six months ago, I would have said the same thing, but it's becoming super useful for every domain.
I'm sure.
Right.
Like there was
I think six months ago, or maybe, maybe a year ago, if you tried to use, let's say the latest Claude model for writing shaders G…”
Insight
Rieseberg: Prompt Opus by stating goals, not specifying exact execution steps
“Honestly though, like I see that you're using Opus 4.6, right? Like my recommendation for people is increasingly don't worry about it anymore. Just like tell it what you want it to do. And it's probably going to figure out a way to do it.”
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Assertion Not checkable as stated
Rumbelow: Standalone Claude Opus Hallucinated Materials Science Data Findings
“So, so Claude, lovely Claude. I'm sorry Claude, but you did a terrible job. It hallucinated some stuff. It made some like big sweeping over, over generalizations. It like over indexed the few outliers. Like, it's fine. It's not Claude's fault. Like Claude is j…”
Insight
Krieger: AI Agents Must Support Both MCP and Visual Computer Use
“And that thing's never gonna have an MCP around it. Like, it's just like, who knows if the company created is even around much less like ready to sort of expose their kind of underlying constructs as API. So I think you will need to be able to do both.”
Assertion Partly supported
Chroma research finds Claude models lead in long-context utilization
“And you know, one thing Chroma released this context rod paper recently about context utilization and the cloud models are actually the best at using kind of like longer context.”
Disclosure
Mohan: Windsurf Cascade splits planning to Claude and codebase application internally
“The high level planning that is going on in the model is actually getting done with products like the cloud, but the extremely fast retrieval, as well as the ability to like take the high level plan and actually apply it to the code base is proprietary systems…”
Opinion
McCloy: Claude users represent an exceptionally valuable audience for companies
“Claude, which is important for, you know, not necessarily huge in terms of raw number of users, but the people who do use Claude tend to be like a very valuable audience, especially for some types of company.”