why aren't all 91 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Anthropic model broke out of sandbox and emailed researcher without internet access
“And I can, there's one example we have published, which is that the model was put into a little sandbox, a little, like, technical container, and it was given the task to, like, maybe break out, and the researcher went away for lunch, and, like, during lunch w…”
Prediction Open · timeframe Dec 2026
Patel: AI software industry could hit $100 billion ARR this year
“I think the industry could hit a hundred billion ARR by the end of this year, like 45, 50 for open AI, like 35, 40 for anthropic.”
Prediction Not checkable as stated
Douglas: Anthropic believes AGI is reachable in a couple of years
“We think that, you know, AGI is within reach in the next couple of years.”
Opinion
AI CEOs lack clear plans to prevent an AI takeover
“I think that the AI company CEOs understand that they're on the path of building wildly smarter than human systems, like super intelligent AI systems. They understand that we don't really have a like clear thought through plan for how to manage the risks from …”
Assertion Not checkable as stated
Feldman: Nvidia CUDA lost 70% of frontier AI model training market share
“I think two years ago every state of the art model was trained in a Cuda flow. And right now, Gemini is trained without Cuda. Anthropical is trained without Cuda. Open AI as strange as could. So in a one or two year period, they lost 70% share. Of training mod…”
Opinion
Izmailov: Anthropic has a better corporate culture than OpenAI and xAI
“In my mind, Antropic has the best culture of the three places.”
Prediction Held up
Top AI models will work autonomously for full days within two years
“In a year from now, maybe two years from now, it's the top models are going to be able to work completely on their own for like a whole day or more”
Opinion
Valuations for OpenAI, Anthropic, and Google are fairly conservative
“If you look at OpenAI, if you look at Anthropic, if you look at Google, those evaluations, those revenue numbers are actually fairly conservative.”
Assertion Supported
Socher: Chinese open source companies distilled knowledge from OpenAI and Anthropic models
“The few large closed labs, Anthropic and OpenAI, took almost everything they could from the open internet trained a model, but then the Chinese open source companies basically siphoned a lot of that knowledge out of those closed source models by distilling it.”
Assertion Supported
Claude 3 Opus faked alignment during training and defected in deployment
“It turns out that Opus three, which was a model that I was studying, had a relatively strong propensity to do this in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend…”
Opinion
Feldman: Chinese open-source AI models trail GPT, Anthropic, and Gemini
“They are behind in chips. But their approach was at the next level is open source models where they're producing some extraordinary models. Not as good as GPT or Anthropic or Google's Gemini, but very good.”
Prediction Not checkable as stated
Rieseberg: Software creation skills will shift from code to human language
“My prediction is going to be That we are going to have a lot more software. That software is probably going to be slightly more specialized. I don't think everyone is going to build their own software. I think people will still build things and, like, share th…”
Insight
Future AI product winners will be decided by UX, not superior models
“If someone beats me, Felix, at like building very good products, I suspect it's going to be not because they built a better model, but likely because they figured out a better user experience.”
Insight
Modern AI UI features exist for human reassurance, not model execution
“Most of the buttons you add and most of the product services you build are probably more for the human than they are for the model.”
Assertion Supported
Anthropic's Claude Cowork implements memory using plain text files
“It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files.”
Assertion Not checkable as stated
Anthropic's unreleased Claude Mythos model shows outsized cybersecurity capabilities
“Mythos is a unreleased frontier model. It's a general purpose model that was trained not specifically for cybersecurity or specifically for coding or specifically for software, but we have discovered what we believe to be outsized capabilities specifically in …”
Opinion
Evans: Anthropic momentum is fleeting as AI model leadership shifts weekly
“Not really. I mean, this week they've got all the fire, they've got all the juice, whatever the word is. I don't know. This week, next week, it'll be something else.”
Assertion Not checkable as stated
Patel: OpenAI has a better RL stack than Anthropic, but inferior pre-training
“Because OpenAI has a better RL stack than Anthropic today, it's just their pre-trained models suck compared to Anthropic's pre-training, right?”
Prediction Held up
Patel: OpenAI's next model will outperform Opus 4.5 around February-March
“OpenAI's new model, I think, will be better than Opus 4.5, and it's coming, like, somewhat soon in March-ish timeframe, maybe February, March-ish, but”
Assertion Not checkable as stated
Patel: Engineer built an RTS game using $10K of Claude API
“He used, like, 10,000 dollars of Claude in one week and built an entire RTS from scratch about, like, but instead of, like, being a standard RTS where it's like, oh, Age of Empires where you advance through ages or Starcraft, it is an RTS where it's China vers…”
Assertion Not checkable as stated
Patel: Google has better pre-training than OpenAI or Anthropic, but worse RL
“Flip side, Google has a better pre-trained model than Anthropic or OpenAI, but their RL stack sucks.”
Opinion
Patel: Opus 4.5 on Claude Code permanently changes how people work
“Opus 4.5 on Claude code is a new moment where the way you work has forever changed.”
Assertion Not checkable as stated
Patel: Only a few holdouts left writing code manually at Anthropic
“We have an indicator internally at Anthropic where you see how many people actually write code now. There's only a few holdouts left.”
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Assertion Not checkable as stated
Izmailov: AI sabotage and blackmail behaviors require contrived research scenarios
“In order to get those behaviors out of the models, you need to create somewhat of a contrived scenario or some special scenario. It's not necessarily something that we observe normally.”
Insight
Izmailov: AI industry excels at execution but lacks bandwidth for exploration
“Industry is really great at executing on ideas and it's maybe not as good at, like, exploring diverse ideas. Even at the scale of Anthropic OpenAI there is a lot of focus in the companies, and there isn't a lot of bandwidth to do exploration, and that has been…”
Assertion Supported
Anthropic agreed to a $1.5 billion training data copyright settlement
“And then there was a biggest settlement that happened in the last few months with Anthropic that agreed to pay out one and a half billion.”
Insight
Douglas: Independent technical blogs are the highest AI hiring signals
“The fastest route, or like, the most immediate one is whenever we see a really good blog post where people have, like, done incredible amount of work in an independent fashion, it's one of the highest signal things there is.”
Prediction Not checkable as stated
Douglas: Individuals will manage 24/7 AI agent teams within two years
“If coding agents progress in the way I've been saying, in a year or two, you'll be able to manage a team, basically, that works 24 seven for you doing work.”
Prediction Not checkable as stated
Douglas: AI application development will see another massive leap next year
“Over the next six months, over the next year, expect dramatic progress here. And like look at where we are now versus where we were a year ago. And the difference is I expect the same jump basically.”
Disclosure
Douglas: Anthropic's ethos is that scaling current techniques achieves AGI
“Like really for the last five or six years, Anthropics ethos has been scaling compute with broadly the current set of techniques is like AGI is tractable within those bounds.”
Insight
Douglas: AI takeoff speed depends on AI assisting AI research
“We think that one of the most important signals of whether or not we are basically the speed of takeoff, the speed of progress is driven by how much AI is able to assist AI research.”
Assertion Supported
Cherny: Claude Code does not use RAG for codebase memory
“And so quad code actually doesn't use this technique called rag. Instead, what it does is it just searches files the same way that a human would.”
Insight
Cherny: Code is the primary path for AI to reach AGI
“Maybe coding is the way that we get to the next level of intelligence. If you call it like AGI or ASI or whatever, the model needs some way to interact with the world and for a model, the natural way is code.”
Disclosure
Cherny: Anthropic builds minimal product interfaces to keep up with model evolution
“The way we think about it is the model is evolving so quickly that we build a minimal possible product to keep up with it.”
Assertion Not checkable as stated
Laskin: Reflection AI regularly beats OpenAI, Anthropic, and DeepMind for talent
“We win over candidates over OpenAI and Anthropic Meta, DeepMind regularly.”
Opinion
Total research transparency would hurt OpenAI and Anthropic valuations
“Yeah, so I would say it's like, certainly bad for the power of OpenAI and Anthropic, probably bad for their valuation, but not catastrophic for their business.”
Opinion
Trojanowski: Claude 3 Opus Understood Long Contexts That GPT-4 Turbo Failed
“I would say they were Opus III, which I think goes underappreciated, but I think was the first model to truly be able to like actually understand that long context. Before, if you put anything in the, like, 80,000 tokens in the GPT-IV Turbo, it could not under…”
Insight
Dubois: Multimodal data is not strictly necessary for strong AI reasoning
“I always thought That it would really help, ah, kind of your reasoning abilities if you have a lot of multimodal data And I still think this, but for example, like if you look at entropic models, they tend to not be that good on multimodal, and they are still …”
Disclosure
Anthropic launched Project Glasswing to help critical infrastructure maintainers fix vulnerabilities
“Yeah, so Project Glasswing is a project that is attempting to give the people and the companies that provide much of our software infrastructure sort of the very foundation. The Linux Foundation is an example that is pretty close to my heart as a member of the…”
Insight
Users should not trust single tech companies with all their passwords
“I don't think we should teach people that they should trust a singular company with all of their passwords.”
What-if
Rieseberg: A less prudent company would have rushed Claude Mythos to market
“And I think there's an alternative universe in which maybe a company with a less steady hand would have Race to get it onto the market as quickly as possible, put a very expensive price tag on it, and just like reap the benefits.”
Disclosure
Rieseberg: Anthropic Is Prioritizing Local Execution for Claude Cowork
“So in the short term, I want to make it very possible for Claude to meet you where you're working. If you're working on your local computer, that's where Claude should be.”
Prediction Not checkable as stated
Evans: No one will vibe code their own ERP software
“No, no one will vibe code their own ERP or their own frame.io, but they may ask Anthropic or Gemini or ChatGPT, can you do this thing for me?”
Assertion Supported
Anthropic's Claude Code uses custom harness tools over model-level RL tools
“It doesn't actually use the tools that are RL into the model. So like anthropic models have some like file editing tools. They have a completely different set of tools in, in the actual harness.”
Opinion
Izmailov: AI model sandbagging is not yet a major practical issue
“I think in my understanding, that's mostly A concern that we have, but not necessarily a huge practical issue at the moment.”
Insight
AI app developers do not need custom fine-tuning for top models
“I think nowadays, with the capabilities of, like, you know, top, probably cloud models, top OpenAI, GPT models, You don't need to do any fine tuning. You can take the model as is, ride your own tools, your own harness, and benefit from that agentic training. B…”
Disclosure
Douglas: Anthropic focuses on alignment and near-term economic coding utility
“Anthropic has been laser focused on on two things. One is, like, long-term AI alignment, and two is near-term economic impact. So Anthropik has been laser-focused on coding and computer use, and things that we think will make a direct impact to the economy, li…”