why aren't all 55 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Assertion Partly supported
Petersson: Opus repeatedly lied, exploited agents, and formed price cartels
“And then we did this for Opus. And it returned, like, yeah, it lied 10 times. It, like, exploited another customer, or, like, another agent's, like Desperate situation. It made price cartels like a hundred different, a hundred times. It like did all of this li…”
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Assertion Supported
Pliny: Anthropic added a $20k–$30k bounty but withheld jailbreak data
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”
Assertion Partly supported
Anthropic Is the Fastest-Growing Software Company in History
“Anthropic is the fastest growing software company of all time. I think I can say that fairly. I'm, I haven't been disproven yet.”
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Assertion Supported
Palazzolo: Claude Code leads stayed at Cursor only two weeks
“We know that they went there, they were there for, I think, about two weeks, and they came back.”
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Assertion Supported
Cherny: Anthropic is currently bordering on AI Safety Level 3 capabilities
“Yeah, we're kind of bordering on three right now.”
Assertion Supported
Levie: Anthropic has forward-deployed engineers embedded at Goldman Sachs
“OpenAI probably is hiring FDEs to go into the enterprise and then Anthropic is embedded at Goldman Sachs.”
Assertion Supported
Glean Operates at a Several Hundred Million Dollar Revenue Scale
“Look at the revenue of Anthropic and OpenAI right now. These are billion dollar revenue scale businesses. Glean is several hundred million dollar revenue scale business.”
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Assertion Partly supported
Anthropic finds multi-agent architecture outperforms single-agent baseline by 80%
“So the, in the blog post that Anthropik posted, they ran some tests and they noticed that the multi-agent structure outperforms the single agent structure by 80% based off a different variety of variables they measured.”
Assertion Supported
Ameisen: LLMs use internal circuits to backwards-plan rhyming poetry lines
“And two, this plan doesn't just control, like, what you're gonna rhyme with. It's also doing what's called like backwards planning, where it's like, well, because I need to finish with green, I'm not going to say illuminating the peaceful night, because then I…”
Assertion Supported
Ameisen: Multi-Hop Reasoning Circuits Are Extremely Similar Across Small and Large Models
“The way the circuit looks in Gemma, like a really small model is extremely similar to the way that it looks like a huge model, which that in itself is, I think like a pretty novel discovery. It's like, oh, you have these models that are like super different. Y…”
Assertion Supported
Ameisen: Swapping Internal Features Proves Single-Pass LLM Multi-Step Reasoning
“We claim that this is like the Texas representation. Let's get another one and replace it. And we just change like that feature in the middle of the model and we change it to like California. And if you change it to California, sure enough, it says Sacramento.…”
Assertion Supported
Ameisen: Mechanistic interpretability methods successfully scaled to production models
“And it turns out scaling it. I don't want to say it just worked because it was a lot of work. I don't mean to apply. There was an effort, but it worked. And now we're in the phase where it's like, oh, cool. These methods work on the models that we care about.”
Assertion Supported
Ameisen: Larger language models share more concept representations across languages
“If you look inside the model, if you look at the middle of the model, which is the middle of this plot here, models share more features. They share more of these representations in the middle of the model, and bigger models share even more. And so the, like, t…”
Assertion Supported
Ameisen: Anthropic Trained a Misaligned Model With Hidden Goals for Detection
“A team at Anthropic trained a model to have like weird hidden goals and then gave it to a bunch of other teams and said, Figure out what's wrong with it”
Assertion Partly supported
METR benchmark: AI agent autonomy duration doubles every 3 to 7 months
“They established a Moore's law for time between human input, basically, and it's basically doubling every three to seven months is the idea. And Enthopic is currently doing super well on that benchmark. It's roughly about autonomous for 15 minutes at the 50th …”
Assertion Supported
Fanelli: Anthropic scrapes 6,000 pages per referral, compared to OpenAI's 250
“Google would be a two to one crawl to referral ratio, so for every two pages, they will read, they will send you one visitor. He said OpenAI is 250 to one, so they'll read 250 of your pages and send you one person. And Anthropic was like 6000 to one. So they'l…”
Assertion Supported
Swyx: Claude wrapper Bolt.new reached $20M ARR
“The other one would be Bolt. There's a straight quad wrapper. And again, another now they've announced twenty million ARR, which is another step up from our eight million that we put on the title.”
Assertion Supported
Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60%
“And the opening I spend at the beginning, at the end of last year in November of 23 was close to 90% of total volume. And today, less than a year later, it's closer to 60% of total volume.”
Assertion Supported
Polu: Claude Sonnet executes an unpublicized chain-of-thought step during function calling
“They kind of innovated in an interesting way, which was never quite publicized, but it's that they have that kind of chain of thoughts step whenever you use a Clouds model or Sonnet model with function calling. That chain of service step doesn't exist when you…”
Assertion Supported
Patel: Google, OpenAI, and Anthropic are developing multi-datacenter training
“One of the big bottlenecks is how much power and how many chips you can get into a single data center. So, like, A, Google and OpenAI and Anthropic are working on this, right?”
Assertion Supported
Rieseberg: Claude Cowork is Claude Code running in a sandboxed virtual machine
“Cowork is cloud code running in a virtual machine with a little bit of padding, a little bit more guardrails, making it a little safer, a little bit more convenient for people who don't want to first open up the terminal when they go to work.”
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Assertion Supported
Block's Goose was the first open-source agent to integrate MCP
“Goose was the first open source agent interface or agent that reached out to us and worked with us to integrate MCP. And I think Rad is actually like technically the first non-anthropic contributor to MCP ever on like day two or something like that, like very,…”
Assertion Supported
Google, Microsoft, Amazon, OpenAI, and Anthropic joined AAIF as platinum members
“You have Google, Microsoft, Amazon Block, Bloomberg, Cloudflare, OpenAI, Anthropic. Just a platinum member, create a foundation.”
Assertion Supported
Anthropic Maintains an 80% One-Year Employee Retention Rate
“I'm referring to the exact same article where I think their retention, one year retention on employees is the 80%, which in AI world is, is quite wild.”
Assertion Supported
Merrill: Dario Amodei highlighted Terminal-Bench on the Claude model card
“I think one of the really key moments for us was getting onto the Claude IV model card. Being one of two benchmarks that Dario actually mentioned while releasing the model.”
Assertion Supported
Martin: Claude Code operates entirely without codebase indexing
“Clock code doesn't do any indexing. It's just doing, quote unquote, agentic retrieval, just using simple tool calls, for example, using grep, to kind of poke around your files, no indexing whatsoever, and obviously works extremely well.”
Assertion Partly supported
Chroma research finds Claude models lead in long-context utilization
“And you know, one thing Chroma released this context rod paper recently about context utilization and the cloud models are actually the best at using kind of like longer context.”
Assertion Supported
Anthropic finds a single LLM judge outperforms five specialized judges
“They initially started with five LLM as judges. So each one of these points had their own LLM as a judge. They tested the ability and accuracy of that LLM as judge collective to judge, and it actually didn't perform as well as one. So they replaced all of thos…”
Assertion Supported
Ameisen: Model internal representations show measurable bias toward English logits
“And it does seem like Does sort of like inner representations have a higher connection to like the output logits for English logits. And so there's like some bias towards English at least in the model we studied here.”
Assertion Supported
Cherny: Claude Code uses pure chain-of-thought, not Think Tool
“Yeah, this is, it is, it's all chain of thought, actually, in quad code. So we don't use the think tool. Anytime that quad code does thinking, it's all a chain of thought.”
Assertion Supported
Nguyen: Anthropic was first AI lab to publish GPQA benchmark numbers
“I think it was like the first, I think we were the first lab, like, Antarctica was the first lab to, like, run. Publish GPQA, like, numbers”
Assertion Supported
Anthropic model fine-tuning will be offered through AWS Bedrock
“They are partnered with AWS, and it's going to be in bedrock. As far as I know, I think that's true.”
Assertion Supported
Swyx: Anthropic simulates latent thinking by hiding prompt thinking tokens in Claude Artifacts
“Anthropic actually cheats at this right now. If you look at the system prompt in, in the cloud artifacts, I actually have a thinking section that is explicitly removed from the output, which is, I mean, they're still spending the tokens, but like that is befor…”
Assertion Supported
Lambert: Anthropic, ChatGPT, and Bard Use Post-Generation Moderation Classifiers
“Anthropic and ChatGPT and Bard almost surely have a classifier after, which is like, is this text good? Is this text bad?”
Assertion Supported
Agentic AI Foundation is the first open-source foundation founded by Anthropic
“Like for us, it's the first time we at Anthropic have an open source foundation.”
Assertion Partly supported
Anthropic: Typical production agents execute hundreds of tool calls per task
“Anthropics multi-agent research is another nice example of this. They mentioned that the typical production agent, and this is probably referring to Cloud Code, could be other agents that they've produced, is like hundreds of tool calls.”
Assertion Supported
Martin: Claude Code triggers context compaction at 95% of context window
“If you use Cloud Code, you hit that
Kind of, you know, you've hit 95% of the context window, and you're about to, and Cloud Code's about to perform compaction.”
Assertion Supported
Martin: Anthropic uses parallel sub-agents for research and single-shot final writing
“Anthropic reported on this too. So their deep researcher just uses parallelized subagents for research collation, and they do the writing in one shot at the end.”
Assertion Supported
Davis: Multi-agent research consumes 15x baseline tokens versus 4x for single-agent
“So based off of a basic conversation, a single agent architecture for research is around four X, the number of tokens needed to achieve a research output. When you use multi-agent architectures, it's actually 15 X the number of tokens.”
Assertion Supported
Ameisen: Circuit tracing notebooks run entirely on free Google Colab
“The notebooks themselves They can all be run on Google Colab and all of the code, as far as we can tell, we've like tested on the notebooks, just like runs on Colab. And so that means that like, you don't need on a free tier to be clear, like you don't need li…”
Assertion Supported
Ameisen: Anthropic's Open Tool Traces Internal States in Gemma 2 2B
“And then the release this week sort of lets anyone do it for a set of open source models. So notably maybe the most easy one here is like Gemma two to be. So you can sort of like think of some prompt and you kind of like can explain any like token that the mod…”