why aren't all 235 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 3 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Howard: Rapid LLM race created massive technical debt and optimization opportunities
“There's a whole lot of technical debt everywhere, you know, nobody's really figured this stuff out because everybody's been so busy building what we know works as quickly as possible. So, yeah, I think there's a huge amount of opportunity to, you know, I think…”
Prediction Not checkable as stated
LLaMA 2 will shift developers from closed APIs to self-hosting
“And I do see that's going to shift the balance of it. More and more folks are going to be using let's say derivatives of Lama two. More folks are going to Fine-tune and serve their own model instead of calling an API.”
Opinion
Berman: Aggressive quota limits prove Anthropic faced severe compute shortages
“I still do think they were bandwidth limited or they were compute limited because if you look at their quota and how aggressive and reducing it and using it it's gotten much better, but especially two months ago, I mean, you would burn through your quota in a …”
Assertion Not checkable as stated
Anthropic made coding its day-one priority as the mechanism to AGI
“And there, P zero from day one was coding. The reason the mechanism system there was, if we crack coding, Then we will crack AGI. You know, our mission is AGI. We want to get there safely. If we focus on coding, it's such a generally powerful capability that i…”
Assertion Not checkable as stated
Anthropic achieved technical model takeoff during its October 2023 training run
“What happened is, Anthropic basically achieved takeoff in October of last year. That training run.”
Disclosure
Notion Partnered With Anthropic and OpenAI to Build 30% Pass Rate Evals
“And then what we have, what we call Frontier Headroom evals, where we actively want to be at 30% pass rate. And that's actually been a effort that we took in partnership with Anthropic and OpenAI in the past maybe two or three months, because we actually hit a…”
Insight
Rieseberg: Aggressively anthropomorphizing Claude improves agent UX and architecture design
“And in terms of architecture and UX and everything else that we've been working on Anthropic, it often is quite useful for you to like anthropomorphize cloud aggressively and just be like, this is a person. What would you do if you give, if you had a person, r…”
Disclosure
Rieseberg: Claude Cowork assembled existing prototypes rather than starting from scratch
“And what cowork actually became is like, we sort of picked the right pieces out of the many prototypes that we had. Right. And that's maybe also like, I think an important qualifier whenever people mention this like 10 day number, I do think it's important to …”
Prediction Not checkable as stated
Rieseberg: Claude is close to effectively controlling real user computers
“I don't think we're far away from claw being very effective at like using your computer and not just a theoretical computer.”
Disclosure
Anthropic: Claude Cowork executes inside a dedicated lightweight Linux VM
“So we currently run like a, we currently run like a lightweight VM and we put clock code into the VM and we do that for a number of reasons. Safety and security is a big one, but even if you ignore for a second safety and security and you're just like, okay, Y…”
Assertion Supported
Levie: Anthropic has forward-deployed engineers embedded at Goldman Sachs
“OpenAI probably is hiring FDEs to go into the enterprise and then Anthropic is embedded at Goldman Sachs.”
Assertion Not checkable as stated
AAIF began when Block asked Anthropic about donating the Goose agent
“We got approached by our friends at block to discuss because they were looking into like donating goose, I think at the time. And so there was a question around doing something together. And then we approached open AI and they were very, very welcoming and lik…”
Opinion
Pliny: AI labs lack enough researchers to explore latent space alone
“They don't have enough researchers to explore the entire latent space on their own. And so I think many hands make light work”
Disclosure
Ubl: Vercel publishes evals to influence OpenAI and Anthropic models
“I'm Vercel and I publish at Eval. That I want OpenAI and Anthropic to use to make sure when they ship the next model that they're better at the stuff that I care about.”
Opinion
Anthropic's MCP Has Won as the De Facto AI Interop Standard
“Now it's like basically kind of de facto one as the interop layer for all the labs and all the models.”
Insight
Menlo Underwrites Late-Stage AI on Revenue and Margin, Not Market Share
“At that stage, to be very honest with you, at the stage that we invest in Anthropic now, like the only things that would really move the needle on the decision is here's the revenue, here's the margin, and here's the trajectory, and here's the other markets we…”
Assertion Supported
Glean Operates at a Several Hundred Million Dollar Revenue Scale
“Look at the revenue of Anthropic and OpenAI right now. These are billion dollar revenue scale businesses. Glean is several hundred million dollar revenue scale business.”
Prediction Not checkable as stated
Swyx: Tuning reasoning activations could let Anthropic leapfrog OpenAI
“If Anthropic ever found
The activations for reasoning and could break down the different kinds of reasoning and turn, tune them properly.
I think that's the thing that takes Anthropic to leapfrog OpenAI.”
Insight
Webster: Enterprises do not want AI models to be maximally helpful
“OpenAI Anthropic, everyone else, they're all building models that are like maximally helpful. And in Actually, most cases in a corporate environment, you don't want that to be maximum. You don't want the model to be like helpful in every way possible.”
Assertion Not checkable as stated
Sonnet 4.5 ran autonomously for 30 hours versus 7 for Opus 4
“So this, you know, we had a customer and internally, we also got like a 30 hour plus kind of execution versus I think Opus four was seven hours.”
Insight
Martin: Multi-agent systems excel at parallel read-only tasks, not writing tasks
“I like the take that apply multi-agents to problems that are easily parallelizable, that are read-only, for example, context gathering for deep research, and do, like, the final quote-unquote write, in this case report writing, at the end. I think this is tric…”
Prediction Not checkable as stated
Dax Reed predicts Cursor will move faster than Anthropic's Claude Code team.
“I mean, when I saw the acquisition, I was like, oh, this is actually more intense now because the cursor team is going to move faster than the cloud code team. I think that's my feeling given they have to, there's more pressure on them and they're smaller.”
Assertion Not checkable as stated
Hou: Windsurf is among Anthropic and OpenAI's largest consumers
“We've had immense success getting people onto the platform, and we've been very fortunate to have the issue of being some of Anthropic and OpenAI's largest consumers.”
Prediction Not checkable as stated
Rizwan: AI ecosystem will converge around an official Anthropic MCP registry
“I think the entire ecosystem will just converge around whatever they do. They just have such good distribution and they're, yeah.”
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Assertion Partly supported
Anthropic finds multi-agent architecture outperforms single-agent baseline by 80%
“So the, in the blog post that Anthropik posted, they ran some tests and they noticed that the multi-agent structure outperforms the single agent structure by 80% based off a different variety of variables they measured.”
Disclosure
Zach Lloyd: Anthropic, not OpenAI, is Warp's primary model of choice
“ChatDBT isn't even the model of choice at this point. It's anthropic for us.”
Prediction Not checkable as stated
Zach Lloyd: Anthropic might launch a desktop agent harness like Warp
“I think even hearing the cloud code folks on your podcast, like I would not be surprised if Anthropic like launched a thing that looks a little bit more like warp where it's like a harness for running a whole bunch of different cloud codes, but it's an actual …”
Assertion Supported
Ameisen: LLMs use internal circuits to backwards-plan rhyming poetry lines
“And two, this plan doesn't just control, like, what you're gonna rhyme with. It's also doing what's called like backwards planning, where it's like, well, because I need to finish with green, I'm not going to say illuminating the peaceful night, because then I…”
Assertion Supported
Ameisen: Multi-Hop Reasoning Circuits Are Extremely Similar Across Small and Large Models
“The way the circuit looks in Gemma, like a really small model is extremely similar to the way that it looks like a huge model, which that in itself is, I think like a pretty novel discovery. It's like, oh, you have these models that are like super different. Y…”
Disclosure
Ameisen: Anthropic's circuit tracing tool ignores attention heads and only decomposes MLPs
“These are just MLPs. So the model has both attention heads and multi-layer perceptions MLPs. We don't just do it. Like we completely ignore attention or like we don't try to decompose it at all. So there's some prompts where like all of the interesting stuff i…”
Insight
Ameisen: LLMs Execute Parallel Sub-Processes During Math and Hallucinations
“So I think one example of this is like math where the model is like independently computing the like last digit and then the like order of magnitude and then kind of like combining them at the end or like hallucinations are also that where like, there's one si…”
Assertion Supported
Ameisen: Swapping Internal Features Proves Single-Pass LLM Multi-Step Reasoning
“We claim that this is like the Texas representation. Let's get another one and replace it. And we just change like that feature in the middle of the model and we change it to like California. And if you change it to California, sure enough, it says Sacramento.…”
Assertion Supported
Ameisen: Mechanistic interpretability methods successfully scaled to production models
“And it turns out scaling it. I don't want to say it just worked because it was a lot of work. I don't mean to apply. There was an effort, but it worked. And now we're in the phase where it's like, oh, cool. These methods work on the models that we care about.”
Assertion Not checkable as stated
Ameisen: Interpretability researchers lack good methods for analyzing attention layers
“So like, I think that right now we have some pretty good solutions for like understanding what's in the residual stream, understanding what's, is it in MLPs? We don't have good solutions for like attention.”
Disclosure
Ameisen: Anthropic publishes interpretability research to recruit more researchers
“The reason for publishing this is that we think interpretably is important. We think it's tractable, and we think more people should work on it. And so publishing it helps us like accomplish with these goals all these goals, which we think are just like crucia…”
Assertion Supported
Ameisen: Larger language models share more concept representations across languages
“If you look inside the model, if you look at the middle of the model, which is the middle of this plot here, models share more features. They share more of these representations in the middle of the model, and bigger models share even more. And so the, like, t…”
Assertion Supported
Ameisen: Anthropic Trained a Misaligned Model With Hidden Goals for Detection
“A team at Anthropic trained a model to have like weird hidden goals and then gave it to a bunch of other teams and said, Figure out what's wrong with it”
Opinion
Brown: Anthropic treats extended thinking as tool use, not distinct model class
“And it seemed like Anthropik's kind of attitude has been that extended thinking is an instance of tool use and that it's the kind of thing you want to equip the model with the ability to do. But it's not like, oh, it's a thinking model. It's just a sync for th…”
Insight
Brown: Anthropic safety issues stem from conflicting model objectives
“A lot of the kind of headline anthropic like safety results, especially related to reward hacking and kind of deviation and alignment faking, Are all things to me that seem like a rock and a hard play situation where the model has two objectives it's given tha…”
Assertion Not checkable as stated
Cherny: Some Anthropic engineers rack up thousands daily running Claude Code automations
“And there's some people at Anthropic that have been racking up like thousands of dollars a day with this kind of automation.”
Assertion Not checkable as stated
Anthropic's non-coding product designer ships monorepo pull requests using Claude Code
“And she's landing PRs to our console product. So it's not even just, like, building on quad code. It's building, like, across our product suite in our monorepo.”
Assertion Not checkable as stated
Cherny: Claude Code is the thinnest possible wrapper over the underlying model
“All the secret sauce, it's all in the model and this is the thinnest possible wrapper over the model. We literally could not build anything more minimal. This is the most minimal thing.”
Assertion Not checkable as stated
Anthropic estimates Claude wrote 80% to 90% of the Claude Code codebase
“Probably near 80, I'd say.”
Assertion Not checkable as stated
Anthropic observes Claude Code API costs averaging roughly $6 daily per user
“Currently we're seeing costs around, like, six dollars per day per active user, and so it's, like, it does come out to a bit higher over the course of a month in Cursor but I don't think it's, like, out of band, and that's, like, roughly how we're thinking abo…”
Assertion Partly supported
METR benchmark: AI agent autonomy duration doubles every 3 to 7 months
“They established a Moore's law for time between human input, basically, and it's basically doubling every three to seven months is the idea. And Enthopic is currently doing super well on that benchmark. It's roughly about autonomous for 15 minutes at the 50th …”
Assertion Not checkable as stated
Anthropic uses Claude to rewrite Claude Code from scratch every 4 weeks
“We've rewritten it from scratch, yeah, probably every three weeks, four weeks or something, and it just like all the, it's like a ship of Theseus, right? Like every piece keeps getting swapped out, and just because quad is so good at writing its own code.”
Assertion Supported
Fanelli: Anthropic scrapes 6,000 pages per referral, compared to OpenAI's 250
“Google would be a two to one crawl to referral ratio, so for every two pages, they will read, they will send you one visitor. He said OpenAI is 250 to one, so they'll read 250 of your pages and send you one person. And Anthropic was like 6000 to one. So they'l…”