why aren't all 69 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Awais: Claude Code hides 50+ tool-call failures per session on DeepSeek
“In CloudCode you know, they hide a lot of the errors behind control O, right? So you don't even know that, you know, you have like 50 plus tool call failures plus per session. You're just sitting there and you're like, oh, why is DeepSeq so slow?”
Opinion
Shah: Claude Code avoids indexing because vendors profit from high token usage
“They could index the code and cursor does, but plot code does not. And I believe this is also because they are not incentivized to utilize less tokens by the end user, because they're just offsetting the cost of indexing, et cetera, to tokens and time that the…”
Opinion
O'Laughlin: Excel, Bloomberg, and Traditional IDEs Are Obsolete Due to AI Agents
“I believe every one of these IDEs are done. It's dead over and gone and dead. I just think it's, why? Why? It doesn't, like, just imagine the concept of you. Like, I remember when I learned Bloomberg. I had to, like, watch videos to learn about all the random …”
Opinion
Yegge: Claude Code is too hard for most engineers
“And the answer is it's too hard. It's too hard. You have to be able to read, man, most engineers, honestly, like to them, five paragraphs is an essay. Okay. And with cloud code, you've got to read waterfalls of not just information, but also code and diffs, ri…”
Prediction Not checkable as stated
Slack: AI tools like Cursor and Claude Code will peak and decline
“A lot of these other tools that are great, like cloud code and codex and cursor and so on that they've forgotten what made them great and what made them grow so fast, which is building the very best product. And they built it in a way that's too overfit on the…”
Opinion
Zach Lloyd: Warp's coding agent outperforms both Cursor and Windsurf
“Warp can code, and Warp code is at a level that's, like, comparable to cloud code. I think it's actually better than, like, Cursor's agent. It's probably better than Windsurf since they aren't, they don't have Sonnet for.”
Insight
Parakhin: Effective PR review requires largest pro-level models, not fast tools
“At PR review time, you want to run the largest models. That means codex or cloud code is not going to cut it. You need to have pro-level models if you really want to stem the tide of bugs from going into production”
Prediction Not checkable as stated
O'Laughlin: Tech industry may face a CPU shortage from AI coding and RL
“You feel like we might actually be seeing a CPU shortage partially because of this refresh cycle, but partially also because like I legitimately believe the cloud code Cloud code is increasing software creation and then on top of that, there is real demand fro…”
Insight
O'Laughlin: Reviewing AI agents requires past hands-on manual domain experience
“If you didn't pay any like human cognition to get there, I don't think you're going to be a great reviewer. One of the reasons why, you know, what makes that, that human feet, that loop well is because once upon a time you did that and you could make the three…”
Prediction Open · timeframe Dec 2026
O'Laughlin: AI Agents Will Write 25% to 50% of Public Code by Year-End
“I sandbagged the ever-living share of that. I just believe 25 is, Very, like, I, like, it's like a, the rate it's on is like, whatever, 50 or something like that, but I think, I feel, I wanted to give a 95 confidence interval. I think 25 is within the 95 confi…”
Prediction Not checkable as stated
Merrill: Frontier AI labs will center operations around vertical products
“And with the Claude codes and the codec CLIs and the deep researchers, researchers, you starting to see some evidence that the products are going to be a much more central part of how these frontier labs operate.”
Prediction Not checkable as stated
Dax Reed predicts OpenCode will dominate when a competitor beats Claude Sonnet.
“What would change things is if there's a day where either another LM lab or like, you know, an open source model drops that is competitive with Sonnet, maybe even better than Sonnet on that day, open code is going to be the only way to do this kind of thing. C…”
Disclosure
OpenCode exactly replicated Claude Code by dumping its system prompts and schemas.
“We look at cloud code. We dump all their system prompts. We dump all their tool descriptions. We dump all their tool schemas. We re-implement the tools. And when you're using an anthropic model, we basically have the exact same implementation.”
Disclosure
Anthropic avoided building an IDE to target future AI model scaling
“If we want some product that has like very broad a product market fit today, we would build, you know, a cursor or a Windsurf or something like this. Like these are awesome products that so many people use every day. I use them. That's not the product that we …”
Insight
Cherny: AI prototyping informs engineering decisions faster than writing design docs
“Before I would write a big design doc, and I would think about a problem for a long time before I would build it sometimes for some set of problems, and now I'll just ask quad code to prototype, like, three versions of it, And I'll try the feature and see whic…”
Opinion
Cherny: Claude Code delivers up to 10x productivity gains for Anthropic engineers
“Anecdotally for me, it's probably two X my productivity. So I'm just like, I'm an engineer that codes all day, every day. For me, it's probably two X. Yeah. I think there's some engineers at Anthropic where It's probably 10 X their productivity”
Insight
Kolter: Concentrated Foundation Model Usage Creates Systemic Correlated Security Exploits
“And especially when there's the possibility of correlated failures, right? So it's not just that there's a lot of AI systems out there, it's that there's actually a few models that everyone is using. And if you find vulnerabilities in the agents that everyone …”
Disclosure
Anthropic: Claude Cowork executes inside a dedicated lightweight Linux VM
“So we currently run like a, we currently run like a lightweight VM and we put clock code into the VM and we do that for a number of reasons. Safety and security is a big one, but even if you ignore for a second safety and security and you're just like, okay, Y…”
Assertion Supported
Shah: Supermemory outperformed Claude Code and OpenClaw benchmarks by almost 50%
“So the Claude code one performed the worst, and OpenClaw slightly more than that, and SuperMemory is the highest, and you can see that, you know, it's like a pretty significant difference, like almost 50%.”
Insight
Levie: Autonomous agents require hard-mode governance unlike current user-impersonating tools
“So far we've been in, in easy mode. We've hit the easy button with AI, which is the agent just is you. And when you're in cloud code and you're in cursor and you're in codex, you're just, the agent is you, you're offing into your services. It can do everything…”
Opinion
O'Laughlin: Claude for Excel is much worse than Claude Code with Python
“Cloud for Excel is much worse than cloud code using Python to use the Excel skills to then deposit into.”
Assertion Contradicted
O'Laughlin: Claude Code captured 4% of GitHub commits in two weeks
“I love watching exponential trends. And I've never seen one even remotely at this rate. You would art, you know, four percent in like two weeks.”
Assertion Supported
Nina Lopatina: Claude Code saturated Princeton's agentic research benchmark within weeks
“It's a set of benchmarks for really evaluating longer-running agentic tasks, and in this case, there was one where they were evaluating, recreating a research paper, and that benchmark came out in October, and it was saturated earlier this week.”
Prediction Not checkable as stated
Dax Reed predicts Cursor will move faster than Anthropic's Claude Code team.
“I mean, when I saw the acquisition, I was like, oh, this is actually more intense now because the cursor team is going to move faster than the cloud code team. I think that's my feeling given they have to, there's more pressure on them and they're smaller.”
Prediction Not checkable as stated
Zach Lloyd: Anthropic might launch a desktop agent harness like Warp
“I think even hearing the cloud code folks on your podcast, like I would not be surprised if Anthropic like launched a thing that looks a little bit more like warp where it's like a harness for running a whole bunch of different cloud codes, but it's an actual …”
Assertion Not checkable as stated
Cherny: Some Anthropic engineers rack up thousands daily running Claude Code automations
“And there's some people at Anthropic that have been racking up like thousands of dollars a day with this kind of automation.”
Assertion Not checkable as stated
Anthropic's non-coding product designer ships monorepo pull requests using Claude Code
“And she's landing PRs to our console product. So it's not even just, like, building on quad code. It's building, like, across our product suite in our monorepo.”
Assertion Not checkable as stated
Cherny: Claude Code is the thinnest possible wrapper over the underlying model
“All the secret sauce, it's all in the model and this is the thinnest possible wrapper over the model. We literally could not build anything more minimal. This is the most minimal thing.”
Assertion Not checkable as stated
Anthropic estimates Claude wrote 80% to 90% of the Claude Code codebase
“Probably near 80, I'd say.”
Assertion Not checkable as stated
Anthropic observes Claude Code API costs averaging roughly $6 daily per user
“Currently we're seeing costs around, like, six dollars per day per active user, and so it's, like, it does come out to a bit higher over the course of a month in Cursor but I don't think it's, like, out of band, and that's, like, roughly how we're thinking abo…”
Assertion Not checkable as stated
Anthropic uses Claude to rewrite Claude Code from scratch every 4 weeks
“We've rewritten it from scratch, yeah, probably every three weeks, four weeks or something, and it just like all the, it's like a ship of Theseus, right? Like every piece keeps getting swapped out, and just because quad is so good at writing its own code.”
Disclosure
Ludwig: Claude Code Has Overtaken Cursor Internally at Applied Intuition
“Cursor was, I think the hottest tool in the company for a good while. Now, Claude Code, I think, has taken the reign on that.”
Assertion Supported
Rieseberg: Claude Cowork is Claude Code running in a sandboxed virtual machine
“Cowork is cloud code running in a virtual machine with a little bit of padding, a little bit more guardrails, making it a little safer, a little bit more convenient for people who don't want to first open up the terminal when they go to work.”
Assertion Supported
Rieseberg: Claude Code and Cowork share unified GitHub-installable plugin format
“We do have skills as part of this container format, which was just called plugins. And plugins are available both for Cloud Code and Cloud Cowork, the same format. And you can install plugins. This works in Cowork today. You can basically say, I'm going to add…”
Assertion Not checkable as stated
Yegge: Anthropic is hiring over 100 people for Claude Code
“They're hiring like a hundred plus people for cloud code in the next, I don't know, month. I mean, like they're going wild and that's just cloud code.”
Assertion Supported
Martin: Claude Code operates entirely without codebase indexing
“Clock code doesn't do any indexing. It's just doing, quote unquote, agentic retrieval, just using simple tool calls, for example, using grep, to kind of poke around your files, no indexing whatsoever, and obviously works extremely well.”
Insight
Lambert: Context compression is crucial for long-horizon AI agents
“Compressing context, like that's not I don't think that's really a verifiable thing, but that being messed up, like that's a super crucial skill for long context actions and long longer tasks is just compressing well, and that's going to take some training nov…”
Assertion Not checkable as stated
Anthropic engineer built a Slack bot using Claude Code to automate PRs
“There was a really early version of Cloud Code many, many months ago, and this one engineer at Anthropic, Jeremy, built a bot that looked through a particular feedback channel on Slack, and he hooked it up to code to have code automatically put up PRs. With ju…”
Assertion Not checkable as stated
Claude Code's markdown parser was generated by Claude in two prompts
“And so the night before the release at like, 10 PM, I'm like, all right, I'm going to do this. So I just asked Quad to write a markdown parser for me. And they wrote it.
It wasn't quite zero shot, but after, you know, like maybe like one or two prompts, it got…”
Assertion Not checkable as stated
Cherny: Internal DAU for Claude Code went vertical across Anthropic
“Then we gave all the engineers and researchers at Anthropic access and pretty soon everyone was using it every day. And I remember we had this DAU chart for internal users and I was just watching it and it was vertical like for days”
Insight
Cherny: Coding agents require early failure detection over delayed intervention
“And what we find is that if the model is doing something wrong, it's better to identify that earlier and correct it earlier, and then you're gonna have a better time. If you wait for the model to just go down this like totally wrong path and then correct it 10…”
Insight
Cherny: Rich CLI development resembles cross-browser issues of the IE6 era
“So building in this way, it feels to me a little bit like a building for the browser back in the day where you had to think about like,
Internet Explorer six versus Oprah versus like Firefox and whatever.
Like you have to think about these cross-terminal diffe…”
Insight
Cherny: Claude Code is built as a composable Unix utility like grep
“We think of it as like a Unix utility. Right. So it's like the same way that you would compose, you know, grep or cat or oh, cat. Or something like this. The same way you can compose code into workflows.”
Insight
Cherny: Prompt Engineering Skills Determine Success With Coding Agents
“And one thing we find is that people that are really good at prompting models, From whatever context, maybe they're not even technical, but they're just really good at prompting. They're really effective at using code. And if you're not very good at prompting,…”
Disclosure
Cherny: I have not manually written a unit test in months
“So for example, like, I have not manually written a unit test in many months.”
Insight
Cherny: Agentic search trades tokens and latency for better security and accuracy
“So essentially at the cost of latency and tokens, you now have really awesome search. Without security downsides.”
Insight
Cherny: Mandating Docker sandboxing for AI coding creates too much developer friction
“Ideally, the thing that we want is to always run code in a Docker container, and then it has freedom and you can kind of snapshot, you know, with other kind of tools later on top, you can snapshot, rewind, do all this stuff. Unfortunately, working with a Docke…”
Assertion Not checkable as stated
Anthropic runs Claude Code in CI/CD to semantically lint GitHub pull requests
“For quad code internally in the GitHub repo, we have this GitHub action that runs. And the GitHub action invokes quad code with a local slash command. And the slash command is lint. So it just runs a linter using quad. And it's a bunch of things that are prett…”