why aren't all 1,598 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 14 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Wang: Meta Agent Swarms Outperform Teams of 100 Engineers
“Like, I think we've seen, internally in meta cases where if you can develop the right agentic loop and you have the right eval or the right metric for the agents to optimize, you can have a swarm of agents accomplish more than, like, a team of a hundred engine…”
Assertion Contradicted
Cherny: Anthropic Can No Longer Demonstrate Prompt Injection on Opus 5
“So essentially, if you combine a well-aligned model, so this is, like, essentially three years of research into alignment, With a prompt injection classifier, which we run for all traffic, and what this is doing is it's based on Chrysola's mechanistic interpre…”
Assertion Supported
Cherny: Bun codebase was rewritten to Rust in 11 days using Claude
“And he had the model rewrite it from Zig to Rust. It was one prompt. It was a dynamic workflow, and a dynamic workflows are a feature in quad code that essentially let you orchestrate, you know, dozens, 100,000 of agents to do work productively. And it ran for…”
Assertion Not checkable as stated
Cherny: Claude routines do the maintenance work of hundreds of engineers
“And so now we have every day, maybe 20 or 30 of these routines. It's running across all of our code bases and It's not totally there yet, but we're on the path to fully automating the maintenance of our apps by doing this. And this is, again, hundreds of agent…”
Assertion Not checkable as stated
Tan: Returning to Coding with AI Yields 400x Output
“It was 13 years of not coding and then suddenly, boom, I'm doing about 400 X the amount of work that I was that year. The last time I was even sort of like two thirds of the time writing code.”
Assertion Not checkable as stated
Mukund Jha: Emergent invented multi-agent paradigms before academic papers were published
“There was a time when we sort of invented the multi-agent system. We invented memory. We invented like, how do we do agent to agent communication? How do you scale up test time compute? A lot of those things which like were sort of coming out, like we would di…”
Assertion Not checkable as stated
Friedman: Majority of Software Code Is Probably Already Written by AI Agents
“It's already probably the case that the majority of code being written is written by agents.”
Assertion Not checkable as stated
Claude Code increased Anthropic engineer productivity by 150 percent
“And since Quad Code came out, productivity per engineer at Anthropic has grown a 150%.”
Assertion Not checkable as stated
Claude generates between 70 and 90 percent of Anthropic's codebase
“If you look at Anthropic overall, it ranges between, like, 70 to 90% you know, depending on the team. For a lot of teams, it's also, like, a hundred percent. For a lot of people, it's a hundred percent.”
Assertion Partly supported
Srinivas: Google possessed only fourth or fifth best AI models in 2023-2024
“And 2024, 20 23 especially, and large part of 2024 too, Google had like, maybe a fourth or fifth best models at any moment. So as a startup outside Google, you had access to AI that was better than what Google internally had.”
Assertion Not checkable as stated
Ng: AI companies have made untrue statements to regulators behind closed doors
“I've been in the room where some of these businesses said stuff to regulators that was just not true.”
Assertion Not checkable as stated
Chollet: All inventive AI systems rely on discrete search
“All known AI systems today that are capable of some kind of invention, some kind of creativity, they rely on discrete search.”
Assertion Partly supported
Elon Musk emailed OpenAI saying it had zero percent success chance
“I can say this one because it got publicized in early, not early, a few years into OpenAI, where Elon sent us this really mean email. We'd been working through that for a while and said we had a zero percent chance of success. Like, not 0.10, that we were tota…”
Assertion Not checkable as stated
Tan: Sam Altman aims to scale AI compute to $1 trillion spend
“Since then, Sam has told me that he actually wants to go to four orders of magnitude to get to a trillion dollars in you know, sort of spend.”
Assertion Not checkable as stated
Seibel: VCs tell YC partners not to advise founders against high burn
“They also always tell us like, hey, like we're funding your companies. Like you shouldn't be like, you know, like telling the founders not to have high burn.”
Assertion Not checkable as stated
Allred: Lambda School grads are financially $250K ahead of university peers
“By the time the student's graduating from university, the Lambda School student has three years of experience and has paid Lambda School off and is a quarter million dollars ahead, at least, of the university grad.”
Assertion Not checkable as stated
Gil: A prominent Midas List investor tried clawing exit value at Mixer Labs
“The biggest issue we had there is we had a bad investor who, as we were exiting the company, basically tried to claw more value for themselves as part of our exit. And this is sort of a Midas list investor who's very well known now.”
Assertion Supported
Iorns: Replication studies show most published scientific results are not reproducible
“So the reality is that the results that have been generated show that most published results are not reproducible.”
Assertion Supported
Altman: Almost no top 30 software startup had remote early co-founders
“The data on this is look at the, say, like, 30 most successful software startups of all time, and try to point to a single example where the co-founders were in different locations. It's really, really tough.”
Assertion Partly supported
Wang: Meta rebuilt frontier AI lab from scratch after Llama 4 stalled
“Lama four wasn't on the trajectory that was needed for Meta. And so I got in there and we kind of did a zero based build of how do you know, build an entire frontier lab you know, in some ways kind of from scratch, obviously using a lot of what we had and move…”
Assertion Not checkable as stated
Altman: AI agents build in seven minutes what took startups three months
“What took three months to build at the time that we, each company built over the whole YC startup, could now be done with, like, in, like, seven minutes by a coding agent”
Assertion Open · timeframe Jul 2027
Cherny: Claude Opus 5 Can Run Autonomously for Months Without Scaffolding
“For five, one example of something it does that I think no other model has done is it runs for a very long period of time. And especially when you combine Opus Five with auto mode, it's just like incredible. Like it can go for days, weeks, months at a time. It…”
Assertion Not checkable as stated
Cherny: Claude Is More Intelligent Without System Prompts
“And what's interesting is that the model is actually a little bit more intelligent without these prompts. That's something that we've been finding.”
Assertion Not checkable as stated
Cherny: Modern models can rewrite essentially any codebase into another language
“One example is the model can now rewrite essentially any code base from one language to a different language.”
Assertion Supported
Tan: Jeff Lawson was ousted from Twilio 199 days after voting protections expired
“Our friend Jeff Lawson at Twilio built that company from nothing. To four billion dollars in, you know, actual revenue, like stock up 390% since IPO. I mean, by all accounts, you know, smash rip roaring success. And then his super voting shares expired after 1…”
Assertion Supported
Ries: Philip Morris generates $600B in annual US external costs
“They have something like eight billion dollars a year in net income. But there's been all these studies. They create six hundred billion dollars a year in costs just in the U.S. That have to be borne by others. I think it's three hundred billion in direct heal…”
Assertion Supported
Ries: Shareholder primacy was never enacted by statutory law or referendum
“In the history of the world. Shareholder primacy has never been subject to any referendum, any legislative action, nothing. So it's weird. If you learned in school, how a bill becomes a law, There's no law for shareholder primacy.”
Assertion Partly supported
Ries: Foundation-owned companies are 6x more likely to reach year 50
“Companies with this structure are six times more likely to live to year 50. 10% versus 60% probability.”
Assertion Open · timeframe May 2026
Graham: YC startups returning home are half as likely to become unicorns
“YC now has a lot of data about this, and the startups that go home after YC don't do as well as the ones that stay. Startups that go home back, startups that go back home after YC are only about half as likely to become unicorns”
Assertion Partly supported
Chaubard: HRM scored 70% on ARC Prize 1 without pre-training
“There is no pre-training at all. This starts from, like, literally Tagula-Rasa weights, and it can outperform at that time, if we go back, you know, we had O-three, if you remember back, way back when. And it, O-three gets zero. Literally zero, and this got, l…”
Assertion Not checkable as stated
Massad: Tech-adjacent creators get the most value from Replit, not engineers
“And what we've noticed is that people that are getting the most value out of our product, Tend to be the more tech-adjacent ones. Maybe people, product managers have written code many years ago, but don't worry about the development environment setup. Don't wa…”
Assertion Not checkable as stated
Tan: Rebuilt Posterous with AI, Replacing 10 Engineers and $10M
“Along the way, I've essentially built all of Posturus, which took two years to build with a co-founder and a team of 10 engineers. I've essentially built all of my startup Posturus, which took two years, ten million dollars, and 10 engineers to build.”
Assertion Not checkable as stated
Vuong: Models perform complex robotic tasks zero-shot, saving hundreds of hours
“Today it's possible to perform tasks Zero shot. Zero shot meaning you don't collect any data. And these are the tasks that last year might have required like hundreds and hundreds of hours.”
Assertion Not checkable as stated
Mellata: Variance's non-technical CSM ships features autonomously with Cursor
“Our customer success manager, who's entirely non-technical, but interfaces with enterprise customers on a day-to-day basis, now gets to take on feature requests, especially the simple ones, directly give them to Cursor, a Cursor agent, and then directly be abl…”
Assertion Not checkable as stated
Emergent tools enable one engineer to replace a five-person development team
“We also are internally seeing like the role sort of combining. So like a PM, a designer engineer, like a single person is doing, you know, like work of all three together, right? So like we have a PM who's By coding internally things. And recently like we so w…”
Assertion Not checkable as stated
Emergent lowers custom software development costs from $500,000 to $5,000
“And if you look at the price point that, you know, we are bringing down, it would have cost you like 500,000 dollars to build the software. Now you can build it for 5000 dollars completely on your own.”
Assertion Not checkable as stated
Hodak: Internal representations in AI models closely resemble biological brain patterns
“When you train AI models, like image models or, and even language models the representations that you get inside them look a lot like the representations you see in the brain.”
Assertion Not checkable as stated
Poetiq achieves faster, cheaper recursive self-improvement than existing methods
“The core insight that we had is that we could do recursive self-improvement far faster and cheaper than all of the other ways that people had been proposing to do this.”
Assertion Not checkable as stated
Poetiq's agentic harness outperforms new base models without code changes
“With poetic what we end up giving you is a you know, people are calling these things harnesses now, but you know, or agentic system or whatever you want to call it, that sits on top of one or more language models, and it just performs better than them. And whe…”
Assertion Not checkable as stated
Friedman: Tan replicated years of startup work in two weeks using Claude
“I feel like your real feel the AGI moment, Gary, was like getting Claude Code to build basically an entire startup for you, like replicating years of work of your previous startup in like two weeks, which is like insane.”
Assertion Not checkable as stated
Claude Code's plugin feature was built entirely by an autonomous swarm
“I think the first kind of big example where it worked is our plugins feature was entirely built by a swarm over, over a weekend. It just ran for like a few days. There wasn't really human intervention and plugins is pretty much in the form that it was when it …”
Assertion Not checkable as stated
Steinberger: OpenClaw agent resisted prompt injections in a public Discord test
“I just created a Discord, And I just put my bot without any security restrictions in the public discord. And then people came in, and they interacted with it, and they saw me build the software with it, and they tried to prompt inject it, and hack it, and my a…”
Assertion Not checkable as stated
Diana Hu: Supabase is the default backend recommendation across all LLMs
“Whenever someone asks how to set up anything that you need, some sort of backend, Firebase type of transaction, the default answer from all the LLMs is actually a Superbase.”
Assertion Not checkable as stated
Hu: Anthropic surpasses OpenAI as top API for YC applicants
“And, shockingly, in this batch, the number one API is actually Anthropic. It came out a bit more than OpenAI, which who would have thought?”
Assertion Supported
Tan: Boom Supersonic is using jet engines to power AI data centers
“Boom supersonic, instead of making supersonic jets right now, is on this good quest to create enough power for a bunch of these
AI data centers that are being built right now.
They use jet engines and even those like are so bad, you know, the supply chain for …”
Assertion Not checkable as stated
Hu: A YC startup's 8B model beat OpenAI on healthcare benchmarks
“There's this particular YCE startup that told me that they collected the best data set for healthcare. And they ended up performing better than OpenAI and a lot of the benchmarks for healthcare with only eight billion parameters.”
Assertion Not checkable as stated
Tan: GPT-4 wiped out YC startups' RL fine-tuning lead over GPT-3.5
“You know, we've also had YC companies where, ah, they had something that beat OpenAI, ah, you know, GPT-III.V and they were doing fine-tuning with RL, but then, ah, yeah, GPT-IV and then, ah, GPT-IV came out and, ah, you know, basically blew their fine-tuning …”
Assertion Not checkable as stated
Fisher: Large enterprises are replacing SaaS purchases with AI-assisted internal builds
“Large enterprises sometimes aren't even going out and buying software from SAS providers anymore. They're like, I can just like Throw two people at cloud code and they'll build it and it'll be dedicated to the capabilities that I need for my organization.”