why aren't all 8 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Trojanowski: AI coding agents remain worse than junior engineers over two-week projects
“Even now with coding, like, the agents are not yet they're not human level at being coherent over long periods of time. That's obvious because they can't code like a junior engineer on a project for two weeks. So they can't, they're, that's worse than a human …”
Prediction Not checkable as stated
Trojanowski: Closing the agent self-improvement loop will be close by year-end
“I think that closing the loop is gonna happen pretty fast.
I think you'll have, like, I don't know about the entire loop being closed, but I think you'll be
Relatively close by end of year.”
Prediction Not checkable as stated
Trojanowski: Scaling Verifiable Rewards Won't Reliably Automate Tax Returns
“I don't believe that if you were to train a model, you know, and you scale up the amount of pre-training computing, you scale up the amount of Like post training from like perfectly verifiable rewards that suddenly will output a model that will do a tax return…”
Prediction Not checkable as stated
Trojanowski: AI models will subsume agent harness engineering in 2-5 years
“I think that how long it will take to get swallowed up. I don't know exactly. You know, I think it's probably sub five years. I don't think it's sub two years. I think it's probably sub five years.”
Assertion Supported
Trojanowski: DeepSeek-R1 succeeded by scaling outcome supervision over process supervision
“If you look, you know, if you fast forward a bit and you look at the, like the DeepSeq R-one paper where they effectively laid out, you know, what I think all the labs were doing at that time, or at least OpenAI was doing in terms of you know, RLVR reasoning f…”
Assertion Not checkable as stated
Trojanowski: AI models are far worse at agent engineering than software
“Right now they are far, far, far worse at engineering agent systems than they are at engineering most software.
Far worse.
because by definition, like that kind of work, which is so novel, has not seen a large amount in their training data.
and so they have …”
Assertion Supported
Trojanowski: Over Three Million Accountants Work in the US
“It is, one of, if not the largest knowledge work profession in the country. There are over three million, you know, combined kind of accountants in the country.”
Assertion Supported
Trojanowski: Preparing a Form 1065 tax return takes 20+ hours of human labor
“Performing a 1065 can take a human, you know like 20 plus hours easily of actual work. I don't mean like it took them a day. I meant like literal sitting down work. And it could actually take much longer for very complicated returns.”