why aren't all 7 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Solving Autonomous Coding Is the Direct Path to AGI
“Our core belief is that if you solve this problem, you solve the autonomous coding problem and build a super intelligent coding agent, that that thing will lead to super intelligence more broadly.”
Opinion
Software Engineering Is Ergonomic for LLMs, Making It the Ideal Wedge
“Our belief as a company is that the correct wedge in, the correct starting point to this entire problem is decoding agent, because it's already, you know, software engineering is already what I would call kind of ergonomic for a language model.”
Opinion
Long-Context Attention Will Beat Agentic Localization for Codebase Indexing
“I bet would be on long context understanding and improving the attention mechanism over the long context.”
Opinion
Needle-in-a-Haystack Tests Are Crude for Evaluating True Context Understanding
“So it's not just having long context, it's whether your model truly understands what's inside the context, and needle in the haystack tests are pretty crude and not very effective way of testing this kind of capability.”
Insight
A 90% SWE-Bench Score Can Still Fall Flat in Customer Environments
“Autonomous coding benchmarks, let's say, like Sweetbench, are useful. I'm not going to discount them. They are useful. But let's say, you know, 90% on Sweetbench could still mean something that just falls over flat within a customer setting.”
Prediction Not checkable as stated
AI Coding Agents Will Discover Unexpected 'Move 37' Breakthrough Solutions
“I think I think there are going to be a lot of move 37”
Insight
Reinforcement Learning Fails Without Initial SFT to Seed Rewardable Behaviors
“It can potentially work otherwise, but practically it only works when the agent has interacted with a reward, right? It's received a positive reward for what it's done. Maybe one out of 10 times, one out of 50 times, but if it's getting zero reward, then you d…”