why aren't all 5,957 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 100 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Webster: AI guardrails are a commodity and easy to build
“Building a business on guardrails really scares me because there are so many incumbents that can come in and eat your lunch. And it's like, It's not actually that hard to build a guardrail. I don't know if I'm gonna make people angry by saying that maybe some …”
Opinion
Corbitt: GRPO is likely a dead end due to parallel rollout constraints
“The big downside, the huge downside of GRPO, and I think actually the reason why GRPO actually is likely to be a dead end, and we probably will not be continue using it indefinitely. The fact that you need to have these parallel rollouts in order to train on i…”
Assertion Not checkable as stated
Corbitt: Prompt optimization methods like JEPA failed OpenPipe's agent benchmarks
“It didn't work on the problems we tried it on. It just didn't. It got like a minor boost over the sort of like more naive prompt we had and was just like, it was like, okay, Just kind of like our naive prompt with our model gets maybe like 50% on this benchmar…”
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year.
And my money is on, they have to do a coin.
Like it's, I'm not a crypto guy at all, but like, y…”
Opinion
Lenz: Model providers should not dictate enterprise AI policies
“Right now, if you're using a model, you're taking in their own policy. Even if I want to use GPT-OSS, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave, etc. And I don't thi…”
Opinion
Agarwal: No AI incident troubleshooting competitor genuinely works in production
“And I don't think we've seen any other company in our space having something actually work at production.”
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Disclosure
Ball: Amp core team has 8 people, ships 15 times daily without code reviews
“I think we're around eight people now on the AMP core team, and we still don't do formal code reviews. We still push to main. We still ship 15 times every day.”
Prediction Not checkable as stated
Slack: AI tools like Cursor and Claude Code will peak and decline
“A lot of these other tools that are great, like cloud code and codex and cursor and so on that they've forgotten what made them great and what made them grow so fast, which is building the very best product. And they built it in a way that's too overfit on the…”
Prediction Not checkable as stated
Slack: Top AI labs face a major customer stampede within two months
“I think we are one or two months away from a possible news cycle. That is the foundation model companies have spent billions of dollars in capex and hired like crazy. And now, you know, they're no longer the best in this realm and there's a huge stampede away …”
Prediction Not checkable as stated
Ball: Complex sub-agent workflows will result in user hangovers
“A lot of the features what we see, you know, where people build like elaborate workflows, like I have my custom slash commands and they trigger custom Custom sub-agents and they in turn trigger custom MCP tool calls behind which again another model is doing in…”
Opinion
Slack: AI prompt enhancers are a 'bullshit feature' that does not work
“Yeah, so prompt enhancer, that's a bullshit feature that doesn't actually work. The theory behind it is nuts, because what helps LLMs is not tricks and phrasing your prompt in a certain way. It's fundamentally information that you have in your head that you ca…”
Opinion
Slack: MCP as user-facing tech creates high token costs and bad UX
“As a user facing technology, it is such a common failure mode where a user will go and add in some MCP servers. Auth is a huge pain, but let's say they get over that hurdle. Then they have, I don't know, 50 tools exposed that often are too low level granularit…”
Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer
With a context length of, you know, 256,000 or more, they're not using a true transformer.
What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Opinion
Gorkem: Hosting LLMs is a bad business due to Google search competition
“Language models, hosting language models is not a good business. At the time we thought, okay, we are going to be competing against OpenAI and Anthropic and all these labs. Turned, turned out that it was even worse because the killer application of language mo…”
Opinion
Morcos: The Transformer is just one of many equivalently good architectures
“And one of my like more controversial viewpoints, I think, is that I think the transformer is a great advance to be sure, but I think it's one of a very large Set of equivalently good architectures that we could have found. And there are many, many ways we cou…”
Opinion
Morcos: GPT-4.5 and Llama 4 show limits of naive mega-model scaling
“And I think that's what we've seen to some extent with the failure of the mega models, right? With 4.5 and Lama four and others. I think that there is a challenge of just continuing to do that naively and you have to figure out how to break it.”
Insight
Morcos: Post-training alignment is ineffective long-term compared to pre-training alignment
“Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to tak…”
Opinion
Huber: Silicon Valley treats AGI as a secular religion
“I think AGI is also a religion. It has a problem of evil. We don't have enough intelligence. It has a solution, a deus ex machina. It has the second coming of Christ that AGI, the singularity is going to come. It's going to save humanity because we will now ha…”
Opinion
Sohmers: Hardening silicon for specific AI models is obsolete in months
“Doing any of that, like, hardening for specific Model things. I don't think lasts more than, you know, two or three months at the rate that the industry moves at.”
Opinion
Agrawal: Cerebras and Groq offer cloud APIs because their software struggles
“You know, if you look at Cerebrus and Grok and others, they've really tried to do this, their cloud kind of portal. And for us, you know, it shows two things. One is the difficulty of software for third party to implement that, that that's why they're kind of …”
Prediction Not checkable as stated
Ermon: Diffusion models could become the dominant architecture over autoregressive models
“I'm pretty optimistic about a future where diffusion models
Can become the dominant solution. I've seen it happen before with GANs a few years ago, so I wouldn't be surprised if that's the case also here.”
Opinion
Mohan: Bearish on outsourced eval startups because AI companies must own evaluations
“And I guess maybe one of the things I'm a little bearish on is If another company comes out and solves eval properly for a bunch of different verticals, what was the company that they were selling to really doing? What are they really doing at that point? If t…”
Opinion
Ramachandran: SWE-bench and HumanEval do not reflect real professional software engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have SweeBench, that's cool, no actual Professional work looks like Sweebench, like human eval, same thing.”
Prediction Not checkable as stated
Hou: 99% of AI editor rules file contents will be automatically inferred
“We strongly believe that having a rules file, you know, we do allow users to add a rules file, we strongly believe that a rules file is a crutch. You know, by the end of twenty-twenty-five, 99% of the things that you're gonna put in a rules file will be interp…”
Prediction Open · timeframe Dec 2029
Kamradt Predicts ARC-AGI-3 Benchmark Will Remain Unbeaten For 3 Years
“And then V three, our durability estimate for that is three years. And that's what we're aiming for is 36 months for V three.”
Opinion
Kamradt: Any AGI Definition Involving Profit Has Ulterior Motives
“Any AGI definition that involves money has ulterior motives. I mean, simple as that. Money has nothing to do with intelligence, right?”
What-if
Morris: ChatGPT could likely have been built using RNNs instead of Transformers
“And I think like, we honestly probably could have gotten this with RNNs. I know like the scaling laws paper shows that RNNs have worse curves for scaling, but probably people would have been like, I bet you could have built chat GPT with a very sophisticated R…”
Opinion
Zach Lloyd: Warp's coding agent outperforms both Cursor and Windsurf
“Warp can code, and Warp code is at a level that's, like, comparable to cloud code. I think it's actually better than, like, Cursor's agent. It's probably better than Windsurf since they aren't, they don't have Sonnet for.”
Opinion
Brown: LLMs implicitly develop world models through scale alone
“I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.”
Opinion
Brown: AI models implicitly develop theory of mind through scale
“If these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that like can take actions and have motives and all this stuff. And these models just develop that implicitly with scale and…”
Opinion
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Opinion
Lattner says C++ 'sucks' for modern accelerated compute
“Let me be the first to tell you, and I can say this now, I feel comfortable saying this, that C++ sucks.”
Opinion
Lattner calls vLLM a 'hot mess' due to too many stakeholders
“VLM seems much more like a massive community with a lot of stakeholders, a lot of stuff going on, and it's kind of a hot mess.”
Opinion
Ameisen: Stochastic Parrots Label Ignores Complex Multi-Step LLM Reasoning
“It's, like, activating many different distributed representations, like, combining them, and sort of, like, doing something pretty complicated. And so, yeah, I think it's funny, because in my opinion, that's like, yeah, like, oh god, stochastic parrots is not …”
Insight
Optimal future AI coding interfaces will not evolve from traditional IDEs
“Our take is that it is very unlikely that the optimal UI or the optimal interaction pattern for this new software development where humans spend much less time writing code. I think it's very unlikely that that optimal interaction pattern will be found by iter…”
Insight
Embiricos: Scaffolding-heavy AI agents are limited by developers' mental capacity
“A lot of, like, agents that I see are really impressive, but it's basically, like, part of what's impressive is it's like a bunch of developers building this, like, really bespoke state machine around a bunch of, like, short model calls, and so then the upper …”
Opinion
Bergum: The standalone vector database infrastructure category is dying
“I'm not saying that the companies are dying, right? I'm just saying that the separate infrastructure category is dying, right? Because you have vector search capabilities in almost any DB technology nowadays, right?”
Prediction Not checkable as stated
Bergum: Pinecone will not endure like MongoDB because it is too narrow
“So there's always like this convergence, but MongoDB kind of, it sticks, but I don't think that for Pinecoin that was originally leading that movement, it won't like stick in the same way.
It's too narrow.
It's too, too narrow.”
Prediction Not checkable as stated
Conrad: Hyperscalers will probably lose significant money reselling Nvidia GPUs
“My intuition is that the hyperscalers are probably going to lose a lot of money, and they know they're going to lose a lot of money on reselling NVIDIA GPUs at least.”
Prediction Not checkable as stated
Conrad: Decentralized compute networks will never beat co-located InfiniBand clusters
“I just don't really think this is gonna ever be more efficient than a fully interconnected cluster with Infiniband, or, you know, whatever sort of next spec might be. Like, I could be completely wrong, but Speedolite is really hard to beat. And regardless of w…”
Prediction Not checkable as stated
Conrad: The AI VC bubble will pop and fail to return capital
“So what you've done by not having a future is you've inflated the venture capital market. And that is a bubble that's totally going to pop at some point. Like a lot of the companies are not going to work. And the valuations are not going to work. And what's go…”
Opinion
Colvin: Agent framework engineering quality lags far behind Python ecosystem
“Looking in general at the ecosystem of agent frameworks, the engineering quality is far below that of the rest of the Python ecosystem.”
Insight
Beauchamp: ML model performance plateaus near human-level on an S-curve
“With machine learning, first of all, you see that the performance of the models follows an S-curve. So it's not like it just goes off to infinity, right? And the S curve, it kind of plateaus around human level performance.”
Prediction Not checkable as stated
Beauchamp: Timeline for AGI has certainly been pushed out
“I think the odds, the timeline for AGI has certainly been pushed out, right?”
Assertion Not checkable as stated
Zhang: Meta Failed at Training MoE Models for Llama Series
“The reason why Lama open-sourced the MOE model, because I think they tried to train our MOE model, but they failed. So that, that's why they didn't open source MOE mode for Lama series.”
Assertion Supported
Bryk: Perplexity and ChatGPT Search rely on legacy Google and Bing APIs
“So these systems, there are a few of them now they basically rely on like traditional search engines like Google or Bing, and then they combine them with like LLMs at the end to, you know, output some power graphics answering your question. So they, Like, Sear…”
Opinion
Comfyanonymous criticizes Gradio for coupling interface logic with backends
“Yeah, Gradio, I don't like Gradio. It's bad. Like, the, that's one of the reasons why, like, Automatic was very bad. It's great because the problem with Gradio, it forces you to, well, not forces you, but it kind of makes your interface logic and your back-end…”