why aren't all 790 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Logistics Data Matters More Than Frontier LLM Intelligence in Lab Science
“I think whether GPT, 5.2 codex max or Opus 4.5 is going to do better. It's probably doesn't matter. It's just a matter of like, which one's going to have all the information about what's in the lab and how much will it cost? How long will it take?”
Opinion
White: New AI laboratory risk scenarios have become realistic
“To some extent, there was like a first wave that we thought this could unlock a lot of stuff and I don't think it came to pass. I think there's now an emerging sort of second wave of like, there are some actually new scenarios that were just too farfetched to …”
Opinion
Yi Tay: Today's Models Likely Couldn't Invent the Transformer from Pre-2015 Data
“Even today's models, they might not even be able to invent the transformer. Like, if you freeze the time at a certain time, and even you bring the time, I mean, the model is a transformer, so I just say there's no, assuming there's no leakage.”
Opinion
Yi Tay: The AI Industry Is Stuck in a Transformer Local Minimum
“So now we are like in this local minima of like transformers, everything, everything, right? Maybe it's not easy to like get totally out Of this, because also a lot of people's investment optimization have been done. So the things that play well needs to play …”
Opinion
Yi Tay: AI Research Has Not Entered Diminishing Returns on Ideas
“The number of ideas that actually work is not decreasing compared to the last, like, we're not in the era of diminishing returns yet.”
Opinion
Yang: Unit tests and independent task instances limit SWE-bench's effectiveness
“I don't like unit tests as a form of verification, and I also think there's an issue with Sweepbench where all of the task instances are independent of each other.”
Opinion
Merullo: Current machine unlearning techniques merely suppress data rather than removing it
“I would describe it more as not unlearning, but maybe suppression. I think there's, like, really, like, I guess, guarantees that you've fully removed information from a model is, is, I don't think it's been convincingly showed anywhere yet”
Opinion
Nair: 2017–2022 academic RL breakthroughs failed because researchers overfit to benchmarks
“A lot of the methods that people were really excited about is, like you know, off policy learning, like, value functions, like, these kind of things, and somehow that, that stuff hasn't really panned out, I would say, and it's not exactly clear why, but in the…”
Opinion
Nair: OpenAI model splits happen because it ships its org chart
“OpenAI has a tendency to ship the org chart, basically.”
Opinion
Nair: Frontier AI labs have converged on similar reinforcement learning methods
“Well, it does seem like basically a lot of the labs have kind of like converged onto some similar-ish way of doing RL, and they're all kind of back at the same level of like Frontier again”
Opinion
Other agent protocols currently lack the adoption required for AAIF inclusion
“At least on the protocol side, like a defector standards. And I don't think any of the other protocols, it feels like they're not just there yet, but of course, if they get there, then we are like super open as long as they're like complementary to what's, wha…”
Opinion
Yegge: Graphite is best poised to solve the AI code merge wall
“Merging is the, it's the wall that everyone is hitting right now. I think the company that's best poised to solve it is graphite.”
Opinion
Yegge: Google, Anthropic, and OpenAI are unbelievably chaotic internally
“All three of those companies, Google, Anthropic, and OpenAI are unbelievably chaotic internally right now.”
Opinion
Pliny: Guardrails degrade AI capability and creativity relative to model size
“I do think they're finding clever and clever ways to lock down particular areas sometimes, but I think it's at the expense of capability and creativity. So there's some model providers that aren't prioritizing this and they seem to do better on benchmarks for …”
Opinion
Fanelli: VC Cycles Often Conflict with Real Security Progress
“Once you're in the VC cycle, you kind of need to do things that then get you to the next round. And I think a lot of times those are Opposed to doing things that actually matter and move the needle in the security community.”
Opinion
Gameplay footage teaches spatial reasoning better than YouTube videos, says de Witte
“When you're playing video games, you're actually simulating the optical dynamics with your hand, right? And I think that, like, that's why I think why games are a better representation of switch support reasoning initially than YouTube videos, for instance.”
Opinion
General Intuition publishes openly because its proprietary data moat prevents model replication
“Because we have such a large data mode, we don't have to be as concerned as the LLM companies about publishing, because we don't need the ones to be able to. Exactly. No one can replicate the models, right?”
Opinion
Hezarkhani: MCP is essentially just a three-letter word for API
“I just think that MCP is a three-letter word for API”
Opinion
Jed Borovik: AI coding tools will not reduce software engineer hiring
“I don't really Buy this, like, you know, that we're not gonna hire more software engineers story. I think, like, for a few reasons I mean, this is an example that, that often comes up, but is it, like, kind of the elasticity of the demand for software?”
Opinion
Zuckerberg: EvolutionaryScale is the most talented AI and biology team
“I mean, this is like probably the most talented team working on AI and biology, right? And like at the intersection of doing of like basically good biology background and also, you know, they've just been working on ESM three.”
Opinion
Groq hardware is too inflexible to run alternative architectures like Mamba SSMs
“Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like Designed it at the hardware level, right?”
Opinion
Ubl: Vercel leads the truly open-source deployment business model
“Vercel maybe has not invented this, but certainly kind of is the most successful at a model where you say, okay, I have this software library and it's truly open source. Everyone can run it. It comes with like adapters for every place on the planet, and that m…”
Opinion
Webster: AI evaluation tools are table-stakes commodities facing a feature-parity bloodbath
“I think evals are our table stakes. I think that they're a commodity and everyone should be doing them. And yes, there are companies that are doing great in the eval space, but To me, it just seemed like a bloodbath, you know, like we would just be, had a grea…”
Opinion
Corbitt: Fine-tuning offers poor ROI for 90% of unconstrained use cases
“I would say for 90% of use cases where you aren't forced to a smaller model, then it's still not a good ROI, and you probably shouldn't invest in it today.”
Opinion
Lenz: Most enterprises avoid reasoning models due to high latency
“Most enterprises don't really want to use reasoning models. The latencies is too high”
Opinion
Dwivedi: Claude is superior at agentic tool calling and error unstacking
“For some of the agentic part of the stack, we are shifting towards Anthropic because they're agentic and the tool calling, especially the unstacking part, you know, when you go down the wrong path and you build context that forces you to keep going down the wr…”
Opinion
Field: Vibe-coding will not easily replace complex enterprise software like Workday
“A lot of CIOs would, You know, you love to go, okay, yeah, I've vibe code workday and I just saved my company all this money or whatever, but okay, you actually peek into the hood of a workday or a rippling. These are very complex pieces of software that have …”
Opinion
Field: The AI industry is showing get-rich-quick energy reminiscent of NFTs
“The parallel that I think is kind of interesting is, you know, you compare AI to that, and it's like, there's been a long era of people that I think are very on mission and thinking about the big picture, the risks, the opportunities, the possibilities. And th…”
Opinion
Falk: Decagon and Sierra will always outperform generalist enterprise AI builders
“I can just guarantee that Decagon is always going to be better. Sierra is always going to be better because they spend all day, every day thinking about it.”
Opinion
Taskaya: Flux was the first enterprise-ready generative image model
“The team at Stability left to start Black Forest Labs, which released Flux models. And that was the first model to, you know, breach the barrier of commercially usable, you know, enterprise-ready grade models”
Opinion
Taskaya: Custom ASICs do not make sense given low Nvidia GEMM overhead
“What is the overhead of an NVIDIA GAM instruction, right? It's like 16%. So like you're essentially buying a, Matrix multiplication machine. So, like, it doesn't really make sense to specialize it that much.”
Opinion
Fal CEO: Newer Video Models Are Now Much Better Than OpenAI's Sora
“Now we have video models that are much better than Sora.”
Opinion
Morcos: Data is AI's most under-invested research area relative to impact
“Something I've said before and I'll say again is, is that data is the most under-invested in area of research relative to its impact, and I don't think it's even close.”
Opinion
Morcos: Modern AI capabilities depended entirely on self-supervised learning
“But I do not think there's any way we could get to where we are today without self-supervised learning and the ability to train on unlabeled data. That was the real advance to my mind that enabled us to get these incredible increases in capabilities.”
Opinion
Morcos: Frontier AI lab data teams are systematically under-resourced
“I think you, what you see in all the frontier labs is that they have data teams. And if you talk to the folks that work on those data teams, what you'll kind of systematically hear is that typically they're under resourced relative to the gains that they're de…”
Opinion
Morcos: AI training is commoditized while data curation remains hard
“Mosaic was the first one to really recognize that there was a huge opportunity in making this easy. And now this has largely been commoditized by things like SageMaker and Together and lots of different folks that help you on the training side. But on the data…”
Opinion
Huber: Regex handles 90% of code queries; embeddings add marginal improvement
“My guess is that, like, for code today, it's something like, 90% of queries or 85% of queries can be satisfactorily run with regex. Regex is obviously, like, the dominant pattern used by Google code search, GitHub code search, but you maybe can get, like, 15% …”
Opinion
Brockman: AGI will be a menagerie of models, not a single model
“The flip side, though, is that I think that the evidence has been away from having the final form factor, the AGI itself being a single model. But instead thinking about this menagerie of models that have different strengths and weaknesses.”
Opinion
Brockman: AI infrastructure buildout dwarfs Apollo program and New Deal
“The engineering project that we collectively as a country, as a society, as a world are undergoing right now, right? It's like projects like the New Deal, like pale in comparison, you know, the Apollo program pale in comparison to what we're doing right now.”
Opinion
Dax Reed claims the Gemini CLI codebase is rushed and corrupts edits.
“And Gemini ones were not good. It was initially it looked good, but then I noticed it was getting caught in these crazy loops and it was like editing my files in these totally messed up ways. And I should have known because you can tell when you're looking thr…”
Opinion
Ganatra: Existing AI tool-calling benchmarks are not production-ready
“Tool calling has a lot of different evals. You have seen the function tool calling evals. You have multi-function, like, multi-tune tool calling evals. But, like, none of them are actual, like, production quality. Like, none of them are actually something that…”
Opinion
Lambert: Deep Research relies on modular RL tasks rather than end-to-end outcomes
“I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof, the system works. Which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that whe…”
Opinion
Mohan: Windsurf's real value is large codebase work, not 0-to-1 apps
“We had the technology to go out and build these zero to one apps very quickly, and I think people are using Windsurf to actually do that, and it's like extremely impressive, but the real value, I think, is actually much deeper than that. It's actually that you…”
Opinion
Fanelli: MCP integrations cannibalize SaaS companies' own paid products
“I feel like in a way the MCPs are like, Cannibalizing the products themselves.”
Opinion
LLM-as-a-judge methodology is now reliable enough for production evaluation
“Two years back where we were using our own similar LLMS judge methodology, and we had certain issues and there were some correlation issues at the time, but last year it got pretty strong and this year I feel it's so strong. LLMS judges is so strong that you c…”
Opinion
GPT-4o Mini cannot accurately evaluate complex Claude 3.7 outputs
“GPD for a mini can't find, ah, cannot evaluate the hard outputs that's 3.7 might be doing correctly, right?”
Opinion
Frontier LLMs remain unreliable at realistic multi-turn tool calling
“Our last leaderboard is saying that models are really great at tool calling. So it's like safe, but they actually not, right? They're making mistakes and this is going to recalibrate the expectation of the users that Be careful because they're still not perfec…”
Opinion
Ben Holmes: LLMs may eventually eliminate the need for React Native
“My hot take was, I don't know if we need things like React Native in the future, because you can just ask these LLMs to translate between different native environments.”