why aren't all 790 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Brown: Previous multi-agent research was heuristic and ignored Bitter Lesson
“I think that a lot of the approaches that have been taken have been very heuristic and haven't really been following like the bitter lesson approach to scaling and research.”
Opinion
Brown: Scaling self-play beyond zero-sum games will not be as easy as AlphaGo
“My point is that, like, this is where the AlphaGo analogy breaks down. And, not necessarily breaks down, but, like, it's not going to be as easy as self-play was in AlphaGo.”
Opinion
PyTorch democratized model training, but AI inference remains undemocratized
“And so things like PyTorch came on the scene and I think PyTorch gets all credit for democratizing model training, right? It's taught to pretty much every computer science student that graduates. That's a huge deal, but nobody democratized inference. Inference…”
Opinion
DeepSeek's open releases pulled global AI progress forward by six months
“But I think that what it did is it pulled forward progress in AI by like six months.”
Opinion
Ameisen: Current LLM Chain of Thought Is Unfaithful and Untrustworthy
“So I think there's like a sense in which right now the chain of thought is, is unfaithful, or at least you can't read the chain of thought and trust that that's how the model did it.”
Opinion
Docker stopped innovating after Solomon Hykes left the company
“At some point Docker basically stopped innovating and I left and the ecosystem continued around containers, but it didn't actually Pick up where we left off. Everything towards applying container tech to development kind of stopped.”
Opinion
Hykes: AI agent tooling is trending toward proprietary vertical monoliths
“I'm seeing us heading in the direction of highly integrated, very vertical end to end monoliths, you know, and I don't want to name names, but it's just, it's a, it's definitely a market trend.”
Opinion
Dockerfile was a 2013 stopgap and Compose was a cloned prototype
“Docker file was something we designed as a stopgap prototype. Thinking, oh, we'll clean this up later in, in 2013. You know, it's been more than 10 years. Compose was a clone of a clone that we acquired into the team and, like, stitched on top, and then as soo…”
Opinion
Enterprises demand full AI delegation, not 15% to 20% speed gains
“I think that the product experience of delegation is really, really immature right now. And most enterprises though, see that as the holy grail, not like going 15% or 20% faster.”
Opinion
Cherny: Claude Code delivers up to 10x productivity gains for Anthropic engineers
“Anecdotally for me, it's probably two X my productivity. So I'm just like, I'm an engineer that codes all day, every day. For me, it's probably two X. Yeah. I think there's some engineers at Anthropic where It's probably 10 X their productivity”
Opinion
Sobo: AI editor moats are shrinking rapidly due to better tool-calling LLMs
“So there's a, there was a lot of like integration work to bridge that gap, but now the LLMs, because they've taken on this tool calling and are getting better at tool calling, the story is the integration story has definitely gotten a lot easier. I think it's …”
Opinion
Sobo: Zed's competitive moat forces competitors to build an editor from scratch
“I mean, to me, the moat Z's moat is like implementing really nice product experiences with technical excellence and their performance, like just delivering a really great user experience that in order to compete with, you would have to Essentially build your o…”
Opinion
Bergum: pgvector is outpacing some standalone vector DBs in search capabilities
“So actually, what you can see, PJ Vector is doing more in the capabilities of vector search than some of the real vector database players, right?”
Opinion
Bergum: In-database ML like PostgresML is the wrong architectural direction
“No, I'm not. I'm sorry. I, I'm not. I think this is, yeah, we also seen other players that tries to, you know, move a lot of the logic into the database, agentic embedding inference and whatnot. I think if the right direction is to Keep infrastructure a little…”
Opinion
Conrad: SF Compute is the only true liquid bid-ask GPU market
“Turned into what is today SF compute, which is a compute market, which we think we are the functionally the most liquid GPU market of any capacity. Honestly, I think we're the only thing that actually is like a real market that there's like bids and asks and t…”
Opinion
Shah: The AI industry does not need a new language to replace Python
“He was talking about like, oh, we need a different language than Python or whatever that is like built for built for AI and built. It's like, No, Brett, I don't think we do actually. It's just fine. It deals with just fine, just expressive enough. And it's nic…”
Opinion
Swix: Descript still sucks despite major funding and OpenAI backing
“Descript is so much funding. They had OpenAI invested in them, and they still suck.”
Opinion
Swix: Most podcasts are bad because creators lack real-world experience
“The reason that most podcasts or YouTube videos are shit is they're made by people who don't have life experience, who are not that important in the world. They're not doing important jobs. And so what you want to actually enable is CEOs to each of them make t…”
Opinion
Husain: 80% of LLM-as-a-judge implementations are unhelpful
“I feel like a 75% LMS judge because it's low effort is kind of easy, but I would say out of the 75%, 80% is not helpful.”
Opinion
Klein: OpenAI's brand constraints will prevent them from offering CAPTCHA solving
“I think it's going to be really hard for a company like OpenAI to do things like support CAPTCHA solving or like have proxies. Like, I think it's hard for them structurally. Imagine this New York Times headline, OpenAI CAPTCHA solving. Like, that would be a pr…”
Opinion
Embeddings will not scale for cross-session memory in real-time AI systems
“I don't think that like embeddings are going to be able to scale to, I think they work well for some of this as like kind of the MVP version of the experience, but I think you're going to need a different experience and it's Probably something like really smar…”
Opinion
Swix: Deep research agents are the first agent category with true PMF
“What are the hard problems in this brand of agent that is like probably the first real product market fit agent. I will say more so than the computer use ones. This is the one where like, yeah, people are like, yeah, easily pays for 200 dollars worth a month w…”
Opinion
Sutin: The Humane AI Pin failed due to weight and overheating
“I think the Humane is like pretty incredible. Some of the engineering they did, but like, it wasn't kind of geared towards solving the problem. It was just it's too heavy. The swappable batteries is too much demand. The heat, the thermals is like too much.”
Opinion
Startups Cannot Profitably Pre-Train Foundation Models Due to Massive Compute Costs
“I think the idea that a company can make money sort of pre-training a foundation model is probably not true. It's hard to, you're competing with just you know, unreasonably large capex budgets.”
Opinion
Bret Taylor Mocks Using AI to Write Unsafe, Slow Python Code
“I think that we're bringing the cost of writing code down to zero. So the fact that we're still writing Python with AI cracks me up just cause it's like literally was designed to be ergonomic to write, not, not safe to run or fast to run.”
Opinion
Colvin: Jupyter Notebooks Are Worse Than or Similarly Bad to Excel
“I have very strong opinions about, you know, proper, like Jupyter notebooks, this idea that like you have to run the cells in the right order. I mean, a whole bunch of things. It's basically like worse than Excel or similarly bad to Excel.”
Opinion
Beauchamp: Crypto's only killer use cases are gambling and evading regulations
“As far as a gambling device is like the most fun form of gambling invented in like ever. Super fun. I thought as a way to evade monetary regulations and banking restrictions, I think it's also absolutely amazing. So it has two like Killer use cases.”
Opinion
Beauchamp: AI content should be created and trained by users, not Silicon Valley
“The number one problem for users in AI is this. All the AI is being generated by middle-aged men in Silicon Valley, right? That's all the content. You're interacting with this AI. You're speaking to it for 90 minutes on average. It's being trained by a middle-…”
Opinion
AMD GPUs perform horribly on Windows for generative AI tasks
“AMD works horribly on Windows. Like, on Linux, it works fine. It's lower than the price equivalent NVIDIA GPU. But it works, like, you can use it, generate images, everything works. On Linux, on Windows, You might have a hard time”
Opinion
Swyx: DeepMind has roughly a four-year advantage over OpenAI in world modeling
“So like they have maybe four years advantage on world modeling that OpenAI does not have. Cause OpenAI basically only started Diffusion Transformers last year when they hired Build Peebles. So DeepMind has a bit of advantage here.”
Opinion
Soldani: Web blocking disproportionately benefits incumbent closed AI labs
“And I think the problem is this blocking or ideas really, it impacts people in different ways. It disproportionately helps companies that have a head start, which are usually the closed labs, and it hurts incoming newcomer players where you either have now to …”
Opinion
Robinson: YOLO architectures have hit a performance plateau
“So, for years, yellows have been the dominant way of doing real time object detection, and we can see here that they've essentially stagnated. The performance between 10 and 11 is not meaningfully different. At least, you know, in, in this type of high level c…”
Opinion
Korupati: Vision-Language Models Are Lagging Behind LLMs in Reasoning
“LLMs are showing enormous progress in reasoning, especially with the latest set of models that we've seen, but we're not really seeing, I have a feeling that VLMs are lagging behind, as we can see with these tasks that should be very simple for a human to do t…”
Opinion
Schluntz: AI robotics today is where autonomous driving was 10 years ago
“I think where we are right now is where self-driving cars were 10 years ago. I think we have very cool demos that work. I mean, 10 years ago, you had videos of people driving a car on the highway, driving a car, you know, on a street with a safety driver, but …”
Opinion
Schluntz: High vehicle costs make Waymo's per-car profitability doubtful
“Those cars are expensive. It's not about if you can hit profitability, it's about your cash conversion cycles. Like is building one Waymo, like how cheap can you make that compared to like how much you're earning sort of as the equivalent of what an Uber drive…”
Opinion
Polu: DeepMind IMO breakthrough relied on scaling RL and autoformalization
“I think the DeepMind team just did a good job of scaling. I think there's nothing too magical in their approach, even if it hasn't been published as a Dan Silver talk from seven days ago, where it goes a little bit into more details. It feels like there's noth…”
Opinion
Polu: Anthropic split was driven by disagreement over OpenAI's API commercialization
“What I understood of it is that there was a disagreement of the commercialization of that technology. I think the focal point of the disagreement was the fact that we started working on the API and wanted to make those models available through an API. Is that …”
Opinion
Pullen: SWE-bench is a poor proxy for real-world AI coding competence
“I know Sweebench is, like, the most commonly talked about thing, and honestly, it's a very, it's an amazing project, but one of the things we've learned the most from actually shipping this product to users is, it's a pretty bad proxy at telling us how compete…”
Opinion
Shawn Wang says Devin's breakthrough was Agent-Computer Interfaces, not advanced planning
“The planner is like actually pretty simple, but ACI. That they book through on.”
Opinion
Schulhoff: Hiring Dedicated Prompt Engineers Makes No Sense for Most Companies
“I have always viewed prompt engineering as a skill that everybody should and will have, rather than a specialized role to hire for. That being said, there are definitely times where you do need just a prompt engineer. I think for AI companies, it's definitely …”
Opinion
Howard: Tech builds too many vanity foundation models over fine-tuning
“People are building too many vanity foundation models rather than taking better advantage of fine-tuning”
Opinion
Howard: Building web apps is much worse now than 15 years ago
“Much to my, you know, horror, the story around creating web applications is much worse now than it was 10 or 15 years ago, in terms of, like, if I say to a data scientist, here's how to create and deploy a web application, You know, either you have to learn Ja…”
Opinion
Howard: Cursor and VS Code shoehorn AI into legacy software paradigms
“It's like a convenience over the top of this incredibly complicated system that full-time, sophisticated software engineers have designed over the past few decades in a totally different environment as a way to build software, you know. And so we're trying to,…”
Opinion
Scialom: AI has minted infrastructure unicorns but few successful application companies
“I see like now a lot of fundamental stacks that are like the unicorn of today. Foundational models, foundational like clusters, data notations, things like that. There's a lot, but less successful yet, for now at least, application company. And it's hard to bu…”
Opinion
Yi Tay: Llama 3 shows Meta may have caught up to Google
“So I think I don't really follow, like, fine much, but I think that, like, Lama Tree actually shows that, like, kind of, like, Meta got a pretty, like, a good stack around training these models you know, like, oh, and I've even started to feel like, oh, they a…”
Opinion
Tay: Mixture-of-Experts is fundamentally the right architecture for scaling
“Fundamentally, I just think that MOEs are just, like, the way to go in terms of, like, floppyram ratio, they bring the benefit from the scaling curve, if you do it right, if you, they bring the benefit from the scaling curve, right, and then, Like, that's, lik…”
Opinion
Tay: Hugging Face's Open LLM Leaderboard is a major problem
“The open LM leaderboard is, like, probably, like, the, a big, like, Problem, to be honest.”
Opinion
Albrecht: Training on AWS prevents diagnosing low-level hardware errors
“And if we're just using, you know, AWS or some other cloud provider, These errors are still going to be there, and you're gonna have no way to know and no way to debug this and no way to diagnose what's going wrong.”