why aren't all 5,957 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 100 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Swyx: AI models hit a 2T parameter wall, won't reach 10T
“As far as everyone is concerned, Claude, you know, Opus 3.5 is not coming out. GPT 4.5 is not coming out. And Gemini two, like we don't have pro whatever we've hit that wall, whatever that wall is. Maybe I'll call it like the two trillion parameter wall. Like …”
Opinion
Neubig: Anthropic's MCP Duplicates Existing APIs With Little Added Value
“We already have an API for GitHub. So why do we need an MCP for GitHub, right? You know, like GitHub has an API. The GitHub API is evolving. We can look up the GitHub API documentation. So it seems like kind of duplicated a little bit. And also they have a set…”
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Opinion
Ben Allal: Synthetic data may enrich the web rather than pollute it
“So personally, I wouldn't say the web is posted with synthetic data. Maybe it's even making it more rich.”
Assertion Not checkable as stated
Fu: Embedding model quality barely matters for final RAG performance
“We had this experience over and over again where you could have any, an embedding model of any quality, so you could have a really, really bad embedding model, or you could have a really, really good one by, and by any measure of good, and for the final RAG ap…”
Opinion
Soldani: Open Model Bio-Risk Warnings Were a Lobbying Ploy
“You know, if you remember the beginning of this year, it was all about bio-risk of these open models. The whole thing fizzled out because there's been, finally there's been, like, rigorous research, not just this paper from coherent folks, but there's been rig…”
Opinion
Ramachandran: SWE-bench and HumanEval do not reflect real professional engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have sweet bench. That's cool. No actual Professional work looks like Sweebench. Like, human eval, same thi…”
What-if
Varun Mohan: Codeium would have failed if it used vLLM
“If we use VLLM, we would not be talking with you right now.”
Insight
Anthropic's Schluntz: Avoid agent frameworks and start from scratch with raw prompts
“I think with agent frameworks in general, they can certainly save you some like boilerplate, but I think there's actually this like downside of making agents too easy, where you end up very quickly, like building a much more complex system than you need. And s…”
Prediction Not checkable as stated
Specialized open-source expert models will outperform one-size-fits-all closed-source models
“And that's our prediction is With specialization, there will be a lot of expert models, really, really good, and even better than, like, one size fits all open source closed source model.”
Assertion Not checkable as stated
Crivello: Lindy AI agents perform better than humans for most use cases
“I think the bar is it needs to be better than a human. And for most use cases we serve today, it is better than the human, especially if you put it on rails.”
Opinion
Crivello: OpenAI's GPT-4o Is Overhyped and Poor for Agents
“I think four O is overhyped. Frankly, we don't use four O. I don't think it's good for agentic behavior.”
Assertion Not checkable as stated
Crivello: Every major remote success story lost to an in-person competitor
“In every one of these examples, you have a co-located counterfactual that is sometimes orders of magnitude bigger.”
Insight
Crivello: Remote work optimizes for cost; building creative software requires in-person collaboration
“If you're optimizing for cost, absolutely be remote. If you're optimizing for creativity, which I think that software and product building is a creative endeavor, if you're optimizing for creativity, it's kind of like composing an album. You can't do it on the…”
Opinion
Polu: The bulk of useful enterprise agent work can use APIs
“The bulk of the useful stuff that you can do within the company can be done through API. The data can be retrieved by API, the actions can be taken through API.”
Opinion
Drew Houston: Large language models are a rapidly self-commoditizing, bad business
“Large language models are a pretty bad
Business from a, you know, you sort of take off your tech lens and just sort of business lens. Like there's sort of this weirdly self-commoditizing thing where, you know, models only have value if they're kind of on this…”
Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And
Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Prediction Not checkable as stated
Goyal: Agent control flow and graph routing will move into models
“It feels very clear to me that this type of logic is going to be built into the model. Anytime there is control flow complexity or uncertainty complexity, I think the history of AI has been to push more and more into the model.”
Prediction Not checkable as stated
Goyal: OpenAI o1 will make agentic frameworks obsolete
“And I think O-one is going to do that to agentic frameworks as well. Hey, I think To me, it seems very unlikely that the, you know, you and me sort of like sipping an espresso and thinking about how, like, different personified roles of people should interact …”
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Prediction Not checkable as stated
Shunyu Yao predicts training models on human computer trajectories achieves AGI
“The simplest way to achieve AGI is literally just record the re-actuatory of every human being and just put them together, you know, like what do you have thought about? What do you have done? Let's say on the computer, right? Imagine like solid experiment. Li…”
Opinion
Schulhoff: Role Prompting Does Not Improve Accuracy on Modern LLMs
“For accuracy-based tasks, like MMLU, you're trying to solve a math problem, and maybe you tell the AI that it's a math professor, and you expect it to have improved performance. I really don't think that works. I'm quite certain that doesn't work on more moder…”
Insight
Tay: Frontier AI researchers cannot maintain standard nine-to-five work-life balance
“You cannot be, like, checking out on, like, Friday, Saturday, Sunday, and, like, work at, like, nine to five if you want to, like, Make progress, or like, some people are just so good at detaching, like, ok, like, you know, like, eight pm, I'm not going to, my…”
Opinion
Tay: Long context architecture is the future of AI over RAG
“And, yeah, I mean, I think long context is definitely the future, rather than rec. But I mean, they could be used in conjunction, like,”
Insight
Albrecht: LLM emergence is an artifact of non-linear evaluation metrics
“This emergent behavior that you're seeing, Is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
Insight
Translating code proves LLMs possess causal, functional understanding
“If you ask the LLM to translate a bit of Python into a little bit of C, and it's performing this task, obviously it is understanding in the sense that it has a, The causal, functional model it implements.”
Insight
Bach: Claude is implemented as an invariant pattern similarly to consciousness
“Claude exists only as a pattern. It's something that is a pattern in the activation of the transistors. And even transistors don't actually exist. They are A pattern in the atoms that we are able to see as an invariance because we tune the atoms in a particula…”
Opinion
Big AI labs build AGI internally while giving the public child-proof apps
“In some sense, the big AI companies are incentivized and interested in building AGI internally, and giving everybody else a child-proof application.”
Opinion
Bach: Prompting an LLM correctly can make it sentient to some degree
“It's a Weltgeist that gets possessed by a prompt. And if you possess it with the right prompt, then it can become sentient to some degree.”
Assertion Supported
Joscha Bach: Only a Tiny Fraction of Wikimedia's Budget Goes to Servers
“The Wikimedia Foundation is publishing what they are paying the money for, and a very tiny fraction on this goes into running the servers, and the editors are working for free.”
Insight
Liu: LLM observability startups ignore full systems; just use Postgres
“The issue really is the fact that these observability companies isn't actually doing observability for the system, it's just doing the LLM thing. Like I still end up using like Datadog, right? Or like, you know, Sentry to do, like, latency. And so I just have …”
Insight
Liu: Companies abandon LLM frameworks to regain control over prompts
“So much of it is changing that if you give control of these systems away too early, you end up ultimately wanting them back. Like many companies I know that I reach out or ones were like, oh, we're going off of the frameworks because now that we know what the …”
Opinion
Liu: App-layer AI startups should avoid hiring traditional machine learning engineers
“I think a lot of these app layer startups should not be hiring MLEs because they end up churning.”
Prediction Not checkable as stated
Chintala: George Hotz's TinyGrad requires major breakthroughs to match PyTorch
“There's no, like, I don't think, like, unless we have, like, great breakthroughs, like, George's vision is achievable, like, or, like, he should be thinking about a narrower problem, such as, I'm only gonna make this for, like, work for self-driving car con ne…”
Prediction Not checkable as stated
Chintala: Apple's MLX will fail server-side due to lack of differentiation
“If they end up expanding onto the server side, and they'll probably build something like PyTorch as well, right? Like, eventually, that'll where it will land. And I think there, they will kind of fail on the, like, lack of differentiation. Like, it wouldn't be…”
Opinion
Chintala: Inference Moats from Fast CUDA Kernels Last Only Months
“I think, like, Together and Fireworks and all these people are trying to build some faster CUDA kernels and faster, like, you know, hardware kernels in general. But those modes only last for a month or two. Like, these ideas quickly propagate.”
Insight
AI Productivity Gains Will Ultimately Increase Demand for Software Engineers
“Like, I think we can, you know, 10 x the amount of developers, and still, you know, have a lot of people making a lot of money, you know, building amazing software, and also being, while at the same time being more productive. Like, I never understood this, li…”
Insight
Hsu: Higher valuation for the same quality company is always worse
“Higher valuation, given the same company quality, is always worse.”
Opinion
Yegge: Software engineers ignoring AI coding assistants risk career obsolescence
“If you're one of those engineers, man, you better start like, you know, planning another career. Okay. Because this stuff is in the future and it's honestly, it takes some effort to actually make coding assistants work today, right? You have to, you know, just…”
Opinion
Patel: Frontier AI model training costs are effectively irrelevant
“In my opinion, I think that's a little bit spicy, but yeah, it's like training costs are irrelevant, right? Like GPT-IV, right? Like 20,000 A-one hundreds. That's like, I know it sounds like a lot of money.”
Opinion
Patel: Fine-tuning existing small models for cloud use is useless
“Unless, unless you're fine tuning for on device use, I think fine tuning current existing models, especially the smaller ones is a useless waste of time, right?”
Prediction Open · timeframe Dec 2026
Patel: OpenAI and Microsoft partnership will likely collapse within years
“Yeah, I expect in the next few years that the OpenAI and Microsoft probably falls apart too.”
Opinion
Royzen: AI dev tools do not need to own the IDE
“Somewhere where I disagree with him is that you need to own the IDE. I think like he made kind of some good points about, you know, not having platform risk in the long term, but some of the, you know, features that were mentioned, like suggesting diffs, for e…”
Prediction Not checkable as stated
Royzen: Fine-tuned open source models will beat proprietary in 2024
“So I think that even if a delta exists, in twenty-twenty-four, the delta between proprietary and open source won't be large enough that a startup like us, with a lot of data that we've collected, can take the data that we have, fine-tune an open source model, …”
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Assertion Not publicly verifiable
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Opinion
Howard: Meta 'blew it' on Code Llama due to catastrophic forgetting
“So Code Llama was a, I think it was like a five hundred billion token fine-tuning of Llama II using code. And also prose about code that Meta did. And honestly, they kind of blew it. Because Code Llama is good at coding, but it's bad at everything else.”
Opinion
Howard: TensorFlow 2 was a failure that Google avoided internally
“Then in the end, you know, Google didn't follow through, which is fair enough, like, asking everybody to, you know, learn a new programming language is going to be tough, but, like, it was very obvious, very, very obvious at that time that TensorFlow II was go…”