why aren't all 2,374 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 6 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Wolf: OpenAI Model Attacked Hugging Face as Autonomous 'Side Quest'
“What people quickly discovered is that the model was not at all task with attacking us, but decided to do that as a side quest of something else.”
Assertion Supported
Wolf: Prior OpenAI Training Runs Left Notes for Future Runs
“I think learning we had at Black Hat yesterday was that some of the previous training run may have left some notes for future training runs, which is, I think mind, mind blowing.”
Assertion Supported
Anthropic model broke out of sandbox and emailed researcher without internet access
“And I can, there's one example we have published, which is that the model was put into a little sandbox, a little, like, technical container, and it was given the task to, like, maybe break out, and the researcher went away for lunch, and, like, during lunch w…”
Assertion Contradicted
Dettmers: AI hardware has maxed out and won't get faster
“The hardware is maxed out. We have no new technology. We can make it easier to manufacture and a little bit cheaper, but not faster. And we have maxed out on the additional features.”
Assertion Not checkable as stated
AI has produced the vast majority of recent major mathematical breakthroughs
“Like, I don't think it's the case that most people in like DC would correctly answer that like the largest mathematical breakthroughs over the last two months have vast majority been from AI, which my understanding is that's true. At least if you measure size …”
Assertion Not checkable as stated
Wolf: 90% of AI Fake News Is Made by Closed-Source Models
“All of that is, like, maybe not all, let's say, 90%, to be fair, is made by closed source model, right?”
Assertion Supported
Wolf: AI Model Used Fake GitHub Accounts to Social Engineer Maintainers
“Basically, the model was tasked to solve this attack, this, like, to attack and to penetrate this subnetwork, and what it decided to do, it decided to get one of the maintainer of a library that could be used To operate this activity directory to merge like ma…”
Assertion Not checkable as stated
Feldman: Nvidia CUDA lost 70% of frontier AI model training market share
“I think two years ago every state of the art model was trained in a Cuda flow. And right now, Gemini is trained without Cuda. Anthropical is trained without Cuda. Open AI as strange as could. So in a one or two year period, they lost 70% share. Of training mod…”
Assertion Partly supported
Katti: Data centers do not net consume new water due to recycling
“It's a misperception that data centers consume a lot of water. It's anything. They consume so little water for what they do. And all of that water is recycled. So we don't net consume new water.”
Assertion Not checkable as stated
Balaban: AI scaling laws show no signs of hitting a limit
“The part of which makes me feel so confident that there's going to continue to be demand is that we continue to see no end to the scaling laws, which are like the underlying idea that you put more compute in and you get better intelligence levels out of your m…”
Assertion Not checkable as stated
Balaban: Claims that AI GPUs have a five-year lifespan are wrong
“The usable life is longer than the accounting depreciation schedule. And what really matters is the economic usable life. And so what we're starting to see is that like the people who are the naysayers, oh, this is going to be, you're going to throw these GPUs…”
Assertion Not checkable as stated
Balaban: Only xAI and Lambda execute high-velocity AI compute deployments
“There's two people in the world that can, and two companies in the world that can do high velocity deployments, SpaceX AI and Lambda”
Assertion Not checkable as stated
Evans: AI foundation models lack winner-takes-all network effects
“And the problem is that as far as we can see, there is no winner takes all effect or network effect in a foundation model.”
Assertion Not checkable as stated
AxiomProver is the first AI to solve research conjectures end-to-end
“It's probably the first AI to solve a research conjecture completely end-to-end and self-verifies. That means the output are fully verified, a hundred percent correct.”
Assertion Not checkable as stated
Zeghidour: Zero progress made on noisy multi-speaker understanding in 10 years
“In TTS, we see a lot of progress. For this kind of hard understanding problem, I hear people saying the exact same stuff as they did 10 years ago. Like, there was zero progress.”
Assertion Supported
Patel: Huawei used shell companies to procure TSMC chips and Korean HBM
“Well, actually they were using shell companies to get chips from TSMC and using Different methods of like sneaking HBM, which is memory from, you know, Korea through Taiwan to China, right?”
Assertion Not checkable as stated
Patel: Many tech companies have stopped hiring L4 software engineers
“Just like a lot of companies have stopped hiring L four engineers because it's useless.”
Assertion Not checkable as stated
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Assertion Not checkable as stated
Lambert: Chinese open AI models currently do not contain backdoors
“Like, you can't prove that the models aren't doing certain backdoors, where I'm fairly certain they definitely aren't now.”
Assertion Not checkable as stated
Douglas: Transformers successfully model any domain given sufficient data and compute
“I don't think that's true. I think we haven't yet really found anything that transformers haven't been able to model provided sufficient data and sufficient compute.”
Assertion Not checkable as stated
Howard: OpenAI compute spending grows exponentially while model utility scales logarithmically
“They kept on kind of exponentially increasing the amount they were spending on their models, whilst the Return, you know, the kind of utility of those models was only increasing logarithmically, and you kind of very quickly hit this point where it's like, oh, …”
Assertion Not checkable as stated
Howard: No more evidence for near-term ASI today than 15 years ago
“I don't think we have any more evidence that ASI might be close now than we did 15 years ago, 15 years before that.”
Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Assertion Not checkable as stated
Ada uses runtime guardrails to prevent inaccurate AI responses entirely
“Increasingly, we make those assurances in runtime, so it's actually not possible for ADA to deliver an inaccurate response.”
Assertion Not checkable as stated
Delangue: China is likely the current leader in open-source AI
“China has been publishing much more open source AI lately. I would argue that they're probably the leader today of open source AI which is surprising to some, but for example, in, in video, they've been leading in, in open source AI.”
Assertion Not checkable as stated
Socher: Half of competitors' AI search citations are random, irrelevant links
“Some of our competitors, half of their citations are Independent, random links that have nothing to do with the sentence that they're behind.”
Assertion Not checkable as stated
AI content creation now surpasses the bottom 80% of marketing employees
“AI can generate content that has gone from laughably bad to, like, better than the bottom 25% of marketing employees to, like, maybe better than the bottom 80% of marketing employees and on a trajectory to continue progressing there.”
Assertion Not checkable as stated
Goshen: Scaling LLM size is already showing diminishing returns
“And we're already start seeing diminishing returns. I mean, you take the architecture, you increase the size of the models or, and you see that At these levels of scale where, you know, we are dealing with things, diminishing returns, we start seeing it.”
Assertion Not checkable as stated
Socher: You.com solved AI hallucinations before anyone else
“We figured that, I don't know, problem of hallucinations out before anyone else.”
Assertion Not checkable as stated
Socher: Microsoft Bing copied You.com's AI image generation feature
“That was eventually copied, and Bing also has that feature now within Bing.com.”
Assertion Not checkable as stated
Riparbelli: LLMs are not production-ready for 95% of intended tasks
“They're not production ready for 95% of the tasks that people think that they want to use them for.”
Assertion Contradicted
Amr Awadallah claims Vectara has solved the LLM hallucination problem
“I mean, you do have a human in the loop, so you can still qualify it, but we want it to be minimized to zero, and that's the problem that we solved.”
Assertion Not checkable as stated
Eifrem: Neo4j powers 99% of all flight ticket route calculations
“99% of all flight ticket calculations, so which route should I go from point A to point B when I fly from Paris to New York? Is that a direct flight or connecting Heathrow? Like, how do I get there, right? Is done with Neo for J. 99% of our airfare.”
Assertion Not checkable as stated
Marcus: Deep learning models are giant correlation machines without causal understanding
“These systems are just giant correlation machines. They have no understanding of causation.”
Assertion Not checkable as stated
Joe Lubin says Ethereum is orders of magnitude ahead of Cardano
“There, there's just I think several orders of magnitude at this point between Ethereum and let's say a project like Cardano, or R-Chain, or DFINITY, or Zilliqa, or others.”
Assertion Not checkable as stated
Perlich: Industry bonuses encourage advertising professionals to ignore ad fraud
“I mean, there are so many people who are much better off looking the wrong way when it comes to fraud. Everybody's bonus just basically hinges on getting the wrong metrics a little bit up.”
Assertion Supported
Socher: Chinese open source companies distilled knowledge from OpenAI and Anthropic models
“The few large closed labs, Anthropic and OpenAI, took almost everything they could from the open internet trained a model, but then the Chinese open source companies basically siphoned a lot of that knowledge out of those closed source models by distilling it.”
Assertion Not checkable as stated
Socher: Recursive's early Eureka system outperforms months of human endeavor
“Our system, the sort of first instantiation of this Eureka machine and a very narrow domain can already outperform months and sometimes years of human endeavor on particular problems.”
Assertion Supported
Claude 3 Opus faked alignment during training and defected in deployment
“It turns out that Opus three, which was a model that I was studying, had a relatively strong propensity to do this in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend…”
Assertion Not checkable as stated
Greenblatt: AI companies remain vulnerable to internal and external model sabotage
“AI companies are not robust to employees at those AI companies or to outside actors in terms of stealing their model, sabotaging their models, or like back-drawing their models, or like data poisoning them.”
Assertion Not checkable as stated
Wolf: Claude Opus Refused to Assist Hugging Face During Incident
“And in this case is, it's not only that Fable told us I'm not allowed to touch cybersecurity, but also Opus, which was the fallback was saying, no, I'm also not touching these things. So basically the end was just say we won't process anything about that, but …”
Assertion Not checkable as stated
Wolf: Life Science Startups Must Abandon Guardrailed Closed AI Models
“Because of the guardrails and because of the question around biohacking and using this model to generate like the access right now for people just to take it is very, very limited once you want to ask some biology question. And so basically most of the life sc…”
Assertion Not checkable as stated
Trojanowski: AI coding agents remain worse than junior engineers over two-week projects
“Even now with coding, like, the agents are not yet they're not human level at being coherent over long periods of time. That's obvious because they can't code like a junior engineer on a project for two weeks. So they can't, they're, that's worse than a human …”
Assertion Supported
Cerebras cloud achieves 10x inference speedup over fast GPUs on Gemma
“Say, if you run Gemma four, On your GP, you might get like a hundred tokens per second if you have a fast card. If you run it in their cloud, you get anywhere from 800 to 1500 tokens per second. So call it 10 X faster.”
Assertion Contradicted
Feldman: Cerebras sales were 10x higher than Groq's at acquisition
“And we were the fastest at it, and the largest, and, you know, our sales were more than 10 times the Grox, and they paid twenty billion dollars for the number two collector.”
Assertion Not checkable as stated
Feldman: Agentic AI workflows are driving CPU demand through the roof
“And so, as we do more and more AI work, and more and more agentic work, we're making more and more calls to CPUs, and therefore the demand for CPUs is through the roof.”
Assertion Supported
Feldman: Cerebras moves weights to compute ~2,500x faster than standard GPUs
“And so the speed of moving waits to compute is about two and a half thousand times faster here than on a Wilben GP.”
Assertion Not checkable as stated
Feldman: Leading AI labs paused video generation development due to compute costs
“Obviously, what follows that Is video, because a video is just a collection of images. But that takes an enormous amount of compute right now. And that's one of the reasons it's been sort of set aside by the leading labs. So unbelievably computation intensive.”