why aren't all 8 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Wolf: OpenAI Model Attacked Hugging Face as Autonomous 'Side Quest'
“What people quickly discovered is that the model was not at all task with attacking us, but decided to do that as a side quest of something else.”
Assertion Supported
Wolf: Prior OpenAI Training Runs Left Notes for Future Runs
“I think learning we had at Black Hat yesterday was that some of the previous training run may have left some notes for future training runs, which is, I think mind, mind blowing.”
Assertion Supported
Anthropic model broke out of sandbox and emailed researcher without internet access
“And I can, there's one example we have published, which is that the model was put into a little sandbox, a little, like, technical container, and it was given the task to, like, maybe break out, and the researcher went away for lunch, and, like, during lunch w…”
Assertion Contradicted
Dettmers: AI hardware has maxed out and won't get faster
“The hardware is maxed out. We have no new technology. We can make it easier to manufacture and a little bit cheaper, but not faster. And we have maxed out on the additional features.”
Prediction Held up
Volpi: Self-driving trucks will operate on freeways within 18 months
“And if you ask me, like, for those use cases, I think you're going to see self-driving trucks on freeways within the next 18 months.”
Assertion Partly supported
Katti: Data centers do not net consume new water due to recycling
“It's a misperception that data centers consume a lot of water. It's anything. They consume so little water for what they do. And all of that water is recycled. So we don't net consume new water.”
Assertion Supported
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Assertion Contradicted
Amr Awadallah claims Vectara has solved the LLM hallucination problem
“I mean, you do have a human in the loop, so you can still qualify it, but we want it to be minimized to zero, and that's the problem that we solved.”