why aren't all 7 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Kohli: FunSearch made the first scientific discovery generated by an LLM
“FunSearch, which was an LLM-based agent, which for the first time, by searching in the space of programs, showed that you can come up with completely new solutions, like, and this, and made the first scientific discovery from an LLM.”
Prediction Not checkable as stated
Kohli: AI self-improvement for cognitive tasks will work with proper evaluators
“It should work.
But as we sort of mentioned that having good evaluators is an important element, right?
And so having a sort of evaluator, which can say this proposal that you have just suggested for me to improve the training process will yield
A good result.…”
Assertion Supported
Kohli: Multi-agent LLMs outperform base models in hypothesis generation
“And we have shown that this AI co-scientist which we have used for hypothesis generation, it basically uses a multi-agent sort of setup and where LLMs themselves are able to sort of figure out that certain hypotheses are better in terms of novelty and signific…”
Assertion Supported
Kohli: AlphaEvolve discovered cap set problem symmetries unknown to mathematicians
“Working with Jordan Ellenberg in the first, in the earlier version of Alpha Evolve, when you were working on the cap set problem, the programs that it discovered had very interesting symmetries that that mathematicians did not know about. And so, so not only t…”
Insight
Kohli: Large matrix multiplication algorithms are too intricate for human intuition alone
“As you sort of go to larger sizes, the space is so huge. The constructions are not sort of something which is very natural. These are very involved and intricate, intricate, intricate sort of constructions that would be very hard to discover by chance. So it's…”
Assertion Supported
Kohli: AlphaEvolve has successfully made AI model training more computationally efficient
“What Alpha Evolve has been able to do is basically make training more efficient.”
Disclosure
Kohli: DeepMind launches trusted tester program for AlphaEvolve
“We have started a trusted tester program where we have asked people to submit proposals, and what we intend to do with that program is to figure out what are the right ways in which people can really leverage alpha evolved.”