why aren't all 9 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
White: Verifiers in the loop provide higher signal than hypothesis rankings
“I have a lot more faith in these, like verifier-in-the-loop kind of scenarios where you have either data analysis, literature search, or you're running a unit test, or whatever, you're going and running the experiment. Anything like that, I think, is going to …”
Disclosure
White: FutureHouse bypassed custom foundation models to focus on scientific agents
“For example, at future house, we took the opinion that scientific agents are the future, and that allows us to skip a lot of steps because a lot of other people were like, we need to build a foundation model for X. Yeah. And we just skipped all that.”
Insight
White: RLHF fails on scientific hypotheses by ignoring impact and information gain
“We learned a lot about how bad our LHF is with people, just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what…”
Assertion Partly supported
White: AI hits 60-70% on Bixbench, matching human expert agreement
“We're getting to 60%, 70% correctness on Bixbench, and we found that actually we're at the point where humans disagree at this level. Like, humans only agree 70% of the analysis, and so it's true that, like, when it comes to analyzing data, like, humans do not…”
Assertion Supported
White: Co-Scientist uses LLM ranking while Robin filters with data and literature
“In co-scientists, their filtration process was other LLMs sort of ranking it with rubrics or like personas. And our filtration process was like literature search and data analysis.”
Disclosure
White: Future House Focuses on Discovery Loops Over System Modeling
“We're trying to automate the like cognitive process of scientific discovery, making hypotheses. Choosing experiments to do, analyzing the results from experiments and using it to update your hypotheses or your confidence in those hypotheses, and then leading t…”
Assertion Supported
White: Human expert hypothesis rankings failed to predict wet-lab success
“And one of the things that came out of the Robin paper is that the hypothesis that people thought was best was not the one that led to success in that paper.”
Assertion Supported
White: PaperQA outputs page citations for every generated sentence
“Paper QA was has, like, every sentence that it outputs has a citation to a page, right?”
Disclosure
White: Venture-Backed Startup Edison Was Spun Out of Future House
“Now we have a venture-backed startup called Edison, which we spun out of Future House and, you know, we took a lot of the ideas and we're trying to do this at an even bigger scale right now.”