why aren't all 7 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Insight
Roucher: DeepSeek-R1 ranks slightly below OpenAI o1 on smolagents tasks
“I tried R one, but R one is a bit under O one with small agents. And I think this is also a matter of formatting. Like sometimes the model struggles to just output them, the code snippets in the correct way that we expect.”
Assertion Supported
DeepSeek-R1 researchers found MCTS and Process Reward Models were not useful
“R-one specifically said, yes, we tried MCTS. Yes, we tried PRMs. And none of that is useful.”
Assertion Supported
Packer: Sleep-time compute offers Pareto improvements across Claude 3.7 and DeepSeek
“It's like pretty consistent across like both 3.7 deep seek, three mini, which all like the way you actually scale the x-axis here is fundamentally quite different in each case with 3.7 extended thinking mode. The parameter you provide to scale it is different …”
Insight
Lambert: Reasoning models solved basic skills; planning is the next frontier
“So I came up with four and the foundational one was skills, which is What I would say that we have already done with O-one and R-one, which is you do a lot of RL, you show the inference time scaling works and you get really high benchmark numbers. And then the…”
Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Assertion Supported
Lambert: DeepSeek-R1 starts solving math questions immediately without explicit planning
“If you look at DeepSeq R-One and you ask it a hard math question, it's not like, here's my plan of attack. It just starts.”