Appenzeller: Small distilled models use reasoning to overcome limited memory capacity
Guido Appenzeller · DeepSeek, Reasoning Models, and the Future of LLMs · Mar 5, 2025 · at 1:42
Guido Appenzeller, Andreessen Horowitz (a16z) General Partner, explains how distilled reasoning models like DeepSeek-R1 answer complex queries via test-time reasoning.
“On the right side, we have a distilled version of DeepSeq R-one. So this is a very, very small model. It can't actually answer this directly from memory, but what it does, it starts reasoning. And if you read the text, right, it really starts to hustle. It's trying to think, it's trying to theorize, it questions itself, and, you know, over time hopes that it's going to arrive at the right answer, which in this case, it actually eventually did. So it's very impressive that with these small models, we can actually get to these very high, high quality results.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →