The Wisdom Wall
16 quotable lessons, heuristics and mental models. Every one is playable at the moment it was said. No fortune cookies allowed.
“one point was, of course, the Transformers when it started, but the other point was reasoning models.”
“Pre-training, as I said, I, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put More compute. You will get better losses if you do things right, which is extremely hard, and that's valuable.”
“using more tokens to think increases your capability, and it increases it, given the computation, way faster than pre-training, right?”
“Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge.”
“with the new paradigm of reasoning, you can get much more gains for the same amount of money because it's on this like lower and like, there are just discoveries to be made and these discoveries unlock insane capabilities.”
“So currently, and current for at least the Most basic ways we use it currently, it needs to be fairly verifiable. So there is an, is your answer correct or not? You prepare data for that. You can do that in mathematics, coding very well. Uh, you can do this in science to some extent, right? You can have test questions,…”
“If you learn to think for math, you can, you will sometimes do some, you know, some strategies are the transfer very much like look up on the web and see what they say and use that information. So some of these things are very generic and they start to transfer.”
“I think in, in general, the, the tech Labs are more similar to each other than people think. There are some differences, but, but I think if, if I look at it from the world, you know, from the university in France, the difference between this university and any of the tech labs is much larger than, than between one lab…”
“pre-training has always worked. And, and the beautiful thing is it even stacks with RL. So if you run this thinking RL process on top of a better model, it works even better. Than if you run it on top of a, of a smaller model.”
“So the understanding of what the models are doing on a higher level has progressed a lot, but then it's still an understanding of what smaller models do, not the biggest ones. But it's not so much that these patterns don't apply to bigger models. They do. It's just the bigger models just do so many things at the same…”
“So it was a bit of a brittle technique, but it was a bit of RL that was extremely crucial to, to making the models chat.”
“in deep learning, people laugh that ideas are cheap. Making them work is, is the hard part.”
“The engineering part is the biggest bottleneck. I mean, GPUs are a bottleneck too, when you scale really up. But implementing something that's larger than one machine, it's an experimental research project, so you don't have a team to do that.”
“This thing, like how do models connect with the external world? It's a fundamentally very hard problem because, you know, when you connect in an unlimited way, you can break things in the real world.”
“robotics is probably just an illustration that we are not doing that well in multimodal and that we're not doing that well in general reasoning yet.”
“with distillation, you have the ability to put a number of projects into one model. It's kind of nice that you don't need to wait on all of them to complete at the same time. You can try to periodically put this together, actually make sure that as a product it's nice to the users and good, and do this separately from,…”