Howard: Accurate next-token prediction forces models to learn world models and causality
Jeremy Howard · The End of Finetuning — with Jeremy Howard of Fast.ai · Oct 20, 2023 · at 12:37
Fast.ai co-founder Jeremy Howard discusses the foundational insight behind scaling language models and his pioneering 2018 work on ULMFiT.
“I thought, okay, so if I do this at a much bigger scale, using all of Wikipedia, what would it need to be able to do to finish a sentence in Wikipedia effectively, to do it quite accurately, quite often? I thought, geez, it would actually have to know a lot about the world. You know, it would have to know that there is a world, and that there are objects, and that objects relate to each other through time, and cause each other to react in ways, and that causes, precede effects, and that, you know, when there are animals, and there are people, and that people can be in certain positions during certain time frames, and then you could, you know, all that together, you can then finish a sentence”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →