Insight certainty 4/5 debate potential 2/5

Uszkoreit: LLM compute scales with token length, not problem difficulty

Jakob Uszkoreit · No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit · Aug 24, 2023 · at 8:34

Transformer co-author Jakob Uszkoreit discusses foundational bottlenecks in AI architecture with Elad Gil and Sarah Guo.

0:00 / 0:26exact quote · 26.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Then ultimately, the way you scale that compute depends on the prompt and how much, how long that is. The longer the prompt, the more compute you get. And it depends on, and there's of course many different screws to tweak here, the length of the response. There are many very hard problems where the response is incredibly short. And you can, in many cases, actually formulate those problems very, very succinctly. So you're not going to be using a lot of compute, even though the problem we know is really, really difficult.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Jakob Uszkoreit

Opinion
Uszkoreit doubts humans will ever achieve full mechanistic understanding of biology
“I don't have very high hopes for humanity to develop that conceptual understanding to the level that we would need it in order to do all the interventions we want to do.”
Jakob Uszkoreit Aug 24, 2023 ▶ 16:05 No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
Insight
Uszkoreit: Hardware efficiency is the only proven way to advance deep learning
“At the end of the day, in my mind, that's the one and only thing we know really works. If you want to push deep learning forward is to make it faster and more effective and more efficient on a given piece of Hardware.”
Jakob Uszkoreit Aug 24, 2023 ▶ 1:29 No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
Insight
Uszkoreit: Transformer breakthrough was driven by accelerator hardware fit
“And if you want to look at, say, the biggest differences, for example, between the transformer, as it was described in the attention is all you need paper, and some of its ancestors, like this decomposable attention model, the big difference is just that the t…”
Jakob Uszkoreit Aug 24, 2023 ▶ 3:59 No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
Opinion
Uszkoreit: GPUs are not at the sweet spot for large-scale deep learning
“I don't think GPUs are at the sweet spot when it comes to large-scale deep learning with respect to exactly those trade-offs, and so it may very well be that if we actually try these combinations, we might actually even quickly find something that's better.”
Jakob Uszkoreit Aug 24, 2023 ▶ 5:45 No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
Insight
Uszkoreit: Community optimism drove Transformer adoption and success
“The other main contributor, I think, to the success of this architecture was optimism and hope. So suddenly you were in a situation where, for whatever reason, a bunch of things that people tried with this started to work, and then more started to work, and th…”
Jakob Uszkoreit Aug 24, 2023 ▶ 7:00 No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
Insight
Uszkoreit: Training on synthetic data works by amortizing generation compute
“And ironically, and this comes back to a question that many people ask, I think around, does it make any sense to train on generated data? Because information theory, family information theory, very clearly says, nope, you're not going to get more information …”
Jakob Uszkoreit Aug 24, 2023 ▶ 9:24 No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.