Aug 24, 2023 · 35m · no-priors

No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit

Jakob Uszkoreit · 25m spoken Elad Gil · 4m spoken Sarah Guo · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, hosts Elad Gil and Sarah Guo interview Transformer co-author and Inceptive CEO Jakob Uszkoreit about the computational origins of modern AI architectures, future directions in dynamic compute allocation, and how generative deep learning is transforming programmable RNA therapeutics through empirical wet-lab integration.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.6% of the talking time here. How this is scored →

The hosts as informed peer 5.0 Guest teaching 5.3 Guest disagreement 2.2 The hosts pushing back 1.5
05100:0010:0020:0030:000:05–4:16 · The hosts as informed peer 3/10 Origins and Hardware Fit of the Transformer Architecture Elad asks about the genesis of the Transformer paper. Jakob gently reframes the invention story, explaining the interplay between linguistic tree representations and accelerator hardware parallelism.4:16–7:57 · The hosts as informed peer 4/10 Hardware-Software Co-Design and Accelerator Evolution Elad and Sarah question whether current hardware locks in Transformer dominance. Jakob explains why GPUs are not optimal for deep learning trade-offs and attributes Transformer success partly to community momentum.7:57–15:12 · The hosts as informed peer 6/10 Rethinking Dynamic Compute Allocation and Model Elasticity Sarah brings up specific literature on depth-adaptive transformers and test-time search. Jakob highlights structural flaws in static compute allocation and critiques classical information theory's omission of compute expenditure.15:12–22:36 · The hosts as informed peer 3/10 Inceptive and the Vision for Biological Software Jakob details Inceptive's vision of biological software, framing deep learning as a scalable empirical tool to compile RNA sequences without needing complete mechanistic biological models.22:36–28:27 · The hosts as informed peer 8/10 Empirical Deep Learning vs Mechanistic Discovery in Medicine Elad demonstrates substantial domain knowledge by linking Transformer data assembly to shotgun genomic sequencing and noting that historical drugs like aspirin and metformin succeeded without known mechanisms.28:28–31:56 · The hosts as informed peer 6/10 Human Brain Architecture, Evolutionary Pre-Training, and AGI Jakob firmly reframes Elad's comparison of child token exposure to LLM training by distinguishing evolutionary pre-training from lifetime fine-tuning, while expressing skepticism toward the term AGI.0:05–4:16 · Guest teaching 5/10 Origins and Hardware Fit of the Transformer Architecture Elad asks about the genesis of the Transformer paper. Jakob gently reframes the invention story, explaining the interplay between linguistic tree representations and accelerator hardware parallelism.4:16–7:57 · Guest teaching 5/10 Hardware-Software Co-Design and Accelerator Evolution Elad and Sarah question whether current hardware locks in Transformer dominance. Jakob explains why GPUs are not optimal for deep learning trade-offs and attributes Transformer success partly to community momentum.7:57–15:12 · Guest teaching 6/10 Rethinking Dynamic Compute Allocation and Model Elasticity Sarah brings up specific literature on depth-adaptive transformers and test-time search. Jakob highlights structural flaws in static compute allocation and critiques classical information theory's omission of compute expenditure.15:12–22:36 · Guest teaching 6/10 Inceptive and the Vision for Biological Software Jakob details Inceptive's vision of biological software, framing deep learning as a scalable empirical tool to compile RNA sequences without needing complete mechanistic biological models.22:36–28:27 · Guest teaching 3/10 Empirical Deep Learning vs Mechanistic Discovery in Medicine Elad demonstrates substantial domain knowledge by linking Transformer data assembly to shotgun genomic sequencing and noting that historical drugs like aspirin and metformin succeeded without known mechanisms.28:28–31:56 · Guest teaching 7/10 Human Brain Architecture, Evolutionary Pre-Training, and AGI Jakob firmly reframes Elad's comparison of child token exposure to LLM training by distinguishing evolutionary pre-training from lifetime fine-tuning, while expressing skepticism toward the term AGI.0:05–4:16 · Guest disagreement 2/10 Origins and Hardware Fit of the Transformer Architecture Elad asks about the genesis of the Transformer paper. Jakob gently reframes the invention story, explaining the interplay between linguistic tree representations and accelerator hardware parallelism.4:16–7:57 · Guest disagreement 2/10 Hardware-Software Co-Design and Accelerator Evolution Elad and Sarah question whether current hardware locks in Transformer dominance. Jakob explains why GPUs are not optimal for deep learning trade-offs and attributes Transformer success partly to community momentum.7:57–15:12 · Guest disagreement 3/10 Rethinking Dynamic Compute Allocation and Model Elasticity Sarah brings up specific literature on depth-adaptive transformers and test-time search. Jakob highlights structural flaws in static compute allocation and critiques classical information theory's omission of compute expenditure.15:12–22:36 · Guest disagreement 1/10 Inceptive and the Vision for Biological Software Jakob details Inceptive's vision of biological software, framing deep learning as a scalable empirical tool to compile RNA sequences without needing complete mechanistic biological models.22:36–28:27 · Guest disagreement 1/10 Empirical Deep Learning vs Mechanistic Discovery in Medicine Elad demonstrates substantial domain knowledge by linking Transformer data assembly to shotgun genomic sequencing and noting that historical drugs like aspirin and metformin succeeded without known mechanisms.28:28–31:56 · Guest disagreement 4/10 Human Brain Architecture, Evolutionary Pre-Training, and AGI Jakob firmly reframes Elad's comparison of child token exposure to LLM training by distinguishing evolutionary pre-training from lifetime fine-tuning, while expressing skepticism toward the term AGI.0:05–4:16 · The hosts pushing back 1/10 Origins and Hardware Fit of the Transformer Architecture Elad asks about the genesis of the Transformer paper. Jakob gently reframes the invention story, explaining the interplay between linguistic tree representations and accelerator hardware parallelism.4:16–7:57 · The hosts pushing back 2/10 Hardware-Software Co-Design and Accelerator Evolution Elad and Sarah question whether current hardware locks in Transformer dominance. Jakob explains why GPUs are not optimal for deep learning trade-offs and attributes Transformer success partly to community momentum.7:57–15:12 · The hosts pushing back 2/10 Rethinking Dynamic Compute Allocation and Model Elasticity Sarah brings up specific literature on depth-adaptive transformers and test-time search. Jakob highlights structural flaws in static compute allocation and critiques classical information theory's omission of compute expenditure.15:12–22:36 · The hosts pushing back 0/10 Inceptive and the Vision for Biological Software Jakob details Inceptive's vision of biological software, framing deep learning as a scalable empirical tool to compile RNA sequences without needing complete mechanistic biological models.22:36–28:27 · The hosts pushing back 1/10 Empirical Deep Learning vs Mechanistic Discovery in Medicine Elad demonstrates substantial domain knowledge by linking Transformer data assembly to shotgun genomic sequencing and noting that historical drugs like aspirin and metformin succeeded without known mechanisms.28:28–31:56 · The hosts pushing back 3/10 Human Brain Architecture, Evolutionary Pre-Training, and AGI Jakob firmly reframes Elad's comparison of child token exposure to LLM training by distinguishing evolutionary pre-training from lifetime fine-tuning, while expressing skepticism toward the term AGI.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 35.1% · guest 64.9%0:00 · the hosts 35.1% · guest 64.9%3:00 · the hosts 16.2% · guest 83.8%3:00 · the hosts 16.2% · guest 83.8%6:00 · the hosts 20.3% · guest 79.7%6:00 · the hosts 20.3% · guest 79.7%9:00 · the hosts 3.3% · guest 96.7%9:00 · the hosts 3.3% · guest 96.7%12:00 · the hosts 19.1% · guest 80.9%12:00 · the hosts 19.1% · guest 80.9%15:00 · the hosts 13.1% · guest 86.9%15:00 · the hosts 13.1% · guest 86.9%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 19.3% · guest 80.7%21:00 · the hosts 19.3% · guest 80.7%24:00 · the hosts 48.8% · guest 51.2%24:00 · the hosts 48.8% · guest 51.2%27:00 · the hosts 34.4% · guest 65.6%27:00 · the hosts 34.4% · guest 65.6%30:00 · the hosts 35.8% · guest 64.2%30:00 · the hosts 35.8% · guest 64.2%33:00 · the hosts 13.7% · guest 86.3%33:00 · the hosts 13.7% · guest 86.3%
Sharpest disagreement ▶ 29:40 Reframing LLM vs toddler token efficiency

Jakob interrupts and directly counters Elad's prompt about toddler token exposure by asserting people confuse fine-tuning with evolutionary pre-training.

Hardest push from the hosts ▶ 4:16 Challenging the hardware lottery thesis

Elad directly questions whether the dominance of the Transformer is merely an artifact of hardware accelerator fit preventing alternative architectures from being tested.

Biggest teaching moment ▶ 8:20 Compute amortization and information theory limits

Jakob explains how classical information theory fails to account for energy expenditure and explains why training on generated data amortizes compute over iterations.

The host holds their own ▶ 24:24 Drawing structural parallels to shotgun genomics

Elad demonstrates deep technical expertise by drawing a precise parallel between Transformer attention across text chunks and shotgun genomic sequencing assembly.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins and Hardware Fit of the Transformer Architecture 3521 Elad asks about the genesis of the Transformer paper. Jakob gently reframes the invention story, explaining the interplay between linguistic tree representations and accelerator hardware parallelism.
Hardware-Software Co-Design and Accelerator Evolution 4522 Elad and Sarah question whether current hardware locks in Transformer dominance. Jakob explains why GPUs are not optimal for deep learning trade-offs and attributes Transformer success partly to community momentum.
Rethinking Dynamic Compute Allocation and Model Elasticity 6632 Sarah brings up specific literature on depth-adaptive transformers and test-time search. Jakob highlights structural flaws in static compute allocation and critiques classical information theory's omission of compute expenditure.
Inceptive and the Vision for Biological Software 3610 Jakob details Inceptive's vision of biological software, framing deep learning as a scalable empirical tool to compile RNA sequences without needing complete mechanistic biological models.
Empirical Deep Learning vs Mechanistic Discovery in Medicine 8311 Elad demonstrates substantial domain knowledge by linking Transformer data assembly to shotgun genomic sequencing and noting that historical drugs like aspirin and metformin succeeded without known mechanisms.
Human Brain Architecture, Evolutionary Pre-Training, and AGI 6743 Jakob firmly reframes Elad's comparison of child token exposure to LLM training by distinguishing evolutionary pre-training from lifetime fine-tuning, while expressing skepticism toward the term AGI.

Statements from this episode (17)

Insight
Uszkoreit: Hardware efficiency is the only proven way to advance deep learning
“At the end of the day, in my mind, that's the one and only thing we know really works. If you want to push deep learning forward is to make it faster and more effective and more efficient on a given piece of Hardware.”
Jakob Uszkoreit Aug 24, 2023 ▶ 1:29
Insight
Uszkoreit: Transformer breakthrough was driven by accelerator hardware fit
“And if you want to look at, say, the biggest differences, for example, between the transformer, as it was described in the attention is all you need paper, and some of its ancestors, like this decomposable attention model, the big difference is just that the t…”
Jakob Uszkoreit Aug 24, 2023 ▶ 3:59
Opinion
Uszkoreit: GPUs are not at the sweet spot for large-scale deep learning
“I don't think GPUs are at the sweet spot when it comes to large-scale deep learning with respect to exactly those trade-offs, and so it may very well be that if we actually try these combinations, we might actually even quickly find something that's better.”
Jakob Uszkoreit Aug 24, 2023 ▶ 5:45
Insight
Uszkoreit: Community optimism drove Transformer adoption and success
“The other main contributor, I think, to the success of this architecture was optimism and hope. So suddenly you were in a situation where, for whatever reason, a bunch of things that people tried with this started to work, and then more started to work, and th…”
Jakob Uszkoreit Aug 24, 2023 ▶ 7:00
Insight
Uszkoreit: LLM compute scales with token length, not problem difficulty
“Then ultimately, the way you scale that compute depends on the prompt and how much, how long that is. The longer the prompt, the more compute you get. And it depends on, and there's of course many different screws to tweak here, the length of the response. The…”
Jakob Uszkoreit Aug 24, 2023 ▶ 8:34
Insight
Uszkoreit: Training on synthetic data works by amortizing generation compute
“And ironically, and this comes back to a question that many people ask, I think around, does it make any sense to train on generated data? Because information theory, family information theory, very clearly says, nope, you're not going to get more information …”
Jakob Uszkoreit Aug 24, 2023 ▶ 9:24
Opinion
Uszkoreit: Test-time search is effective but clunky and hard to optimize
“I think it's super effective in test time search. I do think it's clunky because it's not something that you can easily end to end optimize. So, right, basically this is also what I'm, what I was trying to get at a little bit maybe with saying, well, some of t…”
Jakob Uszkoreit Aug 24, 2023 ▶ 13:45
Opinion
Uszkoreit: Universal Transformers have not caught on because they underperform
“In terms of adaptive, adaptive time transformers, et cetera, we tried this universal transformer thing actually a long time ago. It just hasn't caught on. And that's because it just doesn't work right at this point. It doesn't work well enough. It's not like i…”
Jakob Uszkoreit Aug 24, 2023 ▶ 14:34
Opinion
Uszkoreit doubts humans will ever achieve full mechanistic understanding of biology
“I don't have very high hopes for humanity to develop that conceptual understanding to the level that we would need it in order to do all the interventions we want to do.”
Jakob Uszkoreit Aug 24, 2023 ▶ 16:05
Prediction Not checkable as stated
Uszkoreit: RNA could become a top-three drug modality by 2030
“And now we're talking about a modality that might end up before the end of the decade being the second or third biggest modality in terms of revenue and potentially also in terms of impact.”
Jakob Uszkoreit Aug 24, 2023 ▶ 19:22
Prediction Held up
Uszkoreit: Global RNA manufacturing will reach 6-8B doses within two years
“But right now, if you look at RNA, manufacturing and distribution infrastructure, we're going to have six to eight billion doses two years from now, manufacturable and distributable across the globe.”
Jakob Uszkoreit Aug 24, 2023 ▶ 21:57
Opinion
Uszkoreit: Screening approaches cannot handle multi-antigen personalized cancer vaccines
“Because screening is just not going to cut it 10 to the second and 30th, and that's really just one antigen that we're coding for there when we actually want to code for many and update those for any given. [1415] Sarah Guo: For any given, yep. [1416] Jakob Us…”
Jakob Uszkoreit Aug 24, 2023 ▶ 23:25
Opinion
Gil: Regulatory demands for proven mechanisms do not improve drug efficacy
“A lot of the emphasis right now from a regulatory pathway for drugs is, oh, you need a mechanism of function, or you need a proven pathway, and all these things that create hurdles that don't necessarily help with drug efficacy.”
Elad Gil Aug 24, 2023 ▶ 25:55
Insight
Uszkoreit: Interpretable theories for complex deep learning systems exceed human cognitive limits
“There are people trying, and I think it's worth trying. I, I'm not super optimistic about that. I think it'll work for some cases, right, where it's simple enough that we can get it. I think there are many cases where it just isn't, right? Like, say, climate a…”
Jakob Uszkoreit Aug 24, 2023 ▶ 27:15
Opinion
Uszkoreit: Brains lack excess compute for high-bandwidth human augmentation
“I'm very bullish on human augmentation in the very long term, but it's one that I don't see intuitively. I think looking at our brains, even just physically, they seem to be very focused, and this is not surprising, on our IO. And why would there somewhere in …”
Jakob Uszkoreit Aug 24, 2023 ▶ 28:36
Insight
Uszkoreit: Human learning is fine-tuning, while evolution is pre-training
“But I think that's because we confuse fine tuning and pre-training. Pre-training is all of evolution. And then basically you arrive at this thing that It's maybe doing something that's completely, in a certain sense, a completely irrelevant task at first, but …”
Jakob Uszkoreit Aug 24, 2023 ▶ 29:47
Insight
Uszkoreit: Simple single wet-to-dry iterative biological loops are a pipe dream
“There was always this dream of, and I think it's a pipe dream, of having the cycle between experimentation, and then you put that into some, something in silico, something running on computers, and then that informs the experiments, and then you kind of Iterat…”
Jakob Uszkoreit Aug 24, 2023 ▶ 32:54
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.