Aug 24, 2023 · 35m · no-priors
No Priors Ep. 29 | With Inceptive CEO Jakob Uszkoreit
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of No Priors, hosts Elad Gil and Sarah Guo interview Transformer co-author and Inceptive CEO Jakob Uszkoreit about the computational origins of modern AI architectures, future directions in dynamic compute allocation, and how generative deep learning is transforming programmable RNA therapeutics through empirical wet-lab integration.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.6% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jakob interrupts and directly counters Elad's prompt about toddler token exposure by asserting people confuse fine-tuning with evolutionary pre-training.
Hardest push from the hosts ▶ 4:16 Challenging the hardware lottery thesisElad directly questions whether the dominance of the Transformer is merely an artifact of hardware accelerator fit preventing alternative architectures from being tested.
Biggest teaching moment ▶ 8:20 Compute amortization and information theory limitsJakob explains how classical information theory fails to account for energy expenditure and explains why training on generated data amortizes compute over iterations.
The host holds their own ▶ 24:24 Drawing structural parallels to shotgun genomicsElad demonstrates deep technical expertise by drawing a precise parallel between Transformer attention across text chunks and shotgun genomic sequencing assembly.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Origins and Hardware Fit of the Transformer Architecture | 3 | 5 | 2 | 1 | Elad asks about the genesis of the Transformer paper. Jakob gently reframes the invention story, explaining the interplay between linguistic tree representations and accelerator hardware parallelism. | |
| Hardware-Software Co-Design and Accelerator Evolution | 4 | 5 | 2 | 2 | Elad and Sarah question whether current hardware locks in Transformer dominance. Jakob explains why GPUs are not optimal for deep learning trade-offs and attributes Transformer success partly to community momentum. | |
| Rethinking Dynamic Compute Allocation and Model Elasticity | 6 | 6 | 3 | 2 | Sarah brings up specific literature on depth-adaptive transformers and test-time search. Jakob highlights structural flaws in static compute allocation and critiques classical information theory's omission of compute expenditure. | |
| Inceptive and the Vision for Biological Software | 3 | 6 | 1 | 0 | Jakob details Inceptive's vision of biological software, framing deep learning as a scalable empirical tool to compile RNA sequences without needing complete mechanistic biological models. | |
| Empirical Deep Learning vs Mechanistic Discovery in Medicine | 8 | 3 | 1 | 1 | Elad demonstrates substantial domain knowledge by linking Transformer data assembly to shotgun genomic sequencing and noting that historical drugs like aspirin and metformin succeeded without known mechanisms. | |
| Human Brain Architecture, Evolutionary Pre-Training, and AGI | 6 | 7 | 4 | 3 | Jakob firmly reframes Elad's comparison of child token exposure to LLM training by distinguishing evolutionary pre-training from lifetime fine-tuning, while expressing skepticism toward the term AGI. |