Jul 2, 2025 · 1h 18m · latent-space
Information Theory for Language Models: Jack Morris
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Cornell Tech PhD researcher Jack Morris joins swyx to unpack the information-theoretic foundations of language models, exploring vector embedding privacy, empirical parameter storage limits, and how academic researchers can drive high-impact discoveries amidst rapid industry scaling.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jack takes a provocative devil's advocate stance, arguing that web-scale pre-training on sophisticated RNNs could have produced ChatGPT without needing transformers.
Hardest push from the hosts ▶ 1:12:18 Challenging the dataset-only thesis with efficiency multipliersswyx challenges Jack's pure dataset thesis by pointing out that algorithmic and optimizer improvements act as major multipliers equivalent to orders of magnitude more training data.
Biggest teaching moment ▶ 15:25 Defining extractable V-information under computation limitsJack provides a clear theoretical distinction showing why two files with identical Shannon bit-entropy differ vastly in usable, extractable information.
The host holds their own ▶ 12:29 Detailed analysis of Modular Mojo and hardware compilationswyx demonstrates deep technical and industry insight by breaking down Chris Lattner's Mojo compiler approach to CUDA replacement and fast kernel experimentation.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Navigating AI Research Meta and Scaling to 8B Models | 5 | 4 | 2 | 2 | swyx demonstrates solid familiarity with current AI market dynamics and PhD startup valuations. Jack explains the emergence gap between 100M BERT-scale models and 8B parameter models in academic research. | |
| Systems Engineering, GPU Training, and Modern Toolchains | 8 | 2 | 2 | 3 | swyx shares detailed industry advice regarding HPC resources, PyTorch, DeepSpeed, and Modular Mojo. Jack clarifies his public stance on learning CUDA versus higher-level execution frameworks like vLLM and SGLang. | |
| Information Theory Foundations and Usable Information in LLMs | 7 | 6 | 1 | 2 | Jack introduces V-information and the idea of measuring usable information under computational constraints. swyx engages deeply by comparing this to Kolmogorov complexity and Shannon information limits. | |
| Inverting Embeddings and Vector Database Privacy Risks | 6 | 5 | 2 | 2 | Jack details his research on inverting text embeddings and its privacy implications for commercial vector databases. swyx navigates the slides and connects the findings to real-world context leakage and security attacks. | |
| Universal Geometry of Embeddings and the Platonic Hypothesis | 5 | 4 | 2 | 2 | Jack connects embedding geometry to the Platonic Representation Hypothesis and discusses model capacity plateaus. swyx raises questions about context length degradation and mathematical matrix bounds. | |
| Cognitive Core Concept and Emergence of Reasoning | 7 | 3 | 2 | 3 | swyx introduces Karpathy's concept of the cognitive core and contrasts human biological efficiency with dense LLMs. Jack agrees in spirit but notes the technical difficulty of separating reasoning from factual memorization. | |
| Multimodal Adapters and Theoretical Limits of Small Models | 6 | 4 | 2 | 2 | Jack explains how CycleGAN inspired unaligned latent space mapping across distinct model architectures. swyx applies this insight to modular multimodal adapters like Gemma 3n. | |
| Quantifying Model Capacity and Memorization Bounds | 6 | 5 | 2 | 4 | Jack presents his finding that 32-bit parameters only store roughly 3.6 to 3.9 bits of memorized information. swyx questions whether optimizing memorization capacity conflicts with the primary goal of generalization. | |
| Approximating Training Data Directly from Model Weights | 5 | 5 | 1 | 2 | Jack explains his method for approximating proprietary fine-tuning data from the weight delta between base and instruct checkpoints. swyx highlights the novelty of using synthetic checkpoints for data reconstruction. | |
| Kuhnian Paradigm Shifts and Dataset-Driven AI Progress | 6 | 5 | 5 | 4 | Jack presents a contrarian Kuhnian thesis arguing that AI progress is almost entirely dataset-driven rather than architectural. swyx offers pushback, arguing that compute and optimization efficiency act as substantial data multipliers. |