Jul 2, 2025 · 1h 18m · latent-space

Information Theory for Language Models: Jack Morris

Jack Morris · 50m spoken Shawn Wang · 20m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Cornell Tech PhD researcher Jack Morris joins swyx to unpack the information-theoretic foundations of language models, exploring vector embedding privacy, empirical parameter storage limits, and how academic researchers can drive high-impact discoveries amidst rapid industry scaling.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 6.1 Guest teaching 4.3 Guest disagreement 2.1 The hosts pushing back 2.6
05100:0020:0040:001:00:005:12–8:58 · The hosts as informed peer 5/10 Navigating AI Research Meta and Scaling to 8B Models swyx demonstrates solid familiarity with current AI market dynamics and PhD startup valuations. Jack explains the emergence gap between 100M BERT-scale models and 8B parameter models in academic research.8:58–14:31 · The hosts as informed peer 8/10 Systems Engineering, GPU Training, and Modern Toolchains swyx shares detailed industry advice regarding HPC resources, PyTorch, DeepSpeed, and Modular Mojo. Jack clarifies his public stance on learning CUDA versus higher-level execution frameworks like vLLM and SGLang.14:31–22:29 · The hosts as informed peer 7/10 Information Theory Foundations and Usable Information in LLMs Jack introduces V-information and the idea of measuring usable information under computational constraints. swyx engages deeply by comparing this to Kolmogorov complexity and Shannon information limits.22:30–31:17 · The hosts as informed peer 6/10 Inverting Embeddings and Vector Database Privacy Risks Jack details his research on inverting text embeddings and its privacy implications for commercial vector databases. swyx navigates the slides and connects the findings to real-world context leakage and security attacks.31:17–40:09 · The hosts as informed peer 5/10 Universal Geometry of Embeddings and the Platonic Hypothesis Jack connects embedding geometry to the Platonic Representation Hypothesis and discusses model capacity plateaus. swyx raises questions about context length degradation and mathematical matrix bounds.40:09–44:23 · The hosts as informed peer 7/10 Cognitive Core Concept and Emergence of Reasoning swyx introduces Karpathy's concept of the cognitive core and contrasts human biological efficiency with dense LLMs. Jack agrees in spirit but notes the technical difficulty of separating reasoning from factual memorization.44:23–52:38 · The hosts as informed peer 6/10 Multimodal Adapters and Theoretical Limits of Small Models Jack explains how CycleGAN inspired unaligned latent space mapping across distinct model architectures. swyx applies this insight to modular multimodal adapters like Gemma 3n.52:39–1:00:03 · The hosts as informed peer 6/10 Quantifying Model Capacity and Memorization Bounds Jack presents his finding that 32-bit parameters only store roughly 3.6 to 3.9 bits of memorized information. swyx questions whether optimizing memorization capacity conflicts with the primary goal of generalization.1:00:03–1:06:03 · The hosts as informed peer 5/10 Approximating Training Data Directly from Model Weights Jack explains his method for approximating proprietary fine-tuning data from the weight delta between base and instruct checkpoints. swyx highlights the novelty of using synthetic checkpoints for data reconstruction.1:06:03–1:14:39 · The hosts as informed peer 6/10 Kuhnian Paradigm Shifts and Dataset-Driven AI Progress Jack presents a contrarian Kuhnian thesis arguing that AI progress is almost entirely dataset-driven rather than architectural. swyx offers pushback, arguing that compute and optimization efficiency act as substantial data multipliers.5:12–8:58 · Guest teaching 4/10 Navigating AI Research Meta and Scaling to 8B Models swyx demonstrates solid familiarity with current AI market dynamics and PhD startup valuations. Jack explains the emergence gap between 100M BERT-scale models and 8B parameter models in academic research.8:58–14:31 · Guest teaching 2/10 Systems Engineering, GPU Training, and Modern Toolchains swyx shares detailed industry advice regarding HPC resources, PyTorch, DeepSpeed, and Modular Mojo. Jack clarifies his public stance on learning CUDA versus higher-level execution frameworks like vLLM and SGLang.14:31–22:29 · Guest teaching 6/10 Information Theory Foundations and Usable Information in LLMs Jack introduces V-information and the idea of measuring usable information under computational constraints. swyx engages deeply by comparing this to Kolmogorov complexity and Shannon information limits.22:30–31:17 · Guest teaching 5/10 Inverting Embeddings and Vector Database Privacy Risks Jack details his research on inverting text embeddings and its privacy implications for commercial vector databases. swyx navigates the slides and connects the findings to real-world context leakage and security attacks.31:17–40:09 · Guest teaching 4/10 Universal Geometry of Embeddings and the Platonic Hypothesis Jack connects embedding geometry to the Platonic Representation Hypothesis and discusses model capacity plateaus. swyx raises questions about context length degradation and mathematical matrix bounds.40:09–44:23 · Guest teaching 3/10 Cognitive Core Concept and Emergence of Reasoning swyx introduces Karpathy's concept of the cognitive core and contrasts human biological efficiency with dense LLMs. Jack agrees in spirit but notes the technical difficulty of separating reasoning from factual memorization.44:23–52:38 · Guest teaching 4/10 Multimodal Adapters and Theoretical Limits of Small Models Jack explains how CycleGAN inspired unaligned latent space mapping across distinct model architectures. swyx applies this insight to modular multimodal adapters like Gemma 3n.52:39–1:00:03 · Guest teaching 5/10 Quantifying Model Capacity and Memorization Bounds Jack presents his finding that 32-bit parameters only store roughly 3.6 to 3.9 bits of memorized information. swyx questions whether optimizing memorization capacity conflicts with the primary goal of generalization.1:00:03–1:06:03 · Guest teaching 5/10 Approximating Training Data Directly from Model Weights Jack explains his method for approximating proprietary fine-tuning data from the weight delta between base and instruct checkpoints. swyx highlights the novelty of using synthetic checkpoints for data reconstruction.1:06:03–1:14:39 · Guest teaching 5/10 Kuhnian Paradigm Shifts and Dataset-Driven AI Progress Jack presents a contrarian Kuhnian thesis arguing that AI progress is almost entirely dataset-driven rather than architectural. swyx offers pushback, arguing that compute and optimization efficiency act as substantial data multipliers.5:12–8:58 · Guest disagreement 2/10 Navigating AI Research Meta and Scaling to 8B Models swyx demonstrates solid familiarity with current AI market dynamics and PhD startup valuations. Jack explains the emergence gap between 100M BERT-scale models and 8B parameter models in academic research.8:58–14:31 · Guest disagreement 2/10 Systems Engineering, GPU Training, and Modern Toolchains swyx shares detailed industry advice regarding HPC resources, PyTorch, DeepSpeed, and Modular Mojo. Jack clarifies his public stance on learning CUDA versus higher-level execution frameworks like vLLM and SGLang.14:31–22:29 · Guest disagreement 1/10 Information Theory Foundations and Usable Information in LLMs Jack introduces V-information and the idea of measuring usable information under computational constraints. swyx engages deeply by comparing this to Kolmogorov complexity and Shannon information limits.22:30–31:17 · Guest disagreement 2/10 Inverting Embeddings and Vector Database Privacy Risks Jack details his research on inverting text embeddings and its privacy implications for commercial vector databases. swyx navigates the slides and connects the findings to real-world context leakage and security attacks.31:17–40:09 · Guest disagreement 2/10 Universal Geometry of Embeddings and the Platonic Hypothesis Jack connects embedding geometry to the Platonic Representation Hypothesis and discusses model capacity plateaus. swyx raises questions about context length degradation and mathematical matrix bounds.40:09–44:23 · Guest disagreement 2/10 Cognitive Core Concept and Emergence of Reasoning swyx introduces Karpathy's concept of the cognitive core and contrasts human biological efficiency with dense LLMs. Jack agrees in spirit but notes the technical difficulty of separating reasoning from factual memorization.44:23–52:38 · Guest disagreement 2/10 Multimodal Adapters and Theoretical Limits of Small Models Jack explains how CycleGAN inspired unaligned latent space mapping across distinct model architectures. swyx applies this insight to modular multimodal adapters like Gemma 3n.52:39–1:00:03 · Guest disagreement 2/10 Quantifying Model Capacity and Memorization Bounds Jack presents his finding that 32-bit parameters only store roughly 3.6 to 3.9 bits of memorized information. swyx questions whether optimizing memorization capacity conflicts with the primary goal of generalization.1:00:03–1:06:03 · Guest disagreement 1/10 Approximating Training Data Directly from Model Weights Jack explains his method for approximating proprietary fine-tuning data from the weight delta between base and instruct checkpoints. swyx highlights the novelty of using synthetic checkpoints for data reconstruction.1:06:03–1:14:39 · Guest disagreement 5/10 Kuhnian Paradigm Shifts and Dataset-Driven AI Progress Jack presents a contrarian Kuhnian thesis arguing that AI progress is almost entirely dataset-driven rather than architectural. swyx offers pushback, arguing that compute and optimization efficiency act as substantial data multipliers.5:12–8:58 · The hosts pushing back 2/10 Navigating AI Research Meta and Scaling to 8B Models swyx demonstrates solid familiarity with current AI market dynamics and PhD startup valuations. Jack explains the emergence gap between 100M BERT-scale models and 8B parameter models in academic research.8:58–14:31 · The hosts pushing back 3/10 Systems Engineering, GPU Training, and Modern Toolchains swyx shares detailed industry advice regarding HPC resources, PyTorch, DeepSpeed, and Modular Mojo. Jack clarifies his public stance on learning CUDA versus higher-level execution frameworks like vLLM and SGLang.14:31–22:29 · The hosts pushing back 2/10 Information Theory Foundations and Usable Information in LLMs Jack introduces V-information and the idea of measuring usable information under computational constraints. swyx engages deeply by comparing this to Kolmogorov complexity and Shannon information limits.22:30–31:17 · The hosts pushing back 2/10 Inverting Embeddings and Vector Database Privacy Risks Jack details his research on inverting text embeddings and its privacy implications for commercial vector databases. swyx navigates the slides and connects the findings to real-world context leakage and security attacks.31:17–40:09 · The hosts pushing back 2/10 Universal Geometry of Embeddings and the Platonic Hypothesis Jack connects embedding geometry to the Platonic Representation Hypothesis and discusses model capacity plateaus. swyx raises questions about context length degradation and mathematical matrix bounds.40:09–44:23 · The hosts pushing back 3/10 Cognitive Core Concept and Emergence of Reasoning swyx introduces Karpathy's concept of the cognitive core and contrasts human biological efficiency with dense LLMs. Jack agrees in spirit but notes the technical difficulty of separating reasoning from factual memorization.44:23–52:38 · The hosts pushing back 2/10 Multimodal Adapters and Theoretical Limits of Small Models Jack explains how CycleGAN inspired unaligned latent space mapping across distinct model architectures. swyx applies this insight to modular multimodal adapters like Gemma 3n.52:39–1:00:03 · The hosts pushing back 4/10 Quantifying Model Capacity and Memorization Bounds Jack presents his finding that 32-bit parameters only store roughly 3.6 to 3.9 bits of memorized information. swyx questions whether optimizing memorization capacity conflicts with the primary goal of generalization.1:00:03–1:06:03 · The hosts pushing back 2/10 Approximating Training Data Directly from Model Weights Jack explains his method for approximating proprietary fine-tuning data from the weight delta between base and instruct checkpoints. swyx highlights the novelty of using synthetic checkpoints for data reconstruction.1:06:03–1:14:39 · The hosts pushing back 4/10 Kuhnian Paradigm Shifts and Dataset-Driven AI Progress Jack presents a contrarian Kuhnian thesis arguing that AI progress is almost entirely dataset-driven rather than architectural. swyx offers pushback, arguing that compute and optimization efficiency act as substantial data multipliers.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:10:35 Asserting transformers were not strictly necessary

Jack takes a provocative devil's advocate stance, arguing that web-scale pre-training on sophisticated RNNs could have produced ChatGPT without needing transformers.

Hardest push from the hosts ▶ 1:12:18 Challenging the dataset-only thesis with efficiency multipliers

swyx challenges Jack's pure dataset thesis by pointing out that algorithmic and optimizer improvements act as major multipliers equivalent to orders of magnitude more training data.

Biggest teaching moment ▶ 15:25 Defining extractable V-information under computation limits

Jack provides a clear theoretical distinction showing why two files with identical Shannon bit-entropy differ vastly in usable, extractable information.

The host holds their own ▶ 12:29 Detailed analysis of Modular Mojo and hardware compilation

swyx demonstrates deep technical and industry insight by breaking down Chris Lattner's Mojo compiler approach to CUDA replacement and fast kernel experimentation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Navigating AI Research Meta and Scaling to 8B Models 5422 swyx demonstrates solid familiarity with current AI market dynamics and PhD startup valuations. Jack explains the emergence gap between 100M BERT-scale models and 8B parameter models in academic research.
Systems Engineering, GPU Training, and Modern Toolchains 8223 swyx shares detailed industry advice regarding HPC resources, PyTorch, DeepSpeed, and Modular Mojo. Jack clarifies his public stance on learning CUDA versus higher-level execution frameworks like vLLM and SGLang.
Information Theory Foundations and Usable Information in LLMs 7612 Jack introduces V-information and the idea of measuring usable information under computational constraints. swyx engages deeply by comparing this to Kolmogorov complexity and Shannon information limits.
Inverting Embeddings and Vector Database Privacy Risks 6522 Jack details his research on inverting text embeddings and its privacy implications for commercial vector databases. swyx navigates the slides and connects the findings to real-world context leakage and security attacks.
Universal Geometry of Embeddings and the Platonic Hypothesis 5422 Jack connects embedding geometry to the Platonic Representation Hypothesis and discusses model capacity plateaus. swyx raises questions about context length degradation and mathematical matrix bounds.
Cognitive Core Concept and Emergence of Reasoning 7323 swyx introduces Karpathy's concept of the cognitive core and contrasts human biological efficiency with dense LLMs. Jack agrees in spirit but notes the technical difficulty of separating reasoning from factual memorization.
Multimodal Adapters and Theoretical Limits of Small Models 6422 Jack explains how CycleGAN inspired unaligned latent space mapping across distinct model architectures. swyx applies this insight to modular multimodal adapters like Gemma 3n.
Quantifying Model Capacity and Memorization Bounds 6524 Jack presents his finding that 32-bit parameters only store roughly 3.6 to 3.9 bits of memorized information. swyx questions whether optimizing memorization capacity conflicts with the primary goal of generalization.
Approximating Training Data Directly from Model Weights 5512 Jack explains his method for approximating proprietary fine-tuning data from the weight delta between base and instruct checkpoints. swyx highlights the novelty of using synthetic checkpoints for data reconstruction.
Kuhnian Paradigm Shifts and Dataset-Driven AI Progress 6554 Jack presents a contrarian Kuhnian thesis arguing that AI progress is almost entirely dataset-driven rather than architectural. swyx offers pushback, arguing that compute and optimization efficiency act as substantial data multipliers.

Statements from this episode (23)

Assertion Not checkable as stated
Jack Morris: Most AI research was previously open, but is now closed
“Most stuff was open. Now most stuff is not open.”
Jack Morris Jul 2, 2025 ▶ 3:39
Assertion Not checkable as stated
Jack Morris: Fundamental AI science shifted to companies due to academic compute limits
“That's when I think things really started to change in terms of the types of questions you wanted to ask can't always be answered with academic resources. So a lot of the like fundamental kind of like boundary pushing and AI science moved into companies.”
Jack Morris Jul 2, 2025 ▶ 4:22
Insight
Morris: Re-implementing new paradigm shifts quickly is the best AI grad strategy
“Honestly, if I were to give advice to a younger grad student, I think the way to do it would be literally just like sit and wait until the next kind of paradigm shift and then just immediately start working as fast as you can to like re-implement it. Like, I d…”
Jack Morris Jul 2, 2025 ▶ 5:28
Assertion Supported
Swyx: Stanford RL students founded pre-product startup with $500M valuation
“I just saw this morning that one of the recent Stanford grad students who worked on RL, they just started their company and they're worth five hundred million. It's like absolutely bonkers right now. It's just like no product, just three dudes, you know, sitti…”
Shawn Wang Jul 2, 2025 ▶ 6:38
Opinion
Morris: Two years of academic AI research on small models was inconsequential
“There was like kind of two years where everyone in academia was working on like smaller models and none of it really mattered.”
Jack Morris Jul 2, 2025 ▶ 8:36
Assertion Contradicted
Morris: Top AI graduate programs do not teach multi-node distributed training
“Oh, to be clear, they don't teach you anything, like anything, like if you see a paper coming out from even, you know, Stanford, they're probably the best school in AI if you had to choose. And it's not like they're learning how to do like multi-node distribut…”
Jack Morris Jul 2, 2025 ▶ 9:38
Insight
Morris: Deep understanding of GPU architecture makes engineers exceptionally hireable
“That said, if you do it, you're, you've gotta be one of the most hireable people in the world. Like if you like, Really deeply understand the architecture of the new GPUs coming out and how to control it. You're in a very small handful of people and like every…”
Jack Morris Jul 2, 2025 ▶ 12:13
Prediction Not checkable as stated
Morris: vLLM and SGLang are here to stay and will grow more complex
“I also think, ah, VLLM and SGLang seem, like, really good and important and here to stay. Like, they'll probably just get larger and more complex to accommodate future systems”
Jack Morris Jul 2, 2025 ▶ 14:10
Insight
Morris: Information in AI should be measured under computational constraints
“There's this theoretical framework proposed in this paper, which is a theory of usable information under computational constraints from 20, 20. It really doesn't have that much press. They're not aren't as many citations as you would think, but I think it's a …”
Jack Morris Jul 2, 2025 ▶ 16:30
Opinion
Morris: Machine learning lacks a fundamental unit of deep learning information
“I don't think we know what a bit is yet in terms of like deep learning models.”
Jack Morris Jul 2, 2025 ▶ 19:21
Assertion Supported
Morris: New embedding inversion model exactly recovers 90% of source text
“Like we ended up building a system that can do this quite well, like taking an embedding and I think our highlight number is like at a certain length, like a long sentence length, we can get 90% of the text back exactly.”
Jack Morris Jul 2, 2025 ▶ 26:53
Assertion Supported
Morris: Embedding inversion requires access to and repeated queries of the encoder
“Like none of the vector to text stuff works unless you have this assumption of like knowing the encoder and also being able to make a lot of queries to it.”
Jack Morris Jul 2, 2025 ▶ 32:44
Assertion Supported
Morris: Language models hit a hard memorization plateau regardless of dataset scaling
“Like, no matter how you scale the training size, you hit this like perfect, perfect ish plateau in auto memorization, which we call the model capacity.”
Jack Morris Jul 2, 2025 ▶ 38:32
Opinion
Morris: No Evidence We Can Build Pure Reasoning Models Without World Knowledge
“I don't think we have a lot of evidence that we can build a system like this that like is really, really good at reasoning, but really dumb about the world. Like, I don't know if we have the tools.”
Jack Morris Jul 2, 2025 ▶ 41:31
Assertion Supported
Morris: CycleGAN mapping aligns disparate model embeddings without paired data
“We took it and we applied it to model embeddings where instead of zebras and horses, we have like BERT embeddings and GPT embeddings, or like two completely different models with different architectures. So I think these are GTR, which is a T five based retrie…”
Jack Morris Jul 2, 2025 ▶ 46:25
Insight
Morris: Small models should be defined as runnable on a single GPU
“I think that we should establish the definition of small model as being a model that a grad student can inference at reasonable time on a single GPU. Which is probably like seven B maybe. I don't think 27 is small under any reasonable.”
Jack Morris Jul 2, 2025 ▶ 52:05
Assertion Supported
Morris: 32-bit transformer models store only 3.6 to 3.9 bits per parameter
“Transformers that are trained in 32 bit precision, we approximate can store about 3.6 bits of information to maybe 3.9 bits somewhere in there per parameter.”
Jack Morris Jul 2, 2025 ▶ 56:00
Prediction Open · timeframe Jul 2030
Morris: LLaMA architecture will likely store more information per parameter than GPT
“Maybe even if we tested this with LALAMA architecture, like, there's sort of like a GPT++ architecture, like, I would guess that can store better data just because the kind of numerical flow is a little bit better, the nonlinearities are maybe, like, A little …”
Jack Morris Jul 2, 2025 ▶ 57:30
Opinion
Morris: Open model creators do not use differential privacy or anonymization
“I would be extremely surprised if they do any type of like private training. Like there are these mechanisms for doing like differentially private language model training, or even just anonymization in the pre-training pipeline. I bet they don't do any of that…”
Jack Morris Jul 2, 2025 ▶ 1:02:27
Insight
Morris: Weight deltas can reconstruct a competitor's proprietary fine-tuning dataset
“There's some tricks to it, but it's basically just like gradient based selection based on this weight difference. And it seems to be okay. Like it can get us pretty good training data. So I guess if you actually wanted to use this, it would be like your compet…”
Jack Morris Jul 2, 2025 ▶ 1:04:42
Insight
Morris: AI paradigm shifts are driven by novel datasets, not architectures
“I think, like, all of the things that I would consider paradigm shifts in the Kuhnian sense came from a new technique, but trained on new data, and I think the new data is super, super important”
Jack Morris Jul 2, 2025 ▶ 1:09:20
What-if
Morris: ChatGPT could likely have been built using RNNs instead of Transformers
“And I think like, we honestly probably could have gotten this with RNNs. I know like the scaling laws paper shows that RNNs have worse curves for scaling, but probably people would have been like, I bet you could have built chat GPT with a very sophisticated R…”
Jack Morris Jul 2, 2025 ▶ 1:10:28
Prediction Not checkable as stated
Morris: The next AI paradigm shift will stem from an unused data source
“And so whatever the fifth thing is, whether it's Video or embodied AI or some kind of crazy innovation on reasoning models. Whatever comes next will probably be some type of new data source that we're not using yet.”
Jack Morris Jul 2, 2025 ▶ 1:12:04
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.