Aug 17, 2024 · 1h 10m · latent-space
Answer.ai & AI Magic with Jeremy Howard
⌖ your search result is the highlighted band (46:00–46:26). Playback starts there
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Latent Space podcast episode, Jeremy Howard discusses Answer.ai's Public Benefit Corporation mission, breaks down systems engineering breakthroughs for local fine-tuning, challenges decoder-only model dogmas, and introduces new developer tools like FastHTML and Dialogue Engineering.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jeremy directly rejects Swix's premise regarding model merging, firmly clarifying that distributing merged full weight models rather than quantized bases with merged adapters is inefficient.
Hardest push from the hosts ▶ 46:02 Swix challenges Jeremy's dismissal of model mergingSwix openly defends model merging against Jeremy's categorical assertion, pointing out that merging models provides unique value when working without training data.
Biggest teaching moment ▶ 7:28 Structural breakdown of non-profit board misalignmentsJeremy methodically dissects why academic non-profit boards cannot effectively govern commercial entities when employee compensation is tied directly to equity and profit.
The host holds their own ▶ 3:20 Swix frames multi-phase pre-training trendsSwix demonstrates deep industry monitoring by citing Snowflake Arctic's exact web-to-code dataset mix transitions across training phases.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Rethinking Fine-Tuning as a Pre-Training Continuum | 6 | 6 | 3 | 2 | Swix references Jeremy's past hot take on the death of fine-tuning and cites an ICLR paper on data-driven priors. Jeremy clarifies that fine-tuning is part of a pre-training continuum and rejects starting from random weights. | |
| Multi-Phase Pre-Training and Optimization Schedules | 7 | 6 | 5 | 3 | Swix introduces Snowflake Arctic's multi-phase pre-training and Meta's schedule-free optimizers. Jeremy bluntly dismisses schedule-free approaches, advocating for functional learning rate schedules and maximum parameter control. | |
| AI Governance, OpenAI Dynamics, and Public Benefit Corporations | 6 | 7 | 4 | 1 | Alessio and Jeremy explore OpenAI's governance implosion and the mechanics of Public Benefit Corporations. Jeremy breaks down why non-profit boards overseeing equity-driven commercial entities inevitably fail. | |
| Answer.ai's Organizational Structure and Self-Directed Projects | 5 | 5 | 2 | 1 | Jeremy explains Answer.ai's flat structure and organic project origination, citing spontaneous developments like BERT-24 and WebGPU tooling. | |
| Evaluating Practical Talent Over Academic Credentials | 5 | 5 | 3 | 1 | Alessio asks about screening practical talent versus traditional credentials. Jeremy highlights hiring unconventional problem-solvers based on code quality and proven tenacity rather than academic pedigree. | |
| Architectural Bets: The Resurgence of Encoder-Decoder Models | 6 | 7 | 4 | 2 | Swix queries the non-consensus bet on encoder-decoder models. Jeremy explains why decoder-only models are over-indexed and structurally inefficient for classification and fixed representation tasks. | |
| Engineering Challenges in FSDP, QLoRA, and QDoRA Implementation | 6 | 6 | 3 | 1 | Alessio asks about fine-tuning 70B models on consumer hardware. Jeremy details the grueling systems engineering required to unify FSDP, bitsandbytes, and PEFT into a coherent workflow. | |
| Local Inference Optimization, Quantization, and Merged Adapters | 7 | 7 | 6 | 6 | Swix pushes back on Jeremy's assertion that nobody should merge models by citing zero-data merging benefits. Jeremy forcefully clarifies that distributing merged full weights rather than merged adapters is the anti-pattern. | |
| FastHTML: Bringing Web Fundamentals to Pure Python Development | 5 | 6 | 3 | 1 | Jeremy unveils FastHTML, critiquing modern JavaScript SPA complexity and advocating for pure Python web development built directly on web standards and HTMX. | |
| Dialogue Engineering, AI Magic, and Practical Developer Tooling | 6 | 6 | 3 | 2 | Swix asks about AI-assisted development tools. Jeremy introduces dialogue engineering and AI Magic as an interactive middle ground between basic chat teletypes and complex IDEs. |