Jun 25, 2024 · 1h 32m · latent-space
State of the Art: Training 70B LLMs on 10,000 H100 clusters
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Latent Space Podcast, Jonathan Frankle (Databricks) and Josh Albrecht (Imbue) discuss the low-level systems engineering, bare-metal hardware realities, and empirical optimization required to train 70B+ parameter foundation models. They share practical insights on cluster architecture, debugging distributed training bottlenecks, fixing flawed evaluation benchmarks, and building reliable code-centric AI agents.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Jonathan reacts with exasperation ('Oh, for the love of God'), arguing the benchmark fails to test holistic context utilization and measures nothing realistic.
Hardest push from the hosts ▶ 1:15:16 Swyx defending abstract evaluation testsSwyx pushes back against the guests' dismissal of abstract tests like ARC-AGI by pointing out humanity's successful reliance on SATs and IQ tests as capability predictors.
Biggest teaching moment ▶ 55:45 Josh explaining emergence as a metric mirageJosh systematically disproves the premise of sudden model emergence by explaining that discontinuous accuracy jumps are simply log-linear perplexity improvements passing threshold bands.
The host holds their own ▶ 37:49 Swyx challenging bare-metal abstraction choicesDrawing on his AWS background, Swyx challenges Imbue's custom orchestration layer by asserting that Kubernetes provides essential vendor-agnostic portability for scaling.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Introductions and Catching Up on the LLM Training Landscape | 5 | 3 | 1 | 2 | Swyx introduces the guests, tracks Databricks' recent developments, and asks clarifying questions regarding Shutterstock licensing arrangements and data provenance. | |
| DBRX Architecture, Dinosaur Mascots, and the Hair Dye Bet | 6 | 2 | 1 | 1 | Swyx recaps DBRX technical specs (132B total, 36B active, 12T tokens) while Jonathan jokes about team bets and dinosaur mascots. | |
| Imbue Open Sources the 70B Foundation Model Training Playbook | 4 | 6 | 1 | 1 | Josh outlines Imbue's 70B playbook release across infra, evals, and CARBS hyperparameter optimization; Jonathan strongly validates the release as rare and essential. | |
| Hardware Realities and Bizarre GPU Cluster Failure Modes | 4 | 6 | 1 | 1 | Jonathan and Josh share war stories about cluster failures, from stolen InfiniBand cables to memory correction latency and GPUs outputting incorrect arithmetic. | |
| Co-Designing a 4,096 H100 GPU Cluster with Voltage Park | 5 | 5 | 1 | 2 | Swyx inquires about the 4,096 H100 cluster topology and partnership with Voltage Park, prompting Josh to explain their non-standard 3-tier networking choices. | |
| Automated Bare-Metal Provisioning and Low-Level Hardware Health Checks | 3 | 7 | 1 | 1 | Josh walks through low-level automated bringup using Metal-as-a-Service, out-of-band management interfaces, and debugging directly with firmware vendors like Dell. | |
| Full-Stack Infrastructure: DIY Bare-Metal vs. Managed Cloud Platforms | 5 | 5 | 2 | 3 | Swyx compares Imbue's bare-metal approach with Mosaic's cloud abstraction, leading Jonathan to explain why even cloud users end up dealing with low-level physical realities. | |
| Cluster Reliability Metrics, Triage Automation, and Node Churn | 5 | 6 | 2 | 2 | Swyx presses on cluster failure rates; Josh corrects the uniform failure assumption and details automated dmesg anomaly detection on node reboots. | |
| Cluster Bringup Logistics: Small Teams, Cable Management, and Physical Assembly | 4 | 6 | 1 | 1 | Josh details the physical bringup of 24,000 cable connections executed by a small core team in coordination with vendor technicians. | |
| Distributed Training Tools, S3 Local Mirrors, and Uber's Kraken Registry | 5 | 6 | 1 | 1 | Josh discusses memory fragmentation jitter and highlights simple tooling choices like BitTorrent-based Kraken image distribution over heavy cloud filesystems. | |
| Contrasting Infrastructure Philosophies: Minimal Abstractions vs. Standardized APIs | 6 | 5 | 2 | 4 | Swyx pushes back from an AWS perspective arguing Kubernetes standardizes infrastructure, but Josh counters that Kubernetes is mismatched with training job failures. | |
| Training Dynamics: MoE Network Bandwidth, HSDP, and Next-Gen Interconnects | 5 | 7 | 1 | 1 | Jonathan explains how MoE training stresses network bandwidth, requiring HSDP, and compares Google's 3D torus TPU design with Nvidia's grouped NVLink topologies. | |
| Imbue's Specialized Focus on Code, Reasoning, and Text-Only Architectures | 4 | 5 | 1 | 1 | Swyx probes Imbue's decision to focus purely on text and coding rather than vision; Josh explains how agent reasoning can outsource multimodal components. | |
| Evaluation Metric Precision: Multiple-Choice Perplexity and Data Mix Schedules | 4 | 6 | 3 | 1 | Josh explains how CARBS optimizes hyperparameters across cost scales; Jonathan and Josh cordially disagree on whether dynamic data mix schedules provide real gains. | |
| The Mirage of Emergence: Log-Scale Perplexity vs. Metric Discontinuities | 4 | 8 | 2 | 2 | Swyx asks whether emergence breaks CARBS scaling laws; Josh schools on how emergence is an artifact of non-linear accuracy metrics rather than log-scale loss. | |
| Cleaning Open-Source Benchmarks and Addressing Question Ambiguity | 4 | 7 | 1 | 1 | Josh breaks down human annotation efforts to clean benchmark ambiguity, demonstrating that top models reach near total saturation on well-formed questions. | |
| Programmatic Benchmarks, Code Understanding, and Realistic Test Coverage | 5 | 6 | 1 | 1 | Swyx questions the scope of code understanding; Josh and Jonathan discuss SWE-bench and the difficulty of evaluating nuanced code without leaking bad patches. | |
| Navigating Imperfect Evaluations and Modeling Uncertainty in AI Agents | 4 | 7 | 1 | 1 | Jonathan reflects on living with imperfect evaluations when directing research, while Josh outlines Imbue's work on detecting question ambiguity and agent uncertainty. | |
| Evaluating Abstract Reasoning (ARC-AGI) vs. Pragmatic Business Problems | 6 | 6 | 2 | 4 | Josh and Jonathan critique abstract reasoning benchmarks like ARC-AGI as synthetic toys, prompting Swyx to defend standardized abstract testing as historically predictive. | |
| The Limitations of 'Needle in a Haystack' and Long-Context Agent Scenarios | 4 | 8 | 4 | 1 | Jonathan vehemently dismisses Needle in a Haystack as disconnected from real-world utility, while Josh explains why long context requires active relevance filtering. | |
| Code Execution as the Universal Agent Interface vs. Hard-Coded Tool Calling | 5 | 6 | 1 | 1 | Josh argues that general code execution is vastly superior to rigid tool-calling schemas, while Jonathan highlights the value of native SQL execution for structured enterprise data. | |
| Knowledge Graphs, Data Schemas, and Bounding Agent Action Spaces | 5 | 5 | 1 | 2 | Swyx prompts on whether knowledge graphs offer an enduring abstraction; Josh and Jonathan agree they fit specific ontologies but boundary control remains challenging. | |
| Future Roadmaps: Databricks Science Sharing and the Abra-Kadabra Model Lineage | 4 | 4 | 1 | 1 | Josh outlines Imbue's roadmap toward practical coding agents while Jonathan hints at future DBRX evolutionary models (Abra, Kadabra, and Alakazam) before closing. |