Jun 25, 2024 · 1h 32m · latent-space

State of the Art: Training 70B LLMs on 10,000 H100 clusters

Josh Albrecht · 40m spoken Jonathan Frankle · 27m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Podcast, Jonathan Frankle (Databricks) and Josh Albrecht (Imbue) discuss the low-level systems engineering, bare-metal hardware realities, and empirical optimization required to train 70B+ parameter foundation models. They share practical insights on cluster architecture, debugging distributed training bottlenecks, fixing flawed evaluation benchmarks, and building reliable code-centric AI agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.6 Guest teaching 5.7 Guest disagreement 1.4 The hosts pushing back 1.6
05100:0020:0040:001:00:001:20:000:00–4:15 · The hosts as informed peer 5/10 Introductions and Catching Up on the LLM Training Landscape Swyx introduces the guests, tracks Databricks' recent developments, and asks clarifying questions regarding Shutterstock licensing arrangements and data provenance.4:15–6:26 · The hosts as informed peer 6/10 DBRX Architecture, Dinosaur Mascots, and the Hair Dye Bet Swyx recaps DBRX technical specs (132B total, 36B active, 12T tokens) while Jonathan jokes about team bets and dinosaur mascots.6:26–10:14 · The hosts as informed peer 4/10 Imbue Open Sources the 70B Foundation Model Training Playbook Josh outlines Imbue's 70B playbook release across infra, evals, and CARBS hyperparameter optimization; Jonathan strongly validates the release as rare and essential.10:14–13:14 · The hosts as informed peer 4/10 Hardware Realities and Bizarre GPU Cluster Failure Modes Jonathan and Josh share war stories about cluster failures, from stolen InfiniBand cables to memory correction latency and GPUs outputting incorrect arithmetic.13:14–16:41 · The hosts as informed peer 5/10 Co-Designing a 4,096 H100 GPU Cluster with Voltage Park Swyx inquires about the 4,096 H100 cluster topology and partnership with Voltage Park, prompting Josh to explain their non-standard 3-tier networking choices.16:41–21:23 · The hosts as informed peer 3/10 Automated Bare-Metal Provisioning and Low-Level Hardware Health Checks Josh walks through low-level automated bringup using Metal-as-a-Service, out-of-band management interfaces, and debugging directly with firmware vendors like Dell.21:23–24:35 · The hosts as informed peer 5/10 Full-Stack Infrastructure: DIY Bare-Metal vs. Managed Cloud Platforms Swyx compares Imbue's bare-metal approach with Mosaic's cloud abstraction, leading Jonathan to explain why even cloud users end up dealing with low-level physical realities.24:35–28:08 · The hosts as informed peer 5/10 Cluster Reliability Metrics, Triage Automation, and Node Churn Swyx presses on cluster failure rates; Josh corrects the uniform failure assumption and details automated dmesg anomaly detection on node reboots.28:08–32:34 · The hosts as informed peer 4/10 Cluster Bringup Logistics: Small Teams, Cable Management, and Physical Assembly Josh details the physical bringup of 24,000 cable connections executed by a small core team in coordination with vendor technicians.32:34–35:53 · The hosts as informed peer 5/10 Distributed Training Tools, S3 Local Mirrors, and Uber's Kraken Registry Josh discusses memory fragmentation jitter and highlights simple tooling choices like BitTorrent-based Kraken image distribution over heavy cloud filesystems.35:53–39:12 · The hosts as informed peer 6/10 Contrasting Infrastructure Philosophies: Minimal Abstractions vs. Standardized APIs Swyx pushes back from an AWS perspective arguing Kubernetes standardizes infrastructure, but Josh counters that Kubernetes is mismatched with training job failures.39:12–44:41 · The hosts as informed peer 5/10 Training Dynamics: MoE Network Bandwidth, HSDP, and Next-Gen Interconnects Jonathan explains how MoE training stresses network bandwidth, requiring HSDP, and compares Google's 3D torus TPU design with Nvidia's grouped NVLink topologies.44:41–50:16 · The hosts as informed peer 4/10 Imbue's Specialized Focus on Code, Reasoning, and Text-Only Architectures Swyx probes Imbue's decision to focus purely on text and coding rather than vision; Josh explains how agent reasoning can outsource multimodal components.50:16–55:35 · The hosts as informed peer 4/10 Evaluation Metric Precision: Multiple-Choice Perplexity and Data Mix Schedules Josh explains how CARBS optimizes hyperparameters across cost scales; Jonathan and Josh cordially disagree on whether dynamic data mix schedules provide real gains.55:35–57:55 · The hosts as informed peer 4/10 The Mirage of Emergence: Log-Scale Perplexity vs. Metric Discontinuities Swyx asks whether emergence breaks CARBS scaling laws; Josh schools on how emergence is an artifact of non-linear accuracy metrics rather than log-scale loss.57:55–1:02:21 · The hosts as informed peer 4/10 Cleaning Open-Source Benchmarks and Addressing Question Ambiguity Josh breaks down human annotation efforts to clean benchmark ambiguity, demonstrating that top models reach near total saturation on well-formed questions.1:02:21–1:06:52 · The hosts as informed peer 5/10 Programmatic Benchmarks, Code Understanding, and Realistic Test Coverage Swyx questions the scope of code understanding; Josh and Jonathan discuss SWE-bench and the difficulty of evaluating nuanced code without leaking bad patches.1:06:52–1:11:11 · The hosts as informed peer 4/10 Navigating Imperfect Evaluations and Modeling Uncertainty in AI Agents Jonathan reflects on living with imperfect evaluations when directing research, while Josh outlines Imbue's work on detecting question ambiguity and agent uncertainty.1:11:11–1:15:40 · The hosts as informed peer 6/10 Evaluating Abstract Reasoning (ARC-AGI) vs. Pragmatic Business Problems Josh and Jonathan critique abstract reasoning benchmarks like ARC-AGI as synthetic toys, prompting Swyx to defend standardized abstract testing as historically predictive.1:15:40–1:19:39 · The hosts as informed peer 4/10 The Limitations of 'Needle in a Haystack' and Long-Context Agent Scenarios Jonathan vehemently dismisses Needle in a Haystack as disconnected from real-world utility, while Josh explains why long context requires active relevance filtering.1:19:39–1:23:48 · The hosts as informed peer 5/10 Code Execution as the Universal Agent Interface vs. Hard-Coded Tool Calling Josh argues that general code execution is vastly superior to rigid tool-calling schemas, while Jonathan highlights the value of native SQL execution for structured enterprise data.1:23:48–1:26:23 · The hosts as informed peer 5/10 Knowledge Graphs, Data Schemas, and Bounding Agent Action Spaces Swyx prompts on whether knowledge graphs offer an enduring abstraction; Josh and Jonathan agree they fit specific ontologies but boundary control remains challenging.1:26:23–1:29:53 · The hosts as informed peer 4/10 Future Roadmaps: Databricks Science Sharing and the Abra-Kadabra Model Lineage Josh outlines Imbue's roadmap toward practical coding agents while Jonathan hints at future DBRX evolutionary models (Abra, Kadabra, and Alakazam) before closing.0:00–4:15 · Guest teaching 3/10 Introductions and Catching Up on the LLM Training Landscape Swyx introduces the guests, tracks Databricks' recent developments, and asks clarifying questions regarding Shutterstock licensing arrangements and data provenance.4:15–6:26 · Guest teaching 2/10 DBRX Architecture, Dinosaur Mascots, and the Hair Dye Bet Swyx recaps DBRX technical specs (132B total, 36B active, 12T tokens) while Jonathan jokes about team bets and dinosaur mascots.6:26–10:14 · Guest teaching 6/10 Imbue Open Sources the 70B Foundation Model Training Playbook Josh outlines Imbue's 70B playbook release across infra, evals, and CARBS hyperparameter optimization; Jonathan strongly validates the release as rare and essential.10:14–13:14 · Guest teaching 6/10 Hardware Realities and Bizarre GPU Cluster Failure Modes Jonathan and Josh share war stories about cluster failures, from stolen InfiniBand cables to memory correction latency and GPUs outputting incorrect arithmetic.13:14–16:41 · Guest teaching 5/10 Co-Designing a 4,096 H100 GPU Cluster with Voltage Park Swyx inquires about the 4,096 H100 cluster topology and partnership with Voltage Park, prompting Josh to explain their non-standard 3-tier networking choices.16:41–21:23 · Guest teaching 7/10 Automated Bare-Metal Provisioning and Low-Level Hardware Health Checks Josh walks through low-level automated bringup using Metal-as-a-Service, out-of-band management interfaces, and debugging directly with firmware vendors like Dell.21:23–24:35 · Guest teaching 5/10 Full-Stack Infrastructure: DIY Bare-Metal vs. Managed Cloud Platforms Swyx compares Imbue's bare-metal approach with Mosaic's cloud abstraction, leading Jonathan to explain why even cloud users end up dealing with low-level physical realities.24:35–28:08 · Guest teaching 6/10 Cluster Reliability Metrics, Triage Automation, and Node Churn Swyx presses on cluster failure rates; Josh corrects the uniform failure assumption and details automated dmesg anomaly detection on node reboots.28:08–32:34 · Guest teaching 6/10 Cluster Bringup Logistics: Small Teams, Cable Management, and Physical Assembly Josh details the physical bringup of 24,000 cable connections executed by a small core team in coordination with vendor technicians.32:34–35:53 · Guest teaching 6/10 Distributed Training Tools, S3 Local Mirrors, and Uber's Kraken Registry Josh discusses memory fragmentation jitter and highlights simple tooling choices like BitTorrent-based Kraken image distribution over heavy cloud filesystems.35:53–39:12 · Guest teaching 5/10 Contrasting Infrastructure Philosophies: Minimal Abstractions vs. Standardized APIs Swyx pushes back from an AWS perspective arguing Kubernetes standardizes infrastructure, but Josh counters that Kubernetes is mismatched with training job failures.39:12–44:41 · Guest teaching 7/10 Training Dynamics: MoE Network Bandwidth, HSDP, and Next-Gen Interconnects Jonathan explains how MoE training stresses network bandwidth, requiring HSDP, and compares Google's 3D torus TPU design with Nvidia's grouped NVLink topologies.44:41–50:16 · Guest teaching 5/10 Imbue's Specialized Focus on Code, Reasoning, and Text-Only Architectures Swyx probes Imbue's decision to focus purely on text and coding rather than vision; Josh explains how agent reasoning can outsource multimodal components.50:16–55:35 · Guest teaching 6/10 Evaluation Metric Precision: Multiple-Choice Perplexity and Data Mix Schedules Josh explains how CARBS optimizes hyperparameters across cost scales; Jonathan and Josh cordially disagree on whether dynamic data mix schedules provide real gains.55:35–57:55 · Guest teaching 8/10 The Mirage of Emergence: Log-Scale Perplexity vs. Metric Discontinuities Swyx asks whether emergence breaks CARBS scaling laws; Josh schools on how emergence is an artifact of non-linear accuracy metrics rather than log-scale loss.57:55–1:02:21 · Guest teaching 7/10 Cleaning Open-Source Benchmarks and Addressing Question Ambiguity Josh breaks down human annotation efforts to clean benchmark ambiguity, demonstrating that top models reach near total saturation on well-formed questions.1:02:21–1:06:52 · Guest teaching 6/10 Programmatic Benchmarks, Code Understanding, and Realistic Test Coverage Swyx questions the scope of code understanding; Josh and Jonathan discuss SWE-bench and the difficulty of evaluating nuanced code without leaking bad patches.1:06:52–1:11:11 · Guest teaching 7/10 Navigating Imperfect Evaluations and Modeling Uncertainty in AI Agents Jonathan reflects on living with imperfect evaluations when directing research, while Josh outlines Imbue's work on detecting question ambiguity and agent uncertainty.1:11:11–1:15:40 · Guest teaching 6/10 Evaluating Abstract Reasoning (ARC-AGI) vs. Pragmatic Business Problems Josh and Jonathan critique abstract reasoning benchmarks like ARC-AGI as synthetic toys, prompting Swyx to defend standardized abstract testing as historically predictive.1:15:40–1:19:39 · Guest teaching 8/10 The Limitations of 'Needle in a Haystack' and Long-Context Agent Scenarios Jonathan vehemently dismisses Needle in a Haystack as disconnected from real-world utility, while Josh explains why long context requires active relevance filtering.1:19:39–1:23:48 · Guest teaching 6/10 Code Execution as the Universal Agent Interface vs. Hard-Coded Tool Calling Josh argues that general code execution is vastly superior to rigid tool-calling schemas, while Jonathan highlights the value of native SQL execution for structured enterprise data.1:23:48–1:26:23 · Guest teaching 5/10 Knowledge Graphs, Data Schemas, and Bounding Agent Action Spaces Swyx prompts on whether knowledge graphs offer an enduring abstraction; Josh and Jonathan agree they fit specific ontologies but boundary control remains challenging.1:26:23–1:29:53 · Guest teaching 4/10 Future Roadmaps: Databricks Science Sharing and the Abra-Kadabra Model Lineage Josh outlines Imbue's roadmap toward practical coding agents while Jonathan hints at future DBRX evolutionary models (Abra, Kadabra, and Alakazam) before closing.0:00–4:15 · Guest disagreement 1/10 Introductions and Catching Up on the LLM Training Landscape Swyx introduces the guests, tracks Databricks' recent developments, and asks clarifying questions regarding Shutterstock licensing arrangements and data provenance.4:15–6:26 · Guest disagreement 1/10 DBRX Architecture, Dinosaur Mascots, and the Hair Dye Bet Swyx recaps DBRX technical specs (132B total, 36B active, 12T tokens) while Jonathan jokes about team bets and dinosaur mascots.6:26–10:14 · Guest disagreement 1/10 Imbue Open Sources the 70B Foundation Model Training Playbook Josh outlines Imbue's 70B playbook release across infra, evals, and CARBS hyperparameter optimization; Jonathan strongly validates the release as rare and essential.10:14–13:14 · Guest disagreement 1/10 Hardware Realities and Bizarre GPU Cluster Failure Modes Jonathan and Josh share war stories about cluster failures, from stolen InfiniBand cables to memory correction latency and GPUs outputting incorrect arithmetic.13:14–16:41 · Guest disagreement 1/10 Co-Designing a 4,096 H100 GPU Cluster with Voltage Park Swyx inquires about the 4,096 H100 cluster topology and partnership with Voltage Park, prompting Josh to explain their non-standard 3-tier networking choices.16:41–21:23 · Guest disagreement 1/10 Automated Bare-Metal Provisioning and Low-Level Hardware Health Checks Josh walks through low-level automated bringup using Metal-as-a-Service, out-of-band management interfaces, and debugging directly with firmware vendors like Dell.21:23–24:35 · Guest disagreement 2/10 Full-Stack Infrastructure: DIY Bare-Metal vs. Managed Cloud Platforms Swyx compares Imbue's bare-metal approach with Mosaic's cloud abstraction, leading Jonathan to explain why even cloud users end up dealing with low-level physical realities.24:35–28:08 · Guest disagreement 2/10 Cluster Reliability Metrics, Triage Automation, and Node Churn Swyx presses on cluster failure rates; Josh corrects the uniform failure assumption and details automated dmesg anomaly detection on node reboots.28:08–32:34 · Guest disagreement 1/10 Cluster Bringup Logistics: Small Teams, Cable Management, and Physical Assembly Josh details the physical bringup of 24,000 cable connections executed by a small core team in coordination with vendor technicians.32:34–35:53 · Guest disagreement 1/10 Distributed Training Tools, S3 Local Mirrors, and Uber's Kraken Registry Josh discusses memory fragmentation jitter and highlights simple tooling choices like BitTorrent-based Kraken image distribution over heavy cloud filesystems.35:53–39:12 · Guest disagreement 2/10 Contrasting Infrastructure Philosophies: Minimal Abstractions vs. Standardized APIs Swyx pushes back from an AWS perspective arguing Kubernetes standardizes infrastructure, but Josh counters that Kubernetes is mismatched with training job failures.39:12–44:41 · Guest disagreement 1/10 Training Dynamics: MoE Network Bandwidth, HSDP, and Next-Gen Interconnects Jonathan explains how MoE training stresses network bandwidth, requiring HSDP, and compares Google's 3D torus TPU design with Nvidia's grouped NVLink topologies.44:41–50:16 · Guest disagreement 1/10 Imbue's Specialized Focus on Code, Reasoning, and Text-Only Architectures Swyx probes Imbue's decision to focus purely on text and coding rather than vision; Josh explains how agent reasoning can outsource multimodal components.50:16–55:35 · Guest disagreement 3/10 Evaluation Metric Precision: Multiple-Choice Perplexity and Data Mix Schedules Josh explains how CARBS optimizes hyperparameters across cost scales; Jonathan and Josh cordially disagree on whether dynamic data mix schedules provide real gains.55:35–57:55 · Guest disagreement 2/10 The Mirage of Emergence: Log-Scale Perplexity vs. Metric Discontinuities Swyx asks whether emergence breaks CARBS scaling laws; Josh schools on how emergence is an artifact of non-linear accuracy metrics rather than log-scale loss.57:55–1:02:21 · Guest disagreement 1/10 Cleaning Open-Source Benchmarks and Addressing Question Ambiguity Josh breaks down human annotation efforts to clean benchmark ambiguity, demonstrating that top models reach near total saturation on well-formed questions.1:02:21–1:06:52 · Guest disagreement 1/10 Programmatic Benchmarks, Code Understanding, and Realistic Test Coverage Swyx questions the scope of code understanding; Josh and Jonathan discuss SWE-bench and the difficulty of evaluating nuanced code without leaking bad patches.1:06:52–1:11:11 · Guest disagreement 1/10 Navigating Imperfect Evaluations and Modeling Uncertainty in AI Agents Jonathan reflects on living with imperfect evaluations when directing research, while Josh outlines Imbue's work on detecting question ambiguity and agent uncertainty.1:11:11–1:15:40 · Guest disagreement 2/10 Evaluating Abstract Reasoning (ARC-AGI) vs. Pragmatic Business Problems Josh and Jonathan critique abstract reasoning benchmarks like ARC-AGI as synthetic toys, prompting Swyx to defend standardized abstract testing as historically predictive.1:15:40–1:19:39 · Guest disagreement 4/10 The Limitations of 'Needle in a Haystack' and Long-Context Agent Scenarios Jonathan vehemently dismisses Needle in a Haystack as disconnected from real-world utility, while Josh explains why long context requires active relevance filtering.1:19:39–1:23:48 · Guest disagreement 1/10 Code Execution as the Universal Agent Interface vs. Hard-Coded Tool Calling Josh argues that general code execution is vastly superior to rigid tool-calling schemas, while Jonathan highlights the value of native SQL execution for structured enterprise data.1:23:48–1:26:23 · Guest disagreement 1/10 Knowledge Graphs, Data Schemas, and Bounding Agent Action Spaces Swyx prompts on whether knowledge graphs offer an enduring abstraction; Josh and Jonathan agree they fit specific ontologies but boundary control remains challenging.1:26:23–1:29:53 · Guest disagreement 1/10 Future Roadmaps: Databricks Science Sharing and the Abra-Kadabra Model Lineage Josh outlines Imbue's roadmap toward practical coding agents while Jonathan hints at future DBRX evolutionary models (Abra, Kadabra, and Alakazam) before closing.0:00–4:15 · The hosts pushing back 2/10 Introductions and Catching Up on the LLM Training Landscape Swyx introduces the guests, tracks Databricks' recent developments, and asks clarifying questions regarding Shutterstock licensing arrangements and data provenance.4:15–6:26 · The hosts pushing back 1/10 DBRX Architecture, Dinosaur Mascots, and the Hair Dye Bet Swyx recaps DBRX technical specs (132B total, 36B active, 12T tokens) while Jonathan jokes about team bets and dinosaur mascots.6:26–10:14 · The hosts pushing back 1/10 Imbue Open Sources the 70B Foundation Model Training Playbook Josh outlines Imbue's 70B playbook release across infra, evals, and CARBS hyperparameter optimization; Jonathan strongly validates the release as rare and essential.10:14–13:14 · The hosts pushing back 1/10 Hardware Realities and Bizarre GPU Cluster Failure Modes Jonathan and Josh share war stories about cluster failures, from stolen InfiniBand cables to memory correction latency and GPUs outputting incorrect arithmetic.13:14–16:41 · The hosts pushing back 2/10 Co-Designing a 4,096 H100 GPU Cluster with Voltage Park Swyx inquires about the 4,096 H100 cluster topology and partnership with Voltage Park, prompting Josh to explain their non-standard 3-tier networking choices.16:41–21:23 · The hosts pushing back 1/10 Automated Bare-Metal Provisioning and Low-Level Hardware Health Checks Josh walks through low-level automated bringup using Metal-as-a-Service, out-of-band management interfaces, and debugging directly with firmware vendors like Dell.21:23–24:35 · The hosts pushing back 3/10 Full-Stack Infrastructure: DIY Bare-Metal vs. Managed Cloud Platforms Swyx compares Imbue's bare-metal approach with Mosaic's cloud abstraction, leading Jonathan to explain why even cloud users end up dealing with low-level physical realities.24:35–28:08 · The hosts pushing back 2/10 Cluster Reliability Metrics, Triage Automation, and Node Churn Swyx presses on cluster failure rates; Josh corrects the uniform failure assumption and details automated dmesg anomaly detection on node reboots.28:08–32:34 · The hosts pushing back 1/10 Cluster Bringup Logistics: Small Teams, Cable Management, and Physical Assembly Josh details the physical bringup of 24,000 cable connections executed by a small core team in coordination with vendor technicians.32:34–35:53 · The hosts pushing back 1/10 Distributed Training Tools, S3 Local Mirrors, and Uber's Kraken Registry Josh discusses memory fragmentation jitter and highlights simple tooling choices like BitTorrent-based Kraken image distribution over heavy cloud filesystems.35:53–39:12 · The hosts pushing back 4/10 Contrasting Infrastructure Philosophies: Minimal Abstractions vs. Standardized APIs Swyx pushes back from an AWS perspective arguing Kubernetes standardizes infrastructure, but Josh counters that Kubernetes is mismatched with training job failures.39:12–44:41 · The hosts pushing back 1/10 Training Dynamics: MoE Network Bandwidth, HSDP, and Next-Gen Interconnects Jonathan explains how MoE training stresses network bandwidth, requiring HSDP, and compares Google's 3D torus TPU design with Nvidia's grouped NVLink topologies.44:41–50:16 · The hosts pushing back 1/10 Imbue's Specialized Focus on Code, Reasoning, and Text-Only Architectures Swyx probes Imbue's decision to focus purely on text and coding rather than vision; Josh explains how agent reasoning can outsource multimodal components.50:16–55:35 · The hosts pushing back 1/10 Evaluation Metric Precision: Multiple-Choice Perplexity and Data Mix Schedules Josh explains how CARBS optimizes hyperparameters across cost scales; Jonathan and Josh cordially disagree on whether dynamic data mix schedules provide real gains.55:35–57:55 · The hosts pushing back 2/10 The Mirage of Emergence: Log-Scale Perplexity vs. Metric Discontinuities Swyx asks whether emergence breaks CARBS scaling laws; Josh schools on how emergence is an artifact of non-linear accuracy metrics rather than log-scale loss.57:55–1:02:21 · The hosts pushing back 1/10 Cleaning Open-Source Benchmarks and Addressing Question Ambiguity Josh breaks down human annotation efforts to clean benchmark ambiguity, demonstrating that top models reach near total saturation on well-formed questions.1:02:21–1:06:52 · The hosts pushing back 1/10 Programmatic Benchmarks, Code Understanding, and Realistic Test Coverage Swyx questions the scope of code understanding; Josh and Jonathan discuss SWE-bench and the difficulty of evaluating nuanced code without leaking bad patches.1:06:52–1:11:11 · The hosts pushing back 1/10 Navigating Imperfect Evaluations and Modeling Uncertainty in AI Agents Jonathan reflects on living with imperfect evaluations when directing research, while Josh outlines Imbue's work on detecting question ambiguity and agent uncertainty.1:11:11–1:15:40 · The hosts pushing back 4/10 Evaluating Abstract Reasoning (ARC-AGI) vs. Pragmatic Business Problems Josh and Jonathan critique abstract reasoning benchmarks like ARC-AGI as synthetic toys, prompting Swyx to defend standardized abstract testing as historically predictive.1:15:40–1:19:39 · The hosts pushing back 1/10 The Limitations of 'Needle in a Haystack' and Long-Context Agent Scenarios Jonathan vehemently dismisses Needle in a Haystack as disconnected from real-world utility, while Josh explains why long context requires active relevance filtering.1:19:39–1:23:48 · The hosts pushing back 1/10 Code Execution as the Universal Agent Interface vs. Hard-Coded Tool Calling Josh argues that general code execution is vastly superior to rigid tool-calling schemas, while Jonathan highlights the value of native SQL execution for structured enterprise data.1:23:48–1:26:23 · The hosts pushing back 2/10 Knowledge Graphs, Data Schemas, and Bounding Agent Action Spaces Swyx prompts on whether knowledge graphs offer an enduring abstraction; Josh and Jonathan agree they fit specific ontologies but boundary control remains challenging.1:26:23–1:29:53 · The hosts pushing back 1/10 Future Roadmaps: Databricks Science Sharing and the Abra-Kadabra Model Lineage Josh outlines Imbue's roadmap toward practical coding agents while Jonathan hints at future DBRX evolutionary models (Abra, Kadabra, and Alakazam) before closing.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:15:48 Jonathan's dismissal of Needle in a Haystack

Jonathan reacts with exasperation ('Oh, for the love of God'), arguing the benchmark fails to test holistic context utilization and measures nothing realistic.

Hardest push from the hosts ▶ 1:15:16 Swyx defending abstract evaluation tests

Swyx pushes back against the guests' dismissal of abstract tests like ARC-AGI by pointing out humanity's successful reliance on SATs and IQ tests as capability predictors.

Biggest teaching moment ▶ 55:45 Josh explaining emergence as a metric mirage

Josh systematically disproves the premise of sudden model emergence by explaining that discontinuous accuracy jumps are simply log-linear perplexity improvements passing threshold bands.

The host holds their own ▶ 37:49 Swyx challenging bare-metal abstraction choices

Drawing on his AWS background, Swyx challenges Imbue's custom orchestration layer by asserting that Kubernetes provides essential vendor-agnostic portability for scaling.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions and Catching Up on the LLM Training Landscape 5312 Swyx introduces the guests, tracks Databricks' recent developments, and asks clarifying questions regarding Shutterstock licensing arrangements and data provenance.
DBRX Architecture, Dinosaur Mascots, and the Hair Dye Bet 6211 Swyx recaps DBRX technical specs (132B total, 36B active, 12T tokens) while Jonathan jokes about team bets and dinosaur mascots.
Imbue Open Sources the 70B Foundation Model Training Playbook 4611 Josh outlines Imbue's 70B playbook release across infra, evals, and CARBS hyperparameter optimization; Jonathan strongly validates the release as rare and essential.
Hardware Realities and Bizarre GPU Cluster Failure Modes 4611 Jonathan and Josh share war stories about cluster failures, from stolen InfiniBand cables to memory correction latency and GPUs outputting incorrect arithmetic.
Co-Designing a 4,096 H100 GPU Cluster with Voltage Park 5512 Swyx inquires about the 4,096 H100 cluster topology and partnership with Voltage Park, prompting Josh to explain their non-standard 3-tier networking choices.
Automated Bare-Metal Provisioning and Low-Level Hardware Health Checks 3711 Josh walks through low-level automated bringup using Metal-as-a-Service, out-of-band management interfaces, and debugging directly with firmware vendors like Dell.
Full-Stack Infrastructure: DIY Bare-Metal vs. Managed Cloud Platforms 5523 Swyx compares Imbue's bare-metal approach with Mosaic's cloud abstraction, leading Jonathan to explain why even cloud users end up dealing with low-level physical realities.
Cluster Reliability Metrics, Triage Automation, and Node Churn 5622 Swyx presses on cluster failure rates; Josh corrects the uniform failure assumption and details automated dmesg anomaly detection on node reboots.
Cluster Bringup Logistics: Small Teams, Cable Management, and Physical Assembly 4611 Josh details the physical bringup of 24,000 cable connections executed by a small core team in coordination with vendor technicians.
Distributed Training Tools, S3 Local Mirrors, and Uber's Kraken Registry 5611 Josh discusses memory fragmentation jitter and highlights simple tooling choices like BitTorrent-based Kraken image distribution over heavy cloud filesystems.
Contrasting Infrastructure Philosophies: Minimal Abstractions vs. Standardized APIs 6524 Swyx pushes back from an AWS perspective arguing Kubernetes standardizes infrastructure, but Josh counters that Kubernetes is mismatched with training job failures.
Training Dynamics: MoE Network Bandwidth, HSDP, and Next-Gen Interconnects 5711 Jonathan explains how MoE training stresses network bandwidth, requiring HSDP, and compares Google's 3D torus TPU design with Nvidia's grouped NVLink topologies.
Imbue's Specialized Focus on Code, Reasoning, and Text-Only Architectures 4511 Swyx probes Imbue's decision to focus purely on text and coding rather than vision; Josh explains how agent reasoning can outsource multimodal components.
Evaluation Metric Precision: Multiple-Choice Perplexity and Data Mix Schedules 4631 Josh explains how CARBS optimizes hyperparameters across cost scales; Jonathan and Josh cordially disagree on whether dynamic data mix schedules provide real gains.
The Mirage of Emergence: Log-Scale Perplexity vs. Metric Discontinuities 4822 Swyx asks whether emergence breaks CARBS scaling laws; Josh schools on how emergence is an artifact of non-linear accuracy metrics rather than log-scale loss.
Cleaning Open-Source Benchmarks and Addressing Question Ambiguity 4711 Josh breaks down human annotation efforts to clean benchmark ambiguity, demonstrating that top models reach near total saturation on well-formed questions.
Programmatic Benchmarks, Code Understanding, and Realistic Test Coverage 5611 Swyx questions the scope of code understanding; Josh and Jonathan discuss SWE-bench and the difficulty of evaluating nuanced code without leaking bad patches.
Navigating Imperfect Evaluations and Modeling Uncertainty in AI Agents 4711 Jonathan reflects on living with imperfect evaluations when directing research, while Josh outlines Imbue's work on detecting question ambiguity and agent uncertainty.
Evaluating Abstract Reasoning (ARC-AGI) vs. Pragmatic Business Problems 6624 Josh and Jonathan critique abstract reasoning benchmarks like ARC-AGI as synthetic toys, prompting Swyx to defend standardized abstract testing as historically predictive.
The Limitations of 'Needle in a Haystack' and Long-Context Agent Scenarios 4841 Jonathan vehemently dismisses Needle in a Haystack as disconnected from real-world utility, while Josh explains why long context requires active relevance filtering.
Code Execution as the Universal Agent Interface vs. Hard-Coded Tool Calling 5611 Josh argues that general code execution is vastly superior to rigid tool-calling schemas, while Jonathan highlights the value of native SQL execution for structured enterprise data.
Knowledge Graphs, Data Schemas, and Bounding Agent Action Spaces 5512 Swyx prompts on whether knowledge graphs offer an enduring abstraction; Josh and Jonathan agree they fit specific ontologies but boundary control remains challenging.
Future Roadmaps: Databricks Science Sharing and the Abra-Kadabra Model Lineage 4411 Josh outlines Imbue's roadmap toward practical coding agents while Jonathan hints at future DBRX evolutionary models (Abra, Kadabra, and Alakazam) before closing.

Statements from this episode (43)

Disclosure
Frankle: Databricks released text-to-image model with Shutterstock
“Is that we finally released our text image model which has been a year in the making through a collaboration directly with Shutterstock.”
Jonathan Frankle Jun 25, 2024 ▶ 1:32
Assertion Supported
Frankle: OpenAI, Google, Meta, and Apple have data deals with Shutterstock
“And you know, I, at least I've heard in the news, like opening, I Google, Meta Apple have all called Shutterstock and made those deals.”
Jonathan Frankle Jun 25, 2024 ▶ 3:05
Assertion Supported
Frankle: Databricks model is uniquely trained purely on Shutterstock data
“So a lot of models have had Shutterstock data incorporated into them, but this is the only model I know of so far where it was, you know, exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. You know, we …”
Jonathan Frankle Jun 25, 2024 ▶ 3:14
Disclosure
Albrecht: Imbue will not release model weights, but will open-source training tools
“We're not releasing the model. We're not releasing the weights, but we are releasing a bunch of different things that should make it easier for other people to make their own models.”
Josh Albrecht Jun 25, 2024 ▶ 6:28
Disclosure
Albrecht: Imbue is releasing a new reasoning benchmark and 11 cleaned evaluations
“We're releasing a whole bunch of different data there, a new benchmark about code, reasoning, understanding, as well as our own private versions of 11 different open source benchmarks. So things like PoolQ or ANLI, where we've gone through and kind of cleaned …”
Josh Albrecht Jun 25, 2024 ▶ 7:23
Disclosure
Imbue is releasing approximately 450,000 human evaluation judgments
“A final thing that we're releasing there is around 450,000 human judgments about ambiguity and question quality, which we used In the process of cleaning these evaluations”
Josh Albrecht Jun 25, 2024 ▶ 8:02
Insight
Frankle: Fault tolerance is missing from fundamental model training primitives
“Fault tolerance is still not really built into any of the fundamental primitives of training models. And so if something breaks, you have to go figure out what broke your job stops. You have to restart your job. It is a nightmare just to get to the point where…”
Jonathan Frankle Jun 25, 2024 ▶ 9:21
Assertion Not checkable as stated
Frankle: Porch pirates stole MosaicML InfiniBand cables twice before data center delivery
“Our InfiniBand cables getting stolen from the data center twice, like in boxes before they arrived at the data center, like, you know, porch pirate basically had stolen our InfiniBand cables back when those were hard to come by”
Jonathan Frankle Jun 25, 2024 ▶ 10:54
Assertion Not checkable as stated
Frankle: Most AI data centers are retrofitted, not built for high heat
“In data centers that for the most part were not built remotely for this kind of power or heat and have been retrofitted for this. Like failures happen on a good day with normal CPUs. And this is not a good day and not a normal CPU for the most part.”
Jonathan Frankle Jun 25, 2024 ▶ 12:41
Assertion Supported
Albrecht: 4K GPU clusters require 3-tier networking versus standard 1K 2-tier setups
“The normal, the like vanilla setup or, you know, these large clusters as vanilla as it can be is what's normally like a 127 node cluster. So closer to like 10, 24 GPUs instead of 4000. Here we have a larger cluster. As you start to get into the larger clusters…”
Josh Albrecht Jun 25, 2024 ▶ 13:50
Opinion
Albrecht: Training on AWS prevents diagnosing low-level hardware errors
“And if we're just using, you know, AWS or some other cloud provider, These errors are still going to be there, and you're gonna have no way to know and no way to debug this and no way to diagnose what's going wrong.”
Josh Albrecht Jun 25, 2024 ▶ 19:49
Insight
Frankle: Large-scale model training forces teams to debug the full infrastructure stack
“It's kind of impossible if you're doing training to not go all the way through the entire stack, regardless of what happens. Like somehow I'm still chatting with cloud providers about power contracts, even though the whole point of dealing with the cloud provi…”
Jonathan Frankle Jun 25, 2024 ▶ 23:07
Assertion Not checkable as stated
Albrecht: Imbue's GPU cluster failure rate is well below industry 3% benchmark
“The number that we've heard from other people is like they're having about three percent. I don't think we're experiencing failure rates that are that high. I think ours is actually quite a bit lower than that, probably because we've taken the time to like dig…”
Josh Albrecht Jun 25, 2024 ▶ 24:56
Disclosure
Albrecht: Imbue Manages Infrastructure with Three to Six Engineers
“Like our infrastructure team is like You know, it fluctuates from week to week, depending on like how many things are on fire and how much we need to build. But it's like between like three and six people, like it's small. It's not like some huge team of like …”
Josh Albrecht Jun 25, 2024 ▶ 28:28
Assertion Supported
Albrecht: 4,000-GPU Three-Tier Cluster Requires 12,000 Cables and 24,000 Plugs
“Like to bring up this cluster you know, with 4000 GPUs and three tier networking, networking architecture, you have 12,000 cables. So that's 24,000 things that need to be plugged in.”
Josh Albrecht Jun 25, 2024 ▶ 28:56
Insight
Albrecht: Unsynchronized garbage collection across distributed nodes steadily degrades cluster MFU
“Because you have hundreds of machines, they're doing garbage collection at slightly different times, and then they get slightly further apart, and slightly more and more jittered, until eventually they're all happening kind of at random times, and just, like, …”
Josh Albrecht Jun 25, 2024 ▶ 31:09
Disclosure
Albrecht: Imbue avoids Kubernetes to keep cluster infrastructure simple to debug
“Less layers of infrastructure, less layers of abstraction, make it a lot easier to work with. Like we don't use Kubernetes, for example, I would just directly launch these things and it's just been much easier to debug this way.”
Josh Albrecht Jun 25, 2024 ▶ 34:38
Disclosure
Frankle: Databricks runs across six or seven different cloud providers
“Think we're running on like six or seven different clouds right now.”
Jonathan Frankle Jun 25, 2024 ▶ 36:15
Insight
Frankle: Training MoE models with FSDP creates severe network bandwidth bottlenecks
“And those models are very demanding when it comes to network bandwidth, at least if you're training them in kind of FSTP zero three style. Where there's just a lot of parameters getting shuffled back and forth and your ratio of kind of compute to amount of dat…”
Jonathan Frankle Jun 25, 2024 ▶ 39:59
Opinion
Frankle: Networking was the hardest part of training DBRX at scale
“And so actually the networking part of DPRX was the single hardest thing. I think of the entire process, just get MOE training, working at scale across a big cluster.”
Jonathan Frankle Jun 25, 2024 ▶ 40:27
Assertion Supported
Frankle: Google TPUs provide much higher bandwidth-to-compute ratios
“TPUs have a very different network bandwidth to compute ratio. They have a lot more bandwidth just objectively and TPUs per chip tend to be a little bit less compute intensive and have a little bit less memory.”
Jonathan Frankle Jun 25, 2024 ▶ 41:17
Insight
Albrecht: Vision is not essential for most coding and reasoning agent tasks
“And actually we found that for most of the kind of like code writing and reasoning problems that we care about, the visual part isn't really a huge important part of it.”
Josh Albrecht Jun 25, 2024 ▶ 45:17
Insight
Albrecht: CARBS models compute cost per sample for hyperparameter search
“CARBS is, it's maybe a backronym, but it's for Cost Aware Pareto Region Bayesian Search... The point is that it's a cost aware hyperparameter tuner. So most hyperparameter tuners you kind of say, okay, here's this objective function. I want you to make this nu…”
Josh Albrecht Jun 25, 2024 ▶ 47:25
Insight
Albrecht: Cost-aware tuning reveals scaling laws for all hyperparameters
“So by doing that, we can see the scaling laws or not just, you know, the scaling laws from like the, you know, chinchilla paper, the scaling laws for all parameters. We can see how does the number of layers change with this? How does the You know, the learning…”
Josh Albrecht Jun 25, 2024 ▶ 48:50
Assertion Not checkable as stated
Albrecht: Dynamic data mix schedules during LLM training yield negligible gains
“We did some experiments and we've actually talked to a bunch of researchers who were doing work here as well and looking at kind of their experiments on this. And we were originally pretty hopeful because it sounds like something that should work and make sens…”
Josh Albrecht Jun 25, 2024 ▶ 52:36
Assertion Supported
Frankle: Dynamic data mixing during pre-training is effective for domain-specific models
“We've had some surprisingly good luck with this. We just released a paper on it. The details matter a lot and it really matters what you're trying to do with the model. But it's been quite effective for us depending on the setting. And certainly when we're thi…”
Jonathan Frankle Jun 25, 2024 ▶ 53:17
Insight
Albrecht: LLM emergence is an artifact of non-linear evaluation metrics
“This emergent behavior that you're seeing, Is not really emergent behavior, but is really a function of the evaluation metrics that we're using.”
Josh Albrecht Jun 25, 2024 ▶ 55:56
Insight
Frankle: Deep learning log scales can make trends look however you want
“Anything can look however you want it to look if you put it on a log scale to a certain extent. And log, we love our log scales and deep learning for various reasons. Everything looks very clean on a log scale until everything looks very flat on a log scale.”
Jonathan Frankle Jun 25, 2024 ▶ 57:27
Disclosure
Albrecht: Imbue reproduced 500-1,000 examples per dataset to stop eval contamination
“Let's just reproduce, you know, 500 to a thousand examples for every single one of these data sets ourselves and just make sure that this data is definitely not in the, you know, the training set. So we did that and then we're able to like now be confident abo…”
Josh Albrecht Jun 25, 2024 ▶ 1:00:11
Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Josh Albrecht Jun 25, 2024 ▶ 1:01:43
Assertion Supported
Albrecht: AgentBench paper's appendix examples are actually incorrect solutions
“Like we were looking at the agent bench paper, I think just last week for our paper club. And one of the things that we noticed is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not…”
Josh Albrecht Jun 25, 2024 ▶ 1:06:05
Insight
Frankle: Top AI scientists must tolerate broken infrastructure and imperfect evals
“Like the most successful scientists I see are the ones who are okay operating in a world where everything's going to be broken. And yet we can still cobble things together and make something interesting happen.”
Jonathan Frankle Jun 25, 2024 ▶ 1:07:46
Insight
Albrecht: Coding agents communicating uncertainty are far more useful than slightly more accurate ones
“I would much rather have a coding agent that will give me back a thing. And you know, it's actually the code doesn't work like 10% less of the time than some other model, but it will tell me a hundred percent of the time. When it got like when it's not sure, l…”
Josh Albrecht Jun 25, 2024 ▶ 1:10:42
Insight
Albrecht: Optimizing for competitive coding benchmarks does not create useful programmers
“Like, we do a lot of code generation, but we don't really do a lot on, like, code competition problems for the very, very hard ones, so that you can go very far down that route and make something like really good at those problems, but not actually that useful…”
Josh Albrecht Jun 25, 2024 ▶ 1:13:00
Assertion Not checkable as stated
Frankle: No Databricks enterprise customer asks for abstract reasoning AI
“I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems.”
Jonathan Frankle Jun 25, 2024 ▶ 1:13:34
Prediction Not checkable as stated
Frankle: Multimodal AI models will inevitably require massive context windows
“Once you get into multimodal land, you're just going to end up with giant context. It's kind of unavoidable.”
Jonathan Frankle Jun 25, 2024 ▶ 1:16:55
Insight
Frankle: Needle in a Haystack eval fails to measure holistic context usage
“I think the problems with needle in a haystack are well known. You know, it doesn't measure anything real. You're not even testing the model's ability to holistically use the context just to identify one part of the context. So you can do some wacky things to …”
Jonathan Frankle Jun 25, 2024 ▶ 1:17:50
Insight
Albrecht: Code execution expands agent capabilities far beyond hard-coded tool calling
“Instead of worrying about like weird hard coded agents using tools, Like let's just make them able to actually write code robustly and make that code work and be able to debug that code, know if that code is safe to run, like get really good at the like code w…”
Josh Albrecht Jun 25, 2024 ▶ 1:20:40
Assertion Not checkable as stated
Frankle: PDF parsing remains an unsolved problem in 2024
“PDF parsing is still an unsolved problem, even in 20, 24.”
Jonathan Frankle Jun 25, 2024 ▶ 1:21:43
Assertion Not checkable as stated
Frankle: Text-to-SQL is one of the most impactful LLM use cases for enterprise
“Like it's, you know, text to SQL is still, or like having a model be able to make SQL calls in the backend is actually like one of the single most useful things for my customers. It sounds really boring. Models are really good at it and it moves the needle day…”
Jonathan Frankle Jun 25, 2024 ▶ 1:22:12
Insight
Albrecht: Messy real-world data limits knowledge graphs to niche problems
“But I think in the real world, it gets a lot messier than like knowledge graph style of things where it's like, well, is there a relationship between these two nodes? Like, ah, I don't know. Like is, are these two separate nodes? Like those kinds of messy bord…”
Josh Albrecht Jun 25, 2024 ▶ 1:24:50
Disclosure
Frankle: Releasing Open-Source Models Is Not Databricks' Core Bread and Butter
“Releasing models open source is not our day-to-day bread and butter. It's kind of a fun reward that we get to do sometimes when we have something really cool to share and a little bit of time and spare GPUs in our hands. But for the most part, everything is go…”
Jonathan Frankle Jun 25, 2024 ▶ 1:28:40
Disclosure
Databricks codenamed DBRX Kadabra, teasing a third Alakazam model evolution
“The DBRX small model that we still haven't released yet was called Abra. DBRX was called Kadabra and, you know, there's a third Pokemon in that evolution and that's all I'll say for now.”
Jonathan Frankle Jun 25, 2024 ▶ 1:29:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.