Oct 1, 2024 · 1h 16m · a16z

The Quest for Community-Trained Open Source AI Models

Jeff Schmidt · 40m spoken Anjney Midha · 11m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, Nous Research co-founders Jeff Schmidt and Bowen Peng discuss Distro, a revolutionary decentralized training protocol that reduces GPU bandwidth requirements by up to 1000x. By decoupling frontier AI training from centralized high-speed datacenter interconnects, Nous enables globally distributed consumer and edge hardware to collaboratively train state-of-the-art models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 2.3 Guest teaching 2.4 Guest disagreement 0.1 The host pushing back 0.2
05100:0020:0040:001:00:000:27–4:02 · The host as informed peer 0/10 Title Sequence and Legal Disclaimer Anjney asks a standard opening question about Nous Research's roadmap, allowing Jeff and Bowen to explain their mission of open-source AI and individualistic research.4:02–6:41 · The host as informed peer 0/10 Founder Backgrounds: Automotive, Crypto, and AI Anjney asks for founder backgrounds, leading to Jeff detailing his transition from automotive autonomous driving and Ethereum smart contracts into open-source AI.6:41–9:26 · The host as informed peer 0/10 Bowen's ML Origins and Meeting via Reddit Bowen shares his background learning deep learning under Aaron Courville at Mila and how a local llama Reddit post led to cold-emailing Jeff.9:26–14:02 · The host as informed peer 1/10 Key Projects: Hermes Models and YaRN Context Scaling Jeff details Nous Research's key early projects, including the Hermes model series for customizable personas and the widely adopted YaRN context scaling paper.14:02–16:21 · The host as informed peer 1/10 The Centralized Compute Bottleneck in AI Training Jeff describes the centralized compute bottleneck in modern AI training and explains why distributed internet-scale training was previously considered intractable.16:21–19:00 · The host as informed peer 2/10 Selection Criteria and Synthetic Data Breakthroughs Bowen and Jeff describe their research selection criteria, focusing on fundamental mathematical leverage points like synthetic data generation.19:00–24:40 · The host as informed peer 3/10 Introducing Distro: Decentralized Internet Training Anjney synthesizes the core premise of Distro, prompting Bowen and Jeff to clarify that model performance remains equivalent despite requiring 1000x less bandwidth.24:40–29:26 · The host as informed peer 4/10 Democratizing Frontier AI Model Development Anjney accurately identifies how Distro decouples model performance scaling from physical interconnect requirements, which Bowen directly validates.29:26–33:51 · The host as informed peer 3/10 Distro Empirical Results and 1000x Bandwidth Reduction Anjney highlights the published 857x bandwidth reduction metrics, prompting Bowen and Jeff to elaborate on conservative estimates and benchmark metrics like perplexity.33:51–36:14 · The host as informed peer 4/10 Red-Teaming Distro: Addressing Objections and Scaling Anjney intentionally prompts the guests to red-team their own work, leading Bowen and Jeff to openly analyze baseline choices and model scaling concerns.36:14–39:17 · The host as informed peer 3/10 Rigorous Baseline Verification with OLMo Jeff details how Nous threw out their initial setup and re-ran baseline verification using Allen AI's open OLMo framework to prove reproducibility.39:17–43:22 · The host as informed peer 4/10 Open Science and Global Participation Mindset Anjney questions whether Distro's success poses a threat to Nvidia, prompting Jeff and Bowen to explain hardware architecture nuances and VRAM vs interconnect tradeoffs.43:22–48:19 · The host as informed peer 3/10 Harnessing Consumer GPUs and Fault-Tolerant Code Anjney references historical distributed projects like Folding@home and presses on whether high-end H100 GPUs remain strictly necessary for training.48:19–53:00 · The host as informed peer 2/10 Distro Release Roadmap and Tooling Ecosystem Anjney inquires about the practical roadmap and tooling required to transition Distro from academic research to accessible community infrastructure.53:00–56:00 · The host as informed peer 3/10 Regulatory Risk and Community Frontier Timelines Anjney asks about community timeline projections if corporate labs stop open-sourcing frontier models due to regulatory pressures.56:00–1:04:33 · The host as informed peer 3/10 Technical Deep Dive: Bounded Search vs. AllReduce Bowen and Jeff provide a deep technical breakdown of how Distro replaces traditional AllReduce averaging with a bounded, multi-model search space.1:04:33–1:10:41 · The host as informed peer 4/10 Network Topologies and Asynchronous Node Clusters Anjney probes network topology and validator governance dynamics, prompting Jeff to describe multi-tier continental cluster configurations.1:10:41–1:16:12 · The host as informed peer 3/10 Zeroth-Order Optimization, ASICs, and Continuous Learning Bowen and Jeff reveal their early experiments with zeroth-order optimization and explain how forward-pass training could enable ASICs and mobile devices to train continuously.1:16:12–1:16:33 · The host as informed peer 0/10 Conclusion and SETI@Home Vision for AI Brief episode sign-off emphasizing the grand vision of community-driven SETI@home style training for open AI.0:27–4:02 · Guest teaching 1/10 Title Sequence and Legal Disclaimer Anjney asks a standard opening question about Nous Research's roadmap, allowing Jeff and Bowen to explain their mission of open-source AI and individualistic research.4:02–6:41 · Guest teaching 1/10 Founder Backgrounds: Automotive, Crypto, and AI Anjney asks for founder backgrounds, leading to Jeff detailing his transition from automotive autonomous driving and Ethereum smart contracts into open-source AI.6:41–9:26 · Guest teaching 1/10 Bowen's ML Origins and Meeting via Reddit Bowen shares his background learning deep learning under Aaron Courville at Mila and how a local llama Reddit post led to cold-emailing Jeff.9:26–14:02 · Guest teaching 3/10 Key Projects: Hermes Models and YaRN Context Scaling Jeff details Nous Research's key early projects, including the Hermes model series for customizable personas and the widely adopted YaRN context scaling paper.14:02–16:21 · Guest teaching 3/10 The Centralized Compute Bottleneck in AI Training Jeff describes the centralized compute bottleneck in modern AI training and explains why distributed internet-scale training was previously considered intractable.16:21–19:00 · Guest teaching 3/10 Selection Criteria and Synthetic Data Breakthroughs Bowen and Jeff describe their research selection criteria, focusing on fundamental mathematical leverage points like synthetic data generation.19:00–24:40 · Guest teaching 2/10 Introducing Distro: Decentralized Internet Training Anjney synthesizes the core premise of Distro, prompting Bowen and Jeff to clarify that model performance remains equivalent despite requiring 1000x less bandwidth.24:40–29:26 · Guest teaching 2/10 Democratizing Frontier AI Model Development Anjney accurately identifies how Distro decouples model performance scaling from physical interconnect requirements, which Bowen directly validates.29:26–33:51 · Guest teaching 3/10 Distro Empirical Results and 1000x Bandwidth Reduction Anjney highlights the published 857x bandwidth reduction metrics, prompting Bowen and Jeff to elaborate on conservative estimates and benchmark metrics like perplexity.33:51–36:14 · Guest teaching 2/10 Red-Teaming Distro: Addressing Objections and Scaling Anjney intentionally prompts the guests to red-team their own work, leading Bowen and Jeff to openly analyze baseline choices and model scaling concerns.36:14–39:17 · Guest teaching 3/10 Rigorous Baseline Verification with OLMo Jeff details how Nous threw out their initial setup and re-ran baseline verification using Allen AI's open OLMo framework to prove reproducibility.39:17–43:22 · Guest teaching 3/10 Open Science and Global Participation Mindset Anjney questions whether Distro's success poses a threat to Nvidia, prompting Jeff and Bowen to explain hardware architecture nuances and VRAM vs interconnect tradeoffs.43:22–48:19 · Guest teaching 3/10 Harnessing Consumer GPUs and Fault-Tolerant Code Anjney references historical distributed projects like Folding@home and presses on whether high-end H100 GPUs remain strictly necessary for training.48:19–53:00 · Guest teaching 2/10 Distro Release Roadmap and Tooling Ecosystem Anjney inquires about the practical roadmap and tooling required to transition Distro from academic research to accessible community infrastructure.53:00–56:00 · Guest teaching 3/10 Regulatory Risk and Community Frontier Timelines Anjney asks about community timeline projections if corporate labs stop open-sourcing frontier models due to regulatory pressures.56:00–1:04:33 · Guest teaching 4/10 Technical Deep Dive: Bounded Search vs. AllReduce Bowen and Jeff provide a deep technical breakdown of how Distro replaces traditional AllReduce averaging with a bounded, multi-model search space.1:04:33–1:10:41 · Guest teaching 3/10 Network Topologies and Asynchronous Node Clusters Anjney probes network topology and validator governance dynamics, prompting Jeff to describe multi-tier continental cluster configurations.1:10:41–1:16:12 · Guest teaching 4/10 Zeroth-Order Optimization, ASICs, and Continuous Learning Bowen and Jeff reveal their early experiments with zeroth-order optimization and explain how forward-pass training could enable ASICs and mobile devices to train continuously.1:16:12–1:16:33 · Guest teaching 0/10 Conclusion and SETI@Home Vision for AI Brief episode sign-off emphasizing the grand vision of community-driven SETI@home style training for open AI.0:27–4:02 · Guest disagreement 0/10 Title Sequence and Legal Disclaimer Anjney asks a standard opening question about Nous Research's roadmap, allowing Jeff and Bowen to explain their mission of open-source AI and individualistic research.4:02–6:41 · Guest disagreement 0/10 Founder Backgrounds: Automotive, Crypto, and AI Anjney asks for founder backgrounds, leading to Jeff detailing his transition from automotive autonomous driving and Ethereum smart contracts into open-source AI.6:41–9:26 · Guest disagreement 0/10 Bowen's ML Origins and Meeting via Reddit Bowen shares his background learning deep learning under Aaron Courville at Mila and how a local llama Reddit post led to cold-emailing Jeff.9:26–14:02 · Guest disagreement 0/10 Key Projects: Hermes Models and YaRN Context Scaling Jeff details Nous Research's key early projects, including the Hermes model series for customizable personas and the widely adopted YaRN context scaling paper.14:02–16:21 · Guest disagreement 0/10 The Centralized Compute Bottleneck in AI Training Jeff describes the centralized compute bottleneck in modern AI training and explains why distributed internet-scale training was previously considered intractable.16:21–19:00 · Guest disagreement 0/10 Selection Criteria and Synthetic Data Breakthroughs Bowen and Jeff describe their research selection criteria, focusing on fundamental mathematical leverage points like synthetic data generation.19:00–24:40 · Guest disagreement 0/10 Introducing Distro: Decentralized Internet Training Anjney synthesizes the core premise of Distro, prompting Bowen and Jeff to clarify that model performance remains equivalent despite requiring 1000x less bandwidth.24:40–29:26 · Guest disagreement 0/10 Democratizing Frontier AI Model Development Anjney accurately identifies how Distro decouples model performance scaling from physical interconnect requirements, which Bowen directly validates.29:26–33:51 · Guest disagreement 0/10 Distro Empirical Results and 1000x Bandwidth Reduction Anjney highlights the published 857x bandwidth reduction metrics, prompting Bowen and Jeff to elaborate on conservative estimates and benchmark metrics like perplexity.33:51–36:14 · Guest disagreement 1/10 Red-Teaming Distro: Addressing Objections and Scaling Anjney intentionally prompts the guests to red-team their own work, leading Bowen and Jeff to openly analyze baseline choices and model scaling concerns.36:14–39:17 · Guest disagreement 0/10 Rigorous Baseline Verification with OLMo Jeff details how Nous threw out their initial setup and re-ran baseline verification using Allen AI's open OLMo framework to prove reproducibility.39:17–43:22 · Guest disagreement 0/10 Open Science and Global Participation Mindset Anjney questions whether Distro's success poses a threat to Nvidia, prompting Jeff and Bowen to explain hardware architecture nuances and VRAM vs interconnect tradeoffs.43:22–48:19 · Guest disagreement 0/10 Harnessing Consumer GPUs and Fault-Tolerant Code Anjney references historical distributed projects like Folding@home and presses on whether high-end H100 GPUs remain strictly necessary for training.48:19–53:00 · Guest disagreement 0/10 Distro Release Roadmap and Tooling Ecosystem Anjney inquires about the practical roadmap and tooling required to transition Distro from academic research to accessible community infrastructure.53:00–56:00 · Guest disagreement 0/10 Regulatory Risk and Community Frontier Timelines Anjney asks about community timeline projections if corporate labs stop open-sourcing frontier models due to regulatory pressures.56:00–1:04:33 · Guest disagreement 0/10 Technical Deep Dive: Bounded Search vs. AllReduce Bowen and Jeff provide a deep technical breakdown of how Distro replaces traditional AllReduce averaging with a bounded, multi-model search space.1:04:33–1:10:41 · Guest disagreement 0/10 Network Topologies and Asynchronous Node Clusters Anjney probes network topology and validator governance dynamics, prompting Jeff to describe multi-tier continental cluster configurations.1:10:41–1:16:12 · Guest disagreement 0/10 Zeroth-Order Optimization, ASICs, and Continuous Learning Bowen and Jeff reveal their early experiments with zeroth-order optimization and explain how forward-pass training could enable ASICs and mobile devices to train continuously.1:16:12–1:16:33 · Guest disagreement 0/10 Conclusion and SETI@Home Vision for AI Brief episode sign-off emphasizing the grand vision of community-driven SETI@home style training for open AI.0:27–4:02 · The host pushing back 0/10 Title Sequence and Legal Disclaimer Anjney asks a standard opening question about Nous Research's roadmap, allowing Jeff and Bowen to explain their mission of open-source AI and individualistic research.4:02–6:41 · The host pushing back 0/10 Founder Backgrounds: Automotive, Crypto, and AI Anjney asks for founder backgrounds, leading to Jeff detailing his transition from automotive autonomous driving and Ethereum smart contracts into open-source AI.6:41–9:26 · The host pushing back 0/10 Bowen's ML Origins and Meeting via Reddit Bowen shares his background learning deep learning under Aaron Courville at Mila and how a local llama Reddit post led to cold-emailing Jeff.9:26–14:02 · The host pushing back 0/10 Key Projects: Hermes Models and YaRN Context Scaling Jeff details Nous Research's key early projects, including the Hermes model series for customizable personas and the widely adopted YaRN context scaling paper.14:02–16:21 · The host pushing back 0/10 The Centralized Compute Bottleneck in AI Training Jeff describes the centralized compute bottleneck in modern AI training and explains why distributed internet-scale training was previously considered intractable.16:21–19:00 · The host pushing back 0/10 Selection Criteria and Synthetic Data Breakthroughs Bowen and Jeff describe their research selection criteria, focusing on fundamental mathematical leverage points like synthetic data generation.19:00–24:40 · The host pushing back 0/10 Introducing Distro: Decentralized Internet Training Anjney synthesizes the core premise of Distro, prompting Bowen and Jeff to clarify that model performance remains equivalent despite requiring 1000x less bandwidth.24:40–29:26 · The host pushing back 0/10 Democratizing Frontier AI Model Development Anjney accurately identifies how Distro decouples model performance scaling from physical interconnect requirements, which Bowen directly validates.29:26–33:51 · The host pushing back 0/10 Distro Empirical Results and 1000x Bandwidth Reduction Anjney highlights the published 857x bandwidth reduction metrics, prompting Bowen and Jeff to elaborate on conservative estimates and benchmark metrics like perplexity.33:51–36:14 · The host pushing back 3/10 Red-Teaming Distro: Addressing Objections and Scaling Anjney intentionally prompts the guests to red-team their own work, leading Bowen and Jeff to openly analyze baseline choices and model scaling concerns.36:14–39:17 · The host pushing back 0/10 Rigorous Baseline Verification with OLMo Jeff details how Nous threw out their initial setup and re-ran baseline verification using Allen AI's open OLMo framework to prove reproducibility.39:17–43:22 · The host pushing back 1/10 Open Science and Global Participation Mindset Anjney questions whether Distro's success poses a threat to Nvidia, prompting Jeff and Bowen to explain hardware architecture nuances and VRAM vs interconnect tradeoffs.43:22–48:19 · The host pushing back 0/10 Harnessing Consumer GPUs and Fault-Tolerant Code Anjney references historical distributed projects like Folding@home and presses on whether high-end H100 GPUs remain strictly necessary for training.48:19–53:00 · The host pushing back 0/10 Distro Release Roadmap and Tooling Ecosystem Anjney inquires about the practical roadmap and tooling required to transition Distro from academic research to accessible community infrastructure.53:00–56:00 · The host pushing back 0/10 Regulatory Risk and Community Frontier Timelines Anjney asks about community timeline projections if corporate labs stop open-sourcing frontier models due to regulatory pressures.56:00–1:04:33 · The host pushing back 0/10 Technical Deep Dive: Bounded Search vs. AllReduce Bowen and Jeff provide a deep technical breakdown of how Distro replaces traditional AllReduce averaging with a bounded, multi-model search space.1:04:33–1:10:41 · The host pushing back 0/10 Network Topologies and Asynchronous Node Clusters Anjney probes network topology and validator governance dynamics, prompting Jeff to describe multi-tier continental cluster configurations.1:10:41–1:16:12 · The host pushing back 0/10 Zeroth-Order Optimization, ASICs, and Continuous Learning Bowen and Jeff reveal their early experiments with zeroth-order optimization and explain how forward-pass training could enable ASICs and mobile devices to train continuously.1:16:12–1:16:33 · The host pushing back 0/10 Conclusion and SETI@Home Vision for AI Brief episode sign-off emphasizing the grand vision of community-driven SETI@home style training for open AI.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%54:00 · the host 0% · guest 100%54:00 · the host 0% · guest 100%57:00 · the host 0% · guest 100%57:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%1:03:00 · the host 0% · guest 100%1:03:00 · the host 0% · guest 100%1:06:00 · the host 0% · guest 100%1:06:00 · the host 0% · guest 100%1:09:00 · the host 0% · guest 100%1:09:00 · the host 0% · guest 100%1:12:00 · the host 0% · guest 100%1:12:00 · the host 0% · guest 100%1:15:00 · the host 0% · guest 100%1:15:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 34:01 Guests self-red-team baseline and scaling flaws

Bowen and Jeff rigorously critique their own findings, acknowledging that incorrect baseline selection or small-scale quirks could render their results invalid.

Hardest push from the host ▶ 33:51 Host demands red-teaming of Distro claims

Anjney refuses to uncritically accept the published metrics and explicitly demands that the guests roleplay as skeptical disbelievers.

Biggest teaching moment ▶ 1:00:02 Reframing gradient synchronization as bounded multi-model search

Jeff reframes standard AI training assumptions by explaining that nodes do not need to synchronize back to a single model via AllReduce, but can instead search independently within a bounded loss landscape.

The host holds their own ▶ 25:41 Host articulates the interconnect decoupling insight

Anjney demonstrates technical mastery by identifying that Distro decouples model performance scaling from physical interconnect requirements, earning immediate validation from Bowen.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Title Sequence and Legal Disclaimer 0100 Anjney asks a standard opening question about Nous Research's roadmap, allowing Jeff and Bowen to explain their mission of open-source AI and individualistic research.
Founder Backgrounds: Automotive, Crypto, and AI 0100 Anjney asks for founder backgrounds, leading to Jeff detailing his transition from automotive autonomous driving and Ethereum smart contracts into open-source AI.
Bowen's ML Origins and Meeting via Reddit 0100 Bowen shares his background learning deep learning under Aaron Courville at Mila and how a local llama Reddit post led to cold-emailing Jeff.
Key Projects: Hermes Models and YaRN Context Scaling 1300 Jeff details Nous Research's key early projects, including the Hermes model series for customizable personas and the widely adopted YaRN context scaling paper.
The Centralized Compute Bottleneck in AI Training 1300 Jeff describes the centralized compute bottleneck in modern AI training and explains why distributed internet-scale training was previously considered intractable.
Selection Criteria and Synthetic Data Breakthroughs 2300 Bowen and Jeff describe their research selection criteria, focusing on fundamental mathematical leverage points like synthetic data generation.
Introducing Distro: Decentralized Internet Training 3200 Anjney synthesizes the core premise of Distro, prompting Bowen and Jeff to clarify that model performance remains equivalent despite requiring 1000x less bandwidth.
Democratizing Frontier AI Model Development 4200 Anjney accurately identifies how Distro decouples model performance scaling from physical interconnect requirements, which Bowen directly validates.
Distro Empirical Results and 1000x Bandwidth Reduction 3300 Anjney highlights the published 857x bandwidth reduction metrics, prompting Bowen and Jeff to elaborate on conservative estimates and benchmark metrics like perplexity.
Red-Teaming Distro: Addressing Objections and Scaling 4213 Anjney intentionally prompts the guests to red-team their own work, leading Bowen and Jeff to openly analyze baseline choices and model scaling concerns.
Rigorous Baseline Verification with OLMo 3300 Jeff details how Nous threw out their initial setup and re-ran baseline verification using Allen AI's open OLMo framework to prove reproducibility.
Open Science and Global Participation Mindset 4301 Anjney questions whether Distro's success poses a threat to Nvidia, prompting Jeff and Bowen to explain hardware architecture nuances and VRAM vs interconnect tradeoffs.
Harnessing Consumer GPUs and Fault-Tolerant Code 3300 Anjney references historical distributed projects like Folding@home and presses on whether high-end H100 GPUs remain strictly necessary for training.
Distro Release Roadmap and Tooling Ecosystem 2200 Anjney inquires about the practical roadmap and tooling required to transition Distro from academic research to accessible community infrastructure.
Regulatory Risk and Community Frontier Timelines 3300 Anjney asks about community timeline projections if corporate labs stop open-sourcing frontier models due to regulatory pressures.
Technical Deep Dive: Bounded Search vs. AllReduce 3400 Bowen and Jeff provide a deep technical breakdown of how Distro replaces traditional AllReduce averaging with a bounded, multi-model search space.
Network Topologies and Asynchronous Node Clusters 4300 Anjney probes network topology and validator governance dynamics, prompting Jeff to describe multi-tier continental cluster configurations.
Zeroth-Order Optimization, ASICs, and Continuous Learning 3400 Bowen and Jeff reveal their early experiments with zeroth-order optimization and explain how forward-pass training could enable ASICs and mobile devices to train continuously.
Conclusion and SETI@Home Vision for AI 0000 Brief episode sign-off emphasizing the grand vision of community-driven SETI@home style training for open AI.

Statements from this episode (25)

Insight
Jeff Schmidt: AI Research Offers Unprecedented Green Field for Groundbreaking Discoveries
“We're lucky right now because the state of AI as a science is incredibly new, and it truly has a wide green field. And unlike lots of other areas of science where if you look at something and you think, why hasn't somebody done X, Y, and Z? Unfortunately, ofte…”
Jeff Schmidt Oct 1, 2024 ▶ 2:50
Assertion Not checkable as stated
Schmidt: ChatGPT, Llama, and DeepSeek use Nous Research's YaRN context extension
“Bone here is the lead author of a method we developed called YARN, which is a context window extension method that we released and did the research on. It is now used by every, every model you use nowadays, everything, everything Chachipiti, Lama, DeepSeq, all…”
Jeff Schmidt Oct 1, 2024 ▶ 12:18
Assertion Supported
Schmidt: Current AI training requires all GPUs in the same datacenter
“When it comes to training models, the current paradigm for training models requires that all of the GPUs that train the model, these, you know, these computers that do the training, they all have to be like in the same room.”
Jeff Schmidt Oct 1, 2024 ▶ 14:23
Insight
Schmidt: AI industry relies on outdated 1990s architecture assumptions ripe for disruption
“It turns out that most assumptions in the AI space right now are a product of that's just how things had been done when there was not nearly as much energy and attention to it. So someone made an assumption maybe in the early nineties that everyone just kind o…”
Jeff Schmidt Oct 1, 2024 ▶ 15:48
Assertion Not checkable as stated
Jeff Schmidt: Hermes pioneered synthetic data training before it was standard
“So Hermes was very early to the idea that you could have synthetic data, which is that you could actually make, you could make a better model by taking an AI model, having it generate words and text, and then training a new a model on that output. This is now …”
Jeff Schmidt Oct 1, 2024 ▶ 17:58
Insight
Schmidt: Relying solely on fine-tuning is an existential threat to open-source AI
“Because that for us, if we're just fine tuning models. Right. That's like an existential threat, right? That's like an actual existential threat because the closed providers will continue to get better and we would be like dead in the water in a lot of sense.”
Jeff Schmidt Oct 1, 2024 ▶ 20:04
Assertion Supported
Schmidt: Elon Musk's xAI has acquired 100,000 NVIDIA H100 GPUs
“I think Elon's got a hundred, a 100,000 H 100 now.”
Jeff Schmidt Oct 1, 2024 ▶ 20:28
Assertion Not checkable as stated
Schmidt: Fewer than ten organizations worldwide can train Llama-scale AI models
“Yeah, I mean, I would, it would probably be in the number of ones on my hand and it probably wouldn't use all my fingers, you know. Yeah, I mean, you basically have, OpenA, Anthropic, Meta, X, Google, and then you have a few Mistral, and then Deep Seek and a c…”
Jeff Schmidt Oct 1, 2024 ▶ 25:00
Insight
Schmidt: Sharing only key signals in distributed training yields equivalent model learning
“We know that like what we, what needs to be communicated between these things, the two, the different nodes are just these few key pieces of information. And that is necessary. That is a necessary condition or rather a sufficient condition To get the equivalen…”
Jeff Schmidt Oct 1, 2024 ▶ 32:42
Assertion Not checkable as stated
Nous Research has trained DisTrO models up to 7 billion parameters
“We've gone up through seven B now, like seven B models.”
Jeff Schmidt Oct 1, 2024 ▶ 35:28
Assertion Not checkable as stated
DisTrO's performance advantage over AdamW widens as models scale up
“What we have seen empirically is that as we make it bigger, the differential between distro and MW actually gets wider.”
Jeff Schmidt Oct 1, 2024 ▶ 35:35
Assertion Supported
Nous Research replicates DisTrO training results using Allen AI's OLMo framework
“And we've re-implemented now a third time in their framework, and we're able to reproduce their training run exactly, and then did it again with Distro, got the exact same results we got with Natron and stuff.”
Jeff Schmidt Oct 1, 2024 ▶ 37:50
Insight
Midha: Uncertainty drives AI labs from open source to closed source
“There's so much uncertainty about who wins and who loses, et cetera, that the natural tendency, I, I've seen a natural tendency for people to go to start sharing less openly with the community when they have a breakthrough or results. You know, things go close…”
Anjney Midha Oct 1, 2024 ▶ 39:20
Prediction Not checkable as stated
Schmidt: Decentralized training will force NVIDIA to redesign chips around VRAM ratios
“What might happen sooner would be a redesign of the types of chips that NVIDIA or someone would make. Okay, under this model, we can dedicate more VRAM versus, there's like this question of how much VRAM versus how much processing power is on a die, and that, …”
Jeff Schmidt Oct 1, 2024 ▶ 42:34
Assertion Partly supported
Schmidt: Compute chips in NVIDIA's RTX 4090 and H100 are almost identical
“I think people don't actually realize that like a forty-ninety and like an H 100 are in a lot of ways the same card. For the non-gamers in the room, explain the forty-ninety. The chip that's inside of them is almost identical. The chip, the actual compute chip…”
Jeff Schmidt Oct 1, 2024 ▶ 44:54
Prediction Not checkable as stated
Schmidt: Consumer gaming GPUs will become the sweet spot for distributed training
“Because you're able to distribute it so wide, I think the gaming GPU angle is really going to be like the sweet spot. If, as long as there's continued to be sort of like higher end gaming GPUs, and those are on comparison with the high end training GPUs, even …”
Jeff Schmidt Oct 1, 2024 ▶ 45:39
Disclosure
Schmidt: Nous is building fault-tolerant training code for heterogeneous devices
“We're making sure that also like the code we're writing to help to actually do this training is agnostic to the hardware and is able to communicate and operate. You can have an Apple device and an NVIDIA device training together. And this is actually the, for …”
Jeff Schmidt Oct 1, 2024 ▶ 46:33
Prediction Didn’t hold up
Schmidt: Decentralized training of 400B parameter AI models is solvable by 2025
“I think it still is, it would still be, you know, like a next year sort of environment thing that we would have to do. There are some scaling problems, or not scaling problems, but technical things about how you shard the model, because at that point you get t…”
Jeff Schmidt Oct 1, 2024 ▶ 54:41
Assertion Not checkable as stated
Jeff Schmidt: Open-source AI lags closed AI providers by 1 to 1.5 years
“It seems that we're in the open source space. We're always like a year playing catch up, like a year, it's like a year and a half, a year and a half behind like the closed providers.”
Jeff Schmidt Oct 1, 2024 ▶ 55:45
Insight
DisTrO trains multiple models in a bounded search space instead of full synchronization
“So with distro, what we found is that rather than bringing everyone back home and averaging it back together, what you want to do is give each of those little nodes that are searching for the lowest point in the lost landscape, the freedom to move around. And …”
Jeff Schmidt Oct 1, 2024 ▶ 1:01:22
Insight
Schmidt: Synchronized AI training stems from PyTorch convenience abstractions, not optimal convergence
“But the one monolithic thing was actually just like a technical bot. Like it was from the fact that like we had PyTorch and then they like, you know, or like at Karis or any of the other ones. And they're like, Well, if you want, you can train on multiple GPUs…”
Jeff Schmidt Oct 1, 2024 ▶ 1:04:01
Insight
Peng: Starting distributed AI optimization with fully asynchronous nodes sacrifices convergence
“A lot of optimization algorithms start, like, with everything being, like, asynchronous. Right. But that's, like, really hard to get because you sacrifice a lot. You sacrifice the speed. You sacrifice the convergence, right? Like, actual efficiency of the algo…”
Bowen Peng Oct 1, 2024 ▶ 1:07:21
Prediction Not checkable as stated
Schmidt: DisTrO will eliminate data centers' InfiniBand dependency before edge AI dominates
“I think there's, you know, immediately coming out. What you'll see is the ability for even centralized actors who might have multiple data centers to now, like just use them in a more efficient way. Like just have N equals two, you know, like anything. And eac…”
Jeff Schmidt Oct 1, 2024 ▶ 1:09:20
Assertion Not checkable as stated
Schmidt: Zeroth-Order Optimization Requires 1,000x More Computation Than Backpropagation
“What we discovered is that backprop is still being like, you really do still need to be doing back propagation to find the optimal point of the loss. And it's just like zeroth order is like, what, like a thousand, like it worked, but it was like, you needed li…”
Jeff Schmidt Oct 1, 2024 ▶ 1:11:36
Assertion Not checkable as stated
Schmidt: 1-Bit Model Architecture BitNet Eliminates Multiplication Operations
“The amazing thing about this method called BitNet is that because all of the weights are either one zero or one, The multiplication disappears and it becomes just addition.”
Jeff Schmidt Oct 1, 2024 ▶ 1:14:07
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.