Aug 5, 2026 · 46m · a16z

How Open Source Became AI's Backbone | Inferact with a16z

Simon Moe · 23m spoken Matt Bornstein · 12m spoken Elena Burger · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Show, vLLM lead maintainer Simon Moe and a16z General Partner Matt Bornstein discuss the evolution of open-source AI infrastructure, the engineering challenges of LLM inference, and why customizable open-weight models are rapidly closing the capability gap with closed enterprise APIs.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 5.9 Guest teaching 5.5 Guest disagreement 1.4 The host pushing back 1.4
05100:0015:0030:0045:001:18–5:10 · The host as informed peer 6/10 Origins of vLLM and LLM Serving Challenges Elena sets the historical context of vLLM starting in 2022 pre-ChatGPT, and Matt demonstrates domain expertise recounting BERT, ResNet, and Hugging Face variants. Simon clearly explains the core technical challenges of non-deterministic output distributions and accelerator scheduling.5:10–8:31 · The host as informed peer 7/10 Open Source as Critical Enterprise Infrastructure Matt outlines how frontier enterprise apps like Cursor and Decagon moved away from thin OpenAI wrappers toward open source for post-training control. Simon explains how open weights became foundational once tools like Copilot became essential.8:31–12:59 · The host as informed peer 5/10 vLLM Architecture and Day Zero Release Engineering Simon walks through how vLLM operates across 1,000+ architectures and coordinates multi-party day-zero releases with hardware vendors and labs. Matt prompts Simon for behind-the-scenes engineering anecdotes such as the Mistral torrent release.12:59–18:50 · The host as informed peer 6/10 Open Weights Advocacy and Cost vs. Control Elena challenges Simon on his essay arguing that economics is besides the point when running frontier open models. Simon reframes the discussion, explaining that open weights provide fine-grained latency modes, zero-retention privacy, and architectural control beyond pure unit cost.18:50–20:54 · The host as informed peer 5/10 Performance Tuning and Open Source Licensing Models Matt brings up licensing evolution, citing Meta's Llama commercial revenue thresholds. Simon explains how model labs moved from permissive Apache 2.0 licenses to customized revenue-sharing terms to ensure financial sustainability.20:54–24:34 · The host as informed peer 7/10 Funding Models and the Economics of AI R&D Matt emphasizes that AI weights are capital-intensive compute artifacts rather than traditional software created by volunteer labor. Simon introduces an analogy to pharmaceutical R&D, which Matt engages with and unpacks in terms of closed patents versus open weights.24:34–30:09 · The host as informed peer 6/10 Open Source Inference Maintenance and Reliability Elena asks what open source AI maintenance entails, prompting Matt to describe cluster log panics during training runs. Simon details the complex post-release community effort required to optimize models across diverse hardware topologies and workloads.30:09–32:21 · The host as informed peer 5/10 The UC Berkeley Ecosystem and Mission Alignment Elena questions how projects like vLLM, Ollama, and OpenRouter emerged simultaneously around 2022-2023. Simon attributes this convergence to UC Berkeley's systems research culture and deep curiosity among mission-aligned engineers.32:21–36:31 · The host as informed peer 7/10 Content Moderation, API Guardrails, and Safety Edge Cases Simon reveals that vLLM engineers abandoned Anthropic models for Kimi K3 after false-positive safety filters blocked GPU kernel memory error research. Matt expands on this with an in-depth analysis of social media safe harbor regulations versus AI API liability.36:31–42:27 · The host as informed peer 5/10 Building Infraact with Ion Stoica Elena inquires about Ion Stoica's role at Infraact and asks whether open models will close the gap with frontier models within five years. Simon rejects the idea of a fundamental capability gap, highlighting architectural breakthroughs like RoPE's inventor simplifying Kimi K3.42:27–45:45 · The host as informed peer 6/10 Debunking Model Distillation and Policy Implications Matt asks if Chinese frontier models rely primarily on distillation from US proprietary models. Simon dismantles this premise, explaining that interactive evaluation environments and algorithmic design cannot be distilled from API outputs.1:18–5:10 · Guest teaching 5/10 Origins of vLLM and LLM Serving Challenges Elena sets the historical context of vLLM starting in 2022 pre-ChatGPT, and Matt demonstrates domain expertise recounting BERT, ResNet, and Hugging Face variants. Simon clearly explains the core technical challenges of non-deterministic output distributions and accelerator scheduling.5:10–8:31 · Guest teaching 4/10 Open Source as Critical Enterprise Infrastructure Matt outlines how frontier enterprise apps like Cursor and Decagon moved away from thin OpenAI wrappers toward open source for post-training control. Simon explains how open weights became foundational once tools like Copilot became essential.8:31–12:59 · Guest teaching 6/10 vLLM Architecture and Day Zero Release Engineering Simon walks through how vLLM operates across 1,000+ architectures and coordinates multi-party day-zero releases with hardware vendors and labs. Matt prompts Simon for behind-the-scenes engineering anecdotes such as the Mistral torrent release.12:59–18:50 · Guest teaching 6/10 Open Weights Advocacy and Cost vs. Control Elena challenges Simon on his essay arguing that economics is besides the point when running frontier open models. Simon reframes the discussion, explaining that open weights provide fine-grained latency modes, zero-retention privacy, and architectural control beyond pure unit cost.18:50–20:54 · Guest teaching 5/10 Performance Tuning and Open Source Licensing Models Matt brings up licensing evolution, citing Meta's Llama commercial revenue thresholds. Simon explains how model labs moved from permissive Apache 2.0 licenses to customized revenue-sharing terms to ensure financial sustainability.20:54–24:34 · Guest teaching 5/10 Funding Models and the Economics of AI R&D Matt emphasizes that AI weights are capital-intensive compute artifacts rather than traditional software created by volunteer labor. Simon introduces an analogy to pharmaceutical R&D, which Matt engages with and unpacks in terms of closed patents versus open weights.24:34–30:09 · Guest teaching 6/10 Open Source Inference Maintenance and Reliability Elena asks what open source AI maintenance entails, prompting Matt to describe cluster log panics during training runs. Simon details the complex post-release community effort required to optimize models across diverse hardware topologies and workloads.30:09–32:21 · Guest teaching 5/10 The UC Berkeley Ecosystem and Mission Alignment Elena questions how projects like vLLM, Ollama, and OpenRouter emerged simultaneously around 2022-2023. Simon attributes this convergence to UC Berkeley's systems research culture and deep curiosity among mission-aligned engineers.32:21–36:31 · Guest teaching 6/10 Content Moderation, API Guardrails, and Safety Edge Cases Simon reveals that vLLM engineers abandoned Anthropic models for Kimi K3 after false-positive safety filters blocked GPU kernel memory error research. Matt expands on this with an in-depth analysis of social media safe harbor regulations versus AI API liability.36:31–42:27 · Guest teaching 6/10 Building Infraact with Ion Stoica Elena inquires about Ion Stoica's role at Infraact and asks whether open models will close the gap with frontier models within five years. Simon rejects the idea of a fundamental capability gap, highlighting architectural breakthroughs like RoPE's inventor simplifying Kimi K3.42:27–45:45 · Guest teaching 6/10 Debunking Model Distillation and Policy Implications Matt asks if Chinese frontier models rely primarily on distillation from US proprietary models. Simon dismantles this premise, explaining that interactive evaluation environments and algorithmic design cannot be distilled from API outputs.1:18–5:10 · Guest disagreement 1/10 Origins of vLLM and LLM Serving Challenges Elena sets the historical context of vLLM starting in 2022 pre-ChatGPT, and Matt demonstrates domain expertise recounting BERT, ResNet, and Hugging Face variants. Simon clearly explains the core technical challenges of non-deterministic output distributions and accelerator scheduling.5:10–8:31 · Guest disagreement 1/10 Open Source as Critical Enterprise Infrastructure Matt outlines how frontier enterprise apps like Cursor and Decagon moved away from thin OpenAI wrappers toward open source for post-training control. Simon explains how open weights became foundational once tools like Copilot became essential.8:31–12:59 · Guest disagreement 1/10 vLLM Architecture and Day Zero Release Engineering Simon walks through how vLLM operates across 1,000+ architectures and coordinates multi-party day-zero releases with hardware vendors and labs. Matt prompts Simon for behind-the-scenes engineering anecdotes such as the Mistral torrent release.12:59–18:50 · Guest disagreement 2/10 Open Weights Advocacy and Cost vs. Control Elena challenges Simon on his essay arguing that economics is besides the point when running frontier open models. Simon reframes the discussion, explaining that open weights provide fine-grained latency modes, zero-retention privacy, and architectural control beyond pure unit cost.18:50–20:54 · Guest disagreement 1/10 Performance Tuning and Open Source Licensing Models Matt brings up licensing evolution, citing Meta's Llama commercial revenue thresholds. Simon explains how model labs moved from permissive Apache 2.0 licenses to customized revenue-sharing terms to ensure financial sustainability.20:54–24:34 · Guest disagreement 1/10 Funding Models and the Economics of AI R&D Matt emphasizes that AI weights are capital-intensive compute artifacts rather than traditional software created by volunteer labor. Simon introduces an analogy to pharmaceutical R&D, which Matt engages with and unpacks in terms of closed patents versus open weights.24:34–30:09 · Guest disagreement 1/10 Open Source Inference Maintenance and Reliability Elena asks what open source AI maintenance entails, prompting Matt to describe cluster log panics during training runs. Simon details the complex post-release community effort required to optimize models across diverse hardware topologies and workloads.30:09–32:21 · Guest disagreement 1/10 The UC Berkeley Ecosystem and Mission Alignment Elena questions how projects like vLLM, Ollama, and OpenRouter emerged simultaneously around 2022-2023. Simon attributes this convergence to UC Berkeley's systems research culture and deep curiosity among mission-aligned engineers.32:21–36:31 · Guest disagreement 2/10 Content Moderation, API Guardrails, and Safety Edge Cases Simon reveals that vLLM engineers abandoned Anthropic models for Kimi K3 after false-positive safety filters blocked GPU kernel memory error research. Matt expands on this with an in-depth analysis of social media safe harbor regulations versus AI API liability.36:31–42:27 · Guest disagreement 2/10 Building Infraact with Ion Stoica Elena inquires about Ion Stoica's role at Infraact and asks whether open models will close the gap with frontier models within five years. Simon rejects the idea of a fundamental capability gap, highlighting architectural breakthroughs like RoPE's inventor simplifying Kimi K3.42:27–45:45 · Guest disagreement 2/10 Debunking Model Distillation and Policy Implications Matt asks if Chinese frontier models rely primarily on distillation from US proprietary models. Simon dismantles this premise, explaining that interactive evaluation environments and algorithmic design cannot be distilled from API outputs.1:18–5:10 · The host pushing back 1/10 Origins of vLLM and LLM Serving Challenges Elena sets the historical context of vLLM starting in 2022 pre-ChatGPT, and Matt demonstrates domain expertise recounting BERT, ResNet, and Hugging Face variants. Simon clearly explains the core technical challenges of non-deterministic output distributions and accelerator scheduling.5:10–8:31 · The host pushing back 2/10 Open Source as Critical Enterprise Infrastructure Matt outlines how frontier enterprise apps like Cursor and Decagon moved away from thin OpenAI wrappers toward open source for post-training control. Simon explains how open weights became foundational once tools like Copilot became essential.8:31–12:59 · The host pushing back 1/10 vLLM Architecture and Day Zero Release Engineering Simon walks through how vLLM operates across 1,000+ architectures and coordinates multi-party day-zero releases with hardware vendors and labs. Matt prompts Simon for behind-the-scenes engineering anecdotes such as the Mistral torrent release.12:59–18:50 · The host pushing back 3/10 Open Weights Advocacy and Cost vs. Control Elena challenges Simon on his essay arguing that economics is besides the point when running frontier open models. Simon reframes the discussion, explaining that open weights provide fine-grained latency modes, zero-retention privacy, and architectural control beyond pure unit cost.18:50–20:54 · The host pushing back 1/10 Performance Tuning and Open Source Licensing Models Matt brings up licensing evolution, citing Meta's Llama commercial revenue thresholds. Simon explains how model labs moved from permissive Apache 2.0 licenses to customized revenue-sharing terms to ensure financial sustainability.20:54–24:34 · The host pushing back 2/10 Funding Models and the Economics of AI R&D Matt emphasizes that AI weights are capital-intensive compute artifacts rather than traditional software created by volunteer labor. Simon introduces an analogy to pharmaceutical R&D, which Matt engages with and unpacks in terms of closed patents versus open weights.24:34–30:09 · The host pushing back 1/10 Open Source Inference Maintenance and Reliability Elena asks what open source AI maintenance entails, prompting Matt to describe cluster log panics during training runs. Simon details the complex post-release community effort required to optimize models across diverse hardware topologies and workloads.30:09–32:21 · The host pushing back 1/10 The UC Berkeley Ecosystem and Mission Alignment Elena questions how projects like vLLM, Ollama, and OpenRouter emerged simultaneously around 2022-2023. Simon attributes this convergence to UC Berkeley's systems research culture and deep curiosity among mission-aligned engineers.32:21–36:31 · The host pushing back 1/10 Content Moderation, API Guardrails, and Safety Edge Cases Simon reveals that vLLM engineers abandoned Anthropic models for Kimi K3 after false-positive safety filters blocked GPU kernel memory error research. Matt expands on this with an in-depth analysis of social media safe harbor regulations versus AI API liability.36:31–42:27 · The host pushing back 2/10 Building Infraact with Ion Stoica Elena inquires about Ion Stoica's role at Infraact and asks whether open models will close the gap with frontier models within five years. Simon rejects the idea of a fundamental capability gap, highlighting architectural breakthroughs like RoPE's inventor simplifying Kimi K3.42:27–45:45 · The host pushing back 1/10 Debunking Model Distillation and Policy Implications Matt asks if Chinese frontier models rely primarily on distillation from US proprietary models. Simon dismantles this premise, explaining that interactive evaluation environments and algorithmic design cannot be distilled from API outputs.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 43:32 Simon refutes distillation assumptions

Simon directly dismisses the common narrative that top open-weight labs rely on distillation, arguing that RL training environments cannot be copied.

Hardest push from the host ▶ 16:40 Elena questions the purpose of expensive open models

Elena directly challenges Simon on his own essay, pressing him on why developers should bother running open models if they match proprietary API costs.

Biggest teaching moment ▶ 33:15 Simon exposes overactive safety filters

Simon educates the hosts on how proprietary guardrails fail in technical work, detailing how Claude flagged low-level GPU kernel memory access errors as security violations.

The host holds their own ▶ 34:22 Matt analyzes Section 230 platform dynamics in AI

Matt demonstrates sharp regulatory expertise, comparing the lack of platform liability exemptions in AI to Section 230 safe harbors in early social media.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Origins of vLLM and LLM Serving Challenges 6511 Elena sets the historical context of vLLM starting in 2022 pre-ChatGPT, and Matt demonstrates domain expertise recounting BERT, ResNet, and Hugging Face variants. Simon clearly explains the core technical challenges of non-deterministic output distributions and accelerator scheduling.
Open Source as Critical Enterprise Infrastructure 7412 Matt outlines how frontier enterprise apps like Cursor and Decagon moved away from thin OpenAI wrappers toward open source for post-training control. Simon explains how open weights became foundational once tools like Copilot became essential.
vLLM Architecture and Day Zero Release Engineering 5611 Simon walks through how vLLM operates across 1,000+ architectures and coordinates multi-party day-zero releases with hardware vendors and labs. Matt prompts Simon for behind-the-scenes engineering anecdotes such as the Mistral torrent release.
Open Weights Advocacy and Cost vs. Control 6623 Elena challenges Simon on his essay arguing that economics is besides the point when running frontier open models. Simon reframes the discussion, explaining that open weights provide fine-grained latency modes, zero-retention privacy, and architectural control beyond pure unit cost.
Performance Tuning and Open Source Licensing Models 5511 Matt brings up licensing evolution, citing Meta's Llama commercial revenue thresholds. Simon explains how model labs moved from permissive Apache 2.0 licenses to customized revenue-sharing terms to ensure financial sustainability.
Funding Models and the Economics of AI R&D 7512 Matt emphasizes that AI weights are capital-intensive compute artifacts rather than traditional software created by volunteer labor. Simon introduces an analogy to pharmaceutical R&D, which Matt engages with and unpacks in terms of closed patents versus open weights.
Open Source Inference Maintenance and Reliability 6611 Elena asks what open source AI maintenance entails, prompting Matt to describe cluster log panics during training runs. Simon details the complex post-release community effort required to optimize models across diverse hardware topologies and workloads.
The UC Berkeley Ecosystem and Mission Alignment 5511 Elena questions how projects like vLLM, Ollama, and OpenRouter emerged simultaneously around 2022-2023. Simon attributes this convergence to UC Berkeley's systems research culture and deep curiosity among mission-aligned engineers.
Content Moderation, API Guardrails, and Safety Edge Cases 7621 Simon reveals that vLLM engineers abandoned Anthropic models for Kimi K3 after false-positive safety filters blocked GPU kernel memory error research. Matt expands on this with an in-depth analysis of social media safe harbor regulations versus AI API liability.
Building Infraact with Ion Stoica 5622 Elena inquires about Ion Stoica's role at Infraact and asks whether open models will close the gap with frontier models within five years. Simon rejects the idea of a fundamental capability gap, highlighting architectural breakthroughs like RoPE's inventor simplifying Kimi K3.
Debunking Model Distillation and Policy Implications 6621 Matt asks if Chinese frontier models rely primarily on distillation from US proprietary models. Simon dismantles this premise, explaining that interactive evaluation environments and algorithmic design cannot be distilled from API outputs.

Statements from this episode (19)

Assertion Not checkable as stated
Burger: vLLM Runs on 500,000 GPUs at Any Moment
“Today we're here with Simon Moe, co-founder of Infraact, and a lead maintainer of VLLM, the open source inference engine, now running on half a million GPUs at any moment.”
Elena Burger Aug 5, 2026 ▶ 1:01
Insight
Moe: LLM serving differs fundamentally from traditional ML workloads
“Serving large language model is a fundamentally different problem. Because serving it requires to run it on accelerators like GPUs or TPUs, and it is a computationally intensive process that will require a lot of engineering and ensuring that for each request,…”
Simon Moe Aug 5, 2026 ▶ 1:53
Assertion Not checkable as stated
Bornstein: Open source was the norm for early frontier AI models
“Open source was the norm for AI models early on, right? I mean, we literally had this company called OpenAI, which, you know, it's become a little bit of a joke. It's not as open as it once was or not nearly as open as it once was, but early on, all, all the f…”
Matt Bornstein Aug 5, 2026 ▶ 3:06
Assertion Not checkable as stated
Simon Moe: BERT was the first model requiring GPUs for efficient inference
“Probably BERT. And before that, it was like ResNet for computation, like images, computer vision classification. So ResNet already need to run on NVIDIA K-eighty, which is kind of one of the first SEU on AWS and other places. And, but way over, but even at thi…”
Simon Moe Aug 5, 2026 ▶ 3:46
Opinion
Bornstein: OpenAI and Anthropic models remain more critical than open source
“At any given point in time, including now, I think models from open AI and Anthropic are kind of more widely used and more critical kind of in general than open source models.”
Matt Bornstein Aug 5, 2026 ▶ 7:07
Assertion Supported
Bornstein: Startups like Cursor and Harvey shifted to open source for customization
“But I do think we passed a threshold in like, I want to say about a year ago where a bunch of smaller companies or like new application companies, as they were trying to figure out, how do I really build an AI? Without just being a wrapper on top of open AI, t…”
Matt Bornstein Aug 5, 2026 ▶ 7:19
Assertion Partly supported
Simon Mo: vLLM supports over 1,000 model architectures
“For VRM, we support more than a thousand model architecture up to today, and a lot of those are proprietary, but also a lot of those are open-weight, right?”
Simon Moe Aug 5, 2026 ▶ 8:54
Assertion Supported
Mo: Major chipmakers use vLLM as an internal benchmark
“And additionally, VLM also work closely with all the hardware vendors. So that means across like NVIDIA, AMD, Google, and Amazon, Intel, and a lot more, their newest chip will make sure VLM can run on them. And then a lot of cases they use VLM as a benchmark t…”
Simon Moe Aug 5, 2026 ▶ 9:29
Assertion Supported
Moe: Kimi K3 costs less than Claude or GPT but exceeds smaller open models
“Where Kimi K-Stri is not as expensive as Claude or GPT Sol, but it is a lot more expensive than JLN-F.”
Simon Moe Aug 5, 2026 ▶ 17:09
Assertion Supported
Moe: Open-weight inference can hit 500 tokens/sec, 2-3x faster than proprietary APIs
“But for open weight, when you are running it, every provider can offer potentially even 10 different levels of speed going from like the slowest mode, which can be a lot cheaper to 400 tokens per second almost up to 500 in many cases for some workloads. And th…”
Simon Moe Aug 5, 2026 ▶ 17:55
Assertion Not checkable as stated
Bornstein: Meta tailored Llama license thresholds to exclude only two companies
“I do remember that the numbers were like specifically chosen at that time that you could go find it was like two companies in the world. That like fit the definition that they had excluded from their license.”
Matt Bornstein Aug 5, 2026 ▶ 20:45
Insight
Bornstein: Volunteer open-source models cannot fund frontier AI development
“Like an AI model is not software at the end of the day. And so open source software used to be supported by people donating their time or big companies kind of authorizing their employees to donate their time. So it was sort of like a bulk in kind You know, do…”
Matt Bornstein Aug 5, 2026 ▶ 21:54
Prediction Not checkable as stated
Bornstein: Chinese AI labs without commercial funding will rely on government backing
“If there's no source of economic, if there's no source of funding for Moonshot to continue to train models, like, we know where the funding will come from instead, and it's not, like, something we, right, you know, it's government and things that, like, are ac…”
Matt Bornstein Aug 5, 2026 ▶ 22:41
Assertion Supported
Bornstein: AlexNet originally ran on only two GPUs
“AlexNet, First, you know, kind of, like, neural network to run on, on GPUs that we care about ran on two GPUs. And that's not like there are no missing decimal points or commas in there. Literally two.”
Matt Bornstein Aug 5, 2026 ▶ 28:09
Assertion Not checkable as stated
Moe: Most AI API services use open-source inference engines under the hood
“And this is where kind of, this is why open source inference is the current leading way right now instead of closed source inference engine. And frankly, right. All the, a lot of the open, a lot of the open, sorry. A lot of the influence cloud and API as a ser…”
Simon Moe Aug 5, 2026 ▶ 29:32
Prediction Not checkable as stated
Simon Moe: Users will default to open-weight AI for trusted use cases
“In the future, we'll also see for the trusted use case, people will go to open way by default because that is where you know for sure that the guardrail is lessened or you can control your guardrail for trusted use cases.”
Simon Moe Aug 5, 2026 ▶ 33:18
Disclosure
Moe: Developers abandon proprietary models due to false-positive safety guardrails
“A lot of our developer within Infrax and for VLM are like retreating from using Fable five because you have a two hour job and you trigger the red line, which is false positive. And then you have to lose all of your work. And so a lot of our developer are usin…”
Simon Moe Aug 5, 2026 ▶ 33:51
Opinion
Moe: Open and closed AI models have no capability gap today
“In the end, there's not much differentiation. It's more about the distribution strategy and go-to-market strategy. And the capability wise, I don't really see a big gap, not even today, because for how these models are coming to being, they're really starting …”
Simon Moe Aug 5, 2026 ▶ 38:48
Opinion
Moe: Model distillation is not the primary driver of Chinese AI progress
“So I really don't think from currently what we're seeing this is a big cornerstone of what's powering the progress today. In the end, what's powering the progress is still just really smart people with very interesting algorithms, data environment, and they wi…”
Simon Moe Aug 5, 2026 ▶ 44:13
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.