Simon Moe

Co-Founder & CEO, Inferact · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutiveengineerscientist@simon_mo_ ↗inferact.ai ↗

Simon Mo is an AI systems researcher from UC Berkeley’s Sky Computing Lab who co-created and serves as the lead maintainer of vLLM. He co-founded Inferact to build commercial inference infrastructure around the vLLM ecosystem.

11statements → 7claims → 4claims resolved → 75%fully supported → 3.91/5average certainty → 2.45/5average debate potential → 1said about them ↓

3 supported 1 partly supported 0 contradicted 3 not checkable as stated how the 7 claims stand · each chip opens the sources

1 prediction · 6 assertions · 2 opinions · 1 insight · 1 disclosure · every statement was checked. The prediction and assertions are the 7 claims: statements the public record can support or contradict. 4 are resolved, and 3 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Simon argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Moe: Open-weight inference can hit 500 tokens/sec, 2-3x faster than proprietary APIs
“But for open weight, when you are running it, every provider can offer potentially even 10 different levels of speed going from like the slowest mode, which can be a lot cheaper to 400 tokens per second almost up to 500 in many cases for some workloads. And th…”
Simon Moe Aug 5, 2026 ▶ 17:55 How Open Source Became AI's Backbone | Inferact with a16z

How they sound: speaking style how? →

225 words/min while actually speaking · 21.2 um and uh per 1k words

No argument clarity score for Simon Moe: no usable question→answer exchanges on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to the audio itself: 4,335 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Simon Moe said on the a16z Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Disclosure
Moe: Developers abandon proprietary models due to false-positive safety guardrails
“A lot of our developer within Infrax and for VLM are like retreating from using Fable five because you have a two hour job and you trigger the red line, which is false positive. And then you have to lose all of your work. And so a lot of our developer are usin…”
Simon Moe Aug 5, 2026 ▶ 33:51 How Open Source Became AI's Backbone | Inferact with a16z
Opinion
Moe: Open and closed AI models have no capability gap today
“In the end, there's not much differentiation. It's more about the distribution strategy and go-to-market strategy. And the capability wise, I don't really see a big gap, not even today, because for how these models are coming to being, they're really starting …”
Simon Moe Aug 5, 2026 ▶ 38:48 How Open Source Became AI's Backbone | Inferact with a16z
Opinion
Moe: Model distillation is not the primary driver of Chinese AI progress
“So I really don't think from currently what we're seeing this is a big cornerstone of what's powering the progress today. In the end, what's powering the progress is still just really smart people with very interesting algorithms, data environment, and they wi…”
Simon Moe Aug 5, 2026 ▶ 44:13 How Open Source Became AI's Backbone | Inferact with a16z
Assertion Supported
Moe: Open-weight inference can hit 500 tokens/sec, 2-3x faster than proprietary APIs
“But for open weight, when you are running it, every provider can offer potentially even 10 different levels of speed going from like the slowest mode, which can be a lot cheaper to 400 tokens per second almost up to 500 in many cases for some workloads. And th…”
Simon Moe Aug 5, 2026 ▶ 17:55 How Open Source Became AI's Backbone | Inferact with a16z
Assertion Not checkable as stated
Moe: Most AI API services use open-source inference engines under the hood
“And this is where kind of, this is why open source inference is the current leading way right now instead of closed source inference engine. And frankly, right. All the, a lot of the open, a lot of the open, sorry. A lot of the influence cloud and API as a ser…”
Simon Moe Aug 5, 2026 ▶ 29:32 How Open Source Became AI's Backbone | Inferact with a16z
Prediction Not checkable as stated
Simon Moe: Users will default to open-weight AI for trusted use cases
“In the future, we'll also see for the trusted use case, people will go to open way by default because that is where you know for sure that the guardrail is lessened or you can control your guardrail for trusted use cases.”
Simon Moe Aug 5, 2026 ▶ 33:18 How Open Source Became AI's Backbone | Inferact with a16z
Insight
Moe: LLM serving differs fundamentally from traditional ML workloads
“Serving large language model is a fundamentally different problem. Because serving it requires to run it on accelerators like GPUs or TPUs, and it is a computationally intensive process that will require a lot of engineering and ensuring that for each request,…”
Simon Moe Aug 5, 2026 ▶ 1:53 How Open Source Became AI's Backbone | Inferact with a16z
Assertion Supported
Moe: Kimi K3 costs less than Claude or GPT but exceeds smaller open models
“Where Kimi K-Stri is not as expensive as Claude or GPT Sol, but it is a lot more expensive than JLN-F.”
Simon Moe Aug 5, 2026 ▶ 17:09 How Open Source Became AI's Backbone | Inferact with a16z
Assertion Not checkable as stated
Simon Moe: BERT was the first model requiring GPUs for efficient inference
“Probably BERT. And before that, it was like ResNet for computation, like images, computer vision classification. So ResNet already need to run on NVIDIA K-eighty, which is kind of one of the first SEU on AWS and other places. And, but way over, but even at thi…”
Simon Moe Aug 5, 2026 ▶ 3:46 How Open Source Became AI's Backbone | Inferact with a16z
Assertion Partly supported
Simon Mo: vLLM supports over 1,000 model architectures
“For VRM, we support more than a thousand model architecture up to today, and a lot of those are proprietary, but also a lot of those are open-weight, right?”
Simon Moe Aug 5, 2026 ▶ 8:54 How Open Source Became AI's Backbone | Inferact with a16z
Assertion Supported
Mo: Major chipmakers use vLLM as an internal benchmark
“And additionally, VLM also work closely with all the hardware vendors. So that means across like NVIDIA, AMD, Google, and Amazon, Intel, and a lot more, their newest chip will make sure VLM can run on them. And then a lot of cases they use VLM as a benchmark t…”
Simon Moe Aug 5, 2026 ▶ 9:29 How Open Source Became AI's Backbone | Inferact with a16z

The other half of the tape: Simon Moe's own voice is left out of every number here. 1 statement on the record names them. every mention, with the transcript →

Statements about Simon Moe, by other people (1)

Assertion Not checkable as stated
Burger: vLLM Runs on 500,000 GPUs at Any Moment
“Today we're here with Simon Moe, co-founder of Infraact, and a lead maintainer of VLLM, the open source inference engine, now running on half a million GPUs at any moment.”
Elena Burger Aug 5, 2026 ▶ 1:01 How Open Source Became AI's Backbone | Inferact with a16z

Appearances (1)

EpisodeDateSpeaking time
How Open Source Became AI's Backbone | Inferact with a16z Aug 5, 2026 23m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.