The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Moe: Developers abandon proprietary models due to false-positive safety guardrails
“A lot of our developer within Infrax and for VLM are like retreating from using Fable five because you have a two hour job and you trigger the red line, which is false positive. And then you have to lose all of your work. And so a lot of our developer are usin…”
Moe: Open and closed AI models have no capability gap today
“In the end, there's not much differentiation. It's more about the distribution strategy and go-to-market strategy. And the capability wise, I don't really see a big gap, not even today, because for how these models are coming to being, they're really starting …”
Moe: Model distillation is not the primary driver of Chinese AI progress
“So I really don't think from currently what we're seeing this is a big cornerstone of what's powering the progress today. In the end, what's powering the progress is still just really smart people with very interesting algorithms, data environment, and they wi…”
Moe: Open-weight inference can hit 500 tokens/sec, 2-3x faster than proprietary APIs
“But for open weight, when you are running it, every provider can offer potentially even 10 different levels of speed going from like the slowest mode, which can be a lot cheaper to 400 tokens per second almost up to 500 in many cases for some workloads. And th…”
Moe: Most AI API services use open-source inference engines under the hood
“And this is where kind of, this is why open source inference is the current leading way right now instead of closed source inference engine. And frankly, right. All the, a lot of the open, a lot of the open, sorry. A lot of the influence cloud and API as a ser…”
Simon Moe: Users will default to open-weight AI for trusted use cases
“In the future, we'll also see for the trusted use case, people will go to open way by default because that is where you know for sure that the guardrail is lessened or you can control your guardrail for trusted use cases.”
Moe: LLM serving differs fundamentally from traditional ML workloads
“Serving large language model is a fundamentally different problem. Because serving it requires to run it on accelerators like GPUs or TPUs, and it is a computationally intensive process that will require a lot of engineering and ensuring that for each request,…”
Moe: Kimi K3 costs less than Claude or GPT but exceeds smaller open models
“Where Kimi K-Stri is not as expensive as Claude or GPT Sol, but it is a lot more expensive than JLN-F.”
Simon Moe: BERT was the first model requiring GPUs for efficient inference
“Probably BERT. And before that, it was like ResNet for computation, like images, computer vision classification. So ResNet already need to run on NVIDIA K-eighty, which is kind of one of the first SEU on AWS and other places. And, but way over, but even at thi…”
Simon Mo: vLLM supports over 1,000 model architectures
“For VRM, we support more than a thousand model architecture up to today, and a lot of those are proprietary, but also a lot of those are open-weight, right?”
Mo: Major chipmakers use vLLM as an internal benchmark
“And additionally, VLM also work closely with all the hardware vendors. So that means across like NVIDIA, AMD, Google, and Amazon, Intel, and a lot more, their newest chip will make sure VLM can run on them. And then a lot of cases they use VLM as a benchmark t…”