Baseten

also referred to as: base-ten

9 statements across 3 episodes · 4 bullish · 2 bearish · 4 people on the record · first statement Jan 19, 2025 by Yining Zhang · said 41 times in 5 episodes since 2025 · across every show →

Mentions by year

brought up most by Shawn Wang (11), Alessio Fanelli (2), Stephanie Palazzolo (1), Philip Kiely (1), Loïc Houssier (1)

tap a year for its mentions
0015230320252026episodesmentions
02320252026episodes it came up in
0051.510320252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Baseten, oldest first

Jan 19, 2025 negative
Assertion Not checkable as stated
Zhang: Llama 405B sees very few enterprise users compared to 70B
“I think at the base time, something like LAMA-Seventy-B is more common. I think LAMA-Seventy-B has released the 400 zero five billion weights, but I think there are just a few users use that.”
Yining Zhang Jan 19, 2025 ▶ 4:39 DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
Dec 11, 2025
Disclosure
Superhuman uses Baseten to run LLaMA and BERT classification models
“We use Base-Ten to run some I would say some LAMA, some BERT model for classification.”
Loïc Houssier Dec 11, 2025 ▶ 30:12 The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
Aug 3, 2026 positive
Assertion Supported
Retrofitting Grouped Query Attention layers cures inference bottlenecks and maintains acceptance
“So we find it better to, like, okay, we're going to replace this, you know, we're going to replace this layer with a layer from another model that's using, like, GQA, for instance. And then just for the right training, you can get it to have the same acceptanc…”
Ali Taha Aug 3, 2026 ▶ 21:14 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 neutral
Insight
AI customers migrate from pay-per-token APIs to dedicated GPUs upon finding fit
“I think that we've increasingly seen a lot of demand for the sort of pay per token APIs just because everyone wants to try open models. And then once they find a use case that's really sticky then they move over to Dedicated.”
Philip Kiely Aug 3, 2026 ▶ 5:03 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 bullish
Disclosure
GLM-5.2 autonomously wrote and guided production GPU kernels for Baseten's inference engine
“Some of the GPU kernels that were on GLM-Five-two within our inference engine is written by GLM-Five-two. And the trace and the kernels were guided by GLM-Five-two as the driver.”
Ali Taha Aug 3, 2026 ▶ 1:34:09 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 positive
Insight
Quantizing more layers can actually improve model fidelity via error cancellation
“It is possible that the model in which I quantized more information is going to perform better because the quantization errors have canceled out. And so what Joshua showed in his mathematical proof where he had like a verifier in is that you can predict which …”
Ali Taha Aug 3, 2026 ▶ 31:54 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026
Disclosure
Baseten terminates or reprocesses generations after four identical repeating tokens
“We have, like, in our endpoint, like, if a model was to output the same exact token, like, four plus times, we just call the generation, we say, like, oh, sorry, this, like, try again, or, like, we will re-process the request. Because we know then, like, if it…”
Ali Taha Aug 3, 2026 ▶ 22:55 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 negative
Insight
Unclosed JSON syntax is the primary cause of model tool hallucinations
“The challenge with tool calling more and more seems to be that the companies want certain tool calling, which is a very sensitive thing to train. And because you're dealing with all of the JSON outputs, if it doesn't like close the end of the request in a very…”
Ali Taha Aug 3, 2026 ▶ 7:57 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 positive
Insight
Hourly GPU instances are cheaper than pay-per-token for high-volume AI inference
“And more often than not, it's like way cheaper if you're pushing like millions of tokens per hour, if you just pay per hour instead of pay per token.”
Ali Taha Aug 3, 2026 ▶ 4:56 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.