GLM 5.2

also referred to as: glm-5.2

part of GLM

4 statements across 1 episodes · 2 bullish · 1 bearish · 2 people on the record · first statement Aug 3, 2026 by Ali Taha · said 25 times in 2 episodes since 2026 · across every show →

Mentions by year

brought up most by Ali Taha (12), Alessio Fanelli (8), Philip Kiely (4), Eiso Kant (1)

tap a year for its mentions
001312522026episodesmentions
0122026episodes it came up in
007.511522026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about GLM 5.2, oldest first

Aug 3, 2026 bullish
Disclosure
GLM-5.2 autonomously wrote and guided production GPU kernels for Baseten's inference engine
“Some of the GPU kernels that were on GLM-Five-two within our inference engine is written by GLM-Five-two. And the trace and the kernels were guided by GLM-Five-two as the driver.”
Ali Taha Aug 3, 2026 ▶ 1:34:09 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 positive
Assertion Partly supported
Baseten's vision-retrofitted GLM-5.2 scored 56% on MMLU Pro without text degradation
“It's not, you know, it got to a 56% on MMLU Pro, I think, so not, not quite Frontier, but if you're running this model, you haven't suffered any loss on your GLM-Five-II quality.”
Philip Kiely Aug 3, 2026 ▶ 19:22 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 negative
Insight
Modifying base LLM weights for vision degrades original text performance
“You don't want to mess with the model weights because you run a chance of making the model dumber at something else for the purpose of giving it vision.”
Philip Kiely Aug 3, 2026 ▶ 17:13 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026
Assertion Not checkable as stated
Unoptimized GLM-5.2 delivers a baseline 30 to 40 tokens per second
“So let's say you have, as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10 X that. So like on GLM 5.2 if you want to get unquantized perhaps on hoppers even and you're just using an off the shelf inference engine with no particular optimiz…”
Philip Kiely Aug 3, 2026 ▶ 37:05 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.