GLM 5.2, every mention
8 scenes · ← back to GLM 5.2
tap a year for its mentions
every year anyone Ali Taha 12Alessio Fanelli 8Philip Kiely 4Eiso Kant 1
Verbatim, from the transcripts: the passages where GLM 5.2 comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 0:00 Ali Taha Jelenflag too is very, very good at writing GPU kernels. 4 times in the scene
- ▶ 13:09 Alessio Fanelli Like, I think it was with, uh, Kimmy K-II .5 or GLM five two, the latest, there was sort of an inference war, right? 5 times in the scene
- ▶ 16:40 Philip Kiely So something that, uh, Haley, a guy on our team, if, if we could take a look at this, um, he, like, kind of grafted the Kimi, um, vision encoder onto GLM 5.2. 3 times in the scene
- ▶ 30:58 Ali Taha One of our research interns, Joshua, um, I think it's a tweet on, on how we have 20% better quantized JLM five two than NVIDIA.
- ▶ 37:12 Philip Kiely So like on GLM 5.2, um, if you want to get unquantized, uh, perhaps on hoppers even, um, and you're just using an off the shelf inference engine with no particular optimizations, no, no speculator, um, nothing, nothing extra around like KV…
- ▶ 39:11 Alessio Fanelli If you break down the two to four X, say, say the example is run GLM five two on B 200, single node, right? 3 times in the scene
- ▶ 1:33:25 Ali Taha We do see it, like, with GLM-Five-T, for instance. 7 times in the scene