GLM, every mention
16 scenes, the whole family · ← back to GLM
tap a year for its mentions
every year anyone Ali Taha 15Alessio Fanelli 10Philip Kiely 4Ronak Malde 1Elie Bakouch 1Eiso Kant 1
Verbatim, from the transcripts: the passages where GLM comes up
The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
- ▶ 43:48 unnamed speaker Give us, give us GLM, give us DeepSea.
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 0:00 Ali Taha Jelenflag too is very, very good at writing GPU kernels. 4 times in the scene
- ▶ 13:09 Alessio Fanelli Like, I think it was with, uh, Kimmy K-II .5 or GLM five two, the latest, there was sort of an inference war, right? 5 times in the scene
- ▶ 16:40 Philip Kiely So something that, uh, Haley, a guy on our team, if, if we could take a look at this, um, he, like, kind of grafted the Kimi, um, vision encoder onto GLM 5.2. 3 times in the scene
- ▶ 20:38 Ali Taha Well, to your point previously, when you were mentioning, like, um, the work that goes into supporting a model when it first comes out, like JLM-VII or Minimax M-III or whatever the case is, sometimes you do have to, like, you do have to… 3 times in the scene
- ▶ 30:58 Ali Taha One of our research interns, Joshua, um, I think it's a tweet on, on how we have 20% better quantized JLM five two than NVIDIA.
- ▶ 37:12 Philip Kiely So like on GLM 5.2, um, if you want to get unquantized, uh, perhaps on hoppers even, um, and you're just using an off the shelf inference engine with no particular optimizations, no, no speculator, um, nothing, nothing extra around like KV…
- ▶ 39:11 Alessio Fanelli If you break down the two to four X, say, say the example is run GLM five two on B 200, single node, right? 3 times in the scene
- ▶ 46:22 Alessio Fanelli So say for GLM.
- ▶ 1:08:32 Alessio Fanelli I think on your guys' end, you see a lot of, okay, one day it's GLM, Kimmy, DeepSeq, uh, Minimax, throw in the others.
- ▶ 1:33:25 Ali Taha We do see it, like, with GLM-Five-T, for instance. 7 times in the scene
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
- ▶ 21:30 unnamed speaker How does it compare to, say, I take the same model, GLM, 5.2 FBA, take off the shelf inference engine, VLM, SGLang, um, you know, get compute of similar capacity, similar cost.
⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
- ▶ 14:46 Ronak Malde Like obviously having one trillion parameter models like Kimi or like an amazing models like GLM and DeepSeq, I don't think we're quite there yet for that size of model.
⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
- ▶ 11:08 Elie Bakouch And, uh, recently there was the GLM paper, there is this, um, uh, King K-II, of course.
⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
- ▶ 13:52 unnamed speaker The GLM just came out today, and I know there's a lot of work that's going on between our companies on, on, on that, on the open model side.