GLM, every mention

16 scenes, the whole family · ← back to GLM

tap a year for its mentions
0020340520252026episodesmentions
03520252026episodes it came up in
0042.58520252026episodesmentions per episode

every year anyone Ali Taha 15Alessio Fanelli 10Philip Kiely 4Ronak Malde 1Elie Bakouch 1Eiso Kant 1

Verbatim, from the transcripts: the passages where GLM comes up

loading…

The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO Sep 2, 2026 · 1 mention

  • ▶ 43:48 unnamed speaker Give us, give us GLM, give us DeepSea.

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten Aug 3, 2026 · 29 mentions

  • ▶ 0:00 Ali Taha Jelenflag too is very, very good at writing GPU kernels. 4 times in the scene
  • ▶ 13:09 Alessio Fanelli Like, I think it was with, uh, Kimmy K-II .5 or GLM five two, the latest, there was sort of an inference war, right? 5 times in the scene
  • ▶ 16:40 Philip Kiely So something that, uh, Haley, a guy on our team, if, if we could take a look at this, um, he, like, kind of grafted the Kimi, um, vision encoder onto GLM 5.2. 3 times in the scene
  • ▶ 20:38 Ali Taha Well, to your point previously, when you were mentioning, like, um, the work that goes into supporting a model when it first comes out, like JLM-VII or Minimax M-III or whatever the case is, sometimes you do have to, like, you do have to… 3 times in the scene
  • ▶ 30:58 Ali Taha One of our research interns, Joshua, um, I think it's a tweet on, on how we have 20% better quantized JLM five two than NVIDIA.
  • ▶ 37:12 Philip Kiely So like on GLM 5.2, um, if you want to get unquantized, uh, perhaps on hoppers even, um, and you're just using an off the shelf inference engine with no particular optimizations, no, no speculator, um, nothing, nothing extra around like KV…
  • ▶ 39:11 Alessio Fanelli If you break down the two to four X, say, say the example is run GLM five two on B 200, single node, right? 3 times in the scene
  • ▶ 46:22 Alessio Fanelli So say for GLM.
  • ▶ 1:08:32 Alessio Fanelli I think on your guys' end, you see a lot of, okay, one day it's GLM, Kimmy, DeepSeq, uh, Minimax, throw in the others.
  • ▶ 1:33:25 Ali Taha We do see it, like, with GLM-Five-T, for instance. 7 times in the scene

The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI Jul 22, 2026 · 1 mention

  • ▶ 32:31 Eiso Kant I think obviously everyone's been talking about Zifu lately, uh, with, with GLM 5.2.

The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO Jul 8, 2026 · 1 mention

  • ▶ 21:30 unnamed speaker How does it compare to, say, I take the same model, GLM, 5.2 FBA, take off the shelf inference engine, VLM, SGLang, um, you know, get compute of similar capacity, similar cost.

⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai Jun 21, 2026 · 1 mention

  • ▶ 14:46 Ronak Malde Like obviously having one trillion parameter models like Kimi or like an amazing models like GLM and DeepSeq, I don't think we're quite there yet for that size of model.

⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF Oct 20, 2025 · 1 mention

  • ▶ 11:08 Elie Bakouch And, uh, recently there was the GLM paper, there is this, um, uh, King K-II, of course.

⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras Oct 1, 2025 · 1 mention

  • ▶ 13:52 unnamed speaker The GLM just came out today, and I know there's a lot of work that's going on between our companies on, on, on that, on the open model side.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.