Gemma, every mention

59 scenes, the whole family · ← back to Gemma

tap a year for its mentions
003056010202420252026episodesmentions
0510202420252026episodes it came up in
0045810202420252026episodesmentions per episode

every year anyone Omar Sanseviero 32Alessio Fanelli 17Shawn Wang 12Rishabh Agarwal 11Vibhu (Viboo) 8Emmanuel Ameisen 4Peter Robicheaux 2Jack Morris 2Sean Lie 1Mark Bissell 1

Verbatim, from the transcripts: the passages where Gemma comes up

loading…

The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO Sep 2, 2026 · 2 mentions

  • ▶ 11:25 Sean Lie And what this ultimately means is you'll be able to run, you know, medium sized models like GPT-OSS or JAMA at speeds up to 10,000 TPS.
  • ▶ 12:15 unnamed speaker Right now, sure, there's a Gemma OSS.

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten Aug 3, 2026 · 3 mentions

The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI Jul 22, 2026 · 1 mention

  • ▶ 24:42 unnamed speaker We're very much like, ah, okay, look, it's like, you know, on par with Kimmy, DeepSeek, whatnot, the small ones, Gemma level.

Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He Jun 1, 2026 · 1 mention

  • ▶ 33:51 unnamed speaker It's a Gemma level model trained on roughly 40 trillion tokens at this many H 200 over this much time, right?

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind May 24, 2026 · 44 mentions

  • ▶ 0:05 Alessio Fanelli Gemma Four, Gemma Three One, Gemma Scope, Med Gemma.
  • ▶ 0:05 Alessio Fanelli Gemma Four, Gemma Three One, Gemma Scope, Med Gemma. 5 times in the scene
  • ▶ 0:05 Alessio Fanelli Gemma Four, Gemma Three One, Gemma Scope, Med Gemma.
  • ▶ 2:13 Omar Sanseviero Yeah, so actually, if you install, like, if you buy a Pixel phone or a high-end Samsung, they come with a Gemini Nano, and Gemini Nano is packed into the operating system, and Gemini Nano is really built on top of Gemma.
  • ▶ 2:25 Omar Sanseviero So last year, we released Gemma three end, which was this architecture really designed for phone use cases, and they use a Gemma three end with some additional training, some additional adaptations to make the model good for, like,… 3 times in the scene
  • ▶ 2:53 Shawn Wang Is it exactly the same in the Gemma IV stuff? 2 times in the scene
  • ▶ 3:21 Omar Sanseviero The Gemma team is actually relatively small. 6 times in the scene
  • ▶ 3:49 Omar Sanseviero So we have almost 50 external partners for every, well, for the Gemma for launch, which has been the most complex launch. 4 times in the scene
  • ▶ 6:35 Omar Sanseviero Yeah, so Gemma four was built on the same research as Gemini three, which pretty much means that we benefited from all of the improvements that happened with Gemini three. 6 times in the scene
  • ▶ 10:44 Omar Sanseviero So, uh, we brought one of the researchers that worked in the Gemma development, in the development of Gemma four.
  • ▶ 14:01 Omar Sanseviero So as I was saying, like for Gemma four, we had 50 5 times in the scene
  • ▶ 16:29 Alessio Fanelli Yeah, I have a question about the bigger Gemma models. 2 times in the scene
  • ▶ 18:59 Omar Sanseviero There are the billion parameters for Gemma two, three, and four, and the intelligence is much higher, right?
  • ▶ 18:59 Omar Sanseviero There are the billion parameters for Gemma two, three, and four, and the intelligence is much higher, right?
  • ▶ 18:59 Omar Sanseviero There are the billion parameters for Gemma two, three, and four, and the intelligence is much higher, right?
  • ▶ 20:09 Alessio Fanelli Gemma Scope.
  • ▶ 20:21 Omar Sanseviero And yeah, the team released, I don't know if it was a couple of terabytes, maybe even up to like one petabyte of data that we had to store because we did that for every single layer across all of the Gemma three models. 2 times in the scene
  • ▶ 29:18 Omar Sanseviero I mean, the way we are doing Gemma, Gemini, and all of our tools is really like based on the feedback from the startups, the community, the developers, that's why you see like Logan, Paige, everyone in the team talking with the community…

The $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied Intuition Apr 27, 2026 · 3 mentions

  • ▶ 48:58 Shawn Wang So it's like Google just released the Gemma to be model that you effective to be model. 3 times in the scene

The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean Feb 12, 2026 · 2 mentions

  • ▶ 50:07 Shawn Wang But, uh, the Gemma models, for example, right?
  • ▶ 54:51 Shawn Wang And for listeners, I think I will highlight the Gemma three end paper where they, there was a little bit of that, I think.

Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell Feb 5, 2026 · 2 mentions

  • ▶ 46:10 unnamed speaker Um, DeepMind has opened a lot of essays on, um, Gemma.
  • ▶ 49:43 Mark Bissell Robotics, I know, like, a lot of the companies just use Gemma as, like, the, the, like, backbone, and then they, like, make it into a VLA that, like, takes these actions.

⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF Oct 20, 2025 · 1 mention

A Technical History of Generative Media Sep 8, 2025 · 1 mention

  • ▶ 38:20 unnamed speaker Then you got Gemma, two hundred seventy million.

Better Data is All You Need — Ari Morcos, Datology Aug 29, 2025 · 2 mentions

  • ▶ 1:07:06 Alessio Fanelli And this is a 4.5 B model, which is par with Gemma four B and a little worse than Quan three, but roughly the same. 2 times in the scene

Information Theory for Language Models: Jack Morris Jul 2, 2025 · 5 mentions

  • ▶ 40:57 Shawn Wang Uh, so Gemma three N is like a really good candidate right now because it's like a four B model that is like claimed to be better than Lama four and GPT 4.1, uh, according to, you know, certain arenas that shall not be named.
  • ▶ 48:55 Shawn Wang Yeah, I would say, uh, okay, I pulled out something very current, uh, which is Gemma three N, which launched, which, uh, sort of was, uh, generally available yesterday. 2 times in the scene
  • ▶ 1:01:58 Jack Morris So like you were just mentioning Gemma three B came out yesterday and you can download it and it takes up a certain amount of space on disk 2 times in the scene

The Utility of Interpretability — Emmanuel Amiesen Jun 6, 2025 · 14 mentions

  • ▶ 1:41 Emmanuel Ameisen So notably maybe the most easy one here is like Gemma two to be.
  • ▶ 2:37 Vibhu (Viboo) What, why should we probe Gemma, Lama? 4 times in the scene
  • ▶ 7:54 Emmanuel Ameisen And so if you say like, thanks for having me on the whatever, like Gemma seems to have pretty consistently guessed that you're like on a podcast, uh, which makes sense, right? 2 times in the scene
  • ▶ 22:35 Vibhu (Viboo) So if I have a base Gemma and I have a chat model, what are differences in their attributions, right? 4 times in the scene
  • ▶ 43:10 unnamed speaker He did Neuronpedia and released a bunch of SAEs for, I think, the Llama models and the Gemma models. 2 times in the scene
  • ▶ 1:27:29 Emmanuel Ameisen There's some of the JAMA models.

The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind Mar 23, 2025 · 11 mentions

  • ▶ 9:59 Rishabh Agarwal And this can, now you can see that this can be easily used for pre-training, and one notable example of this is GemRTool. 3 times in the scene
  • ▶ 17:41 Rishabh Agarwal So one thing we found consistently, so here what we, we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for… 4 times in the scene
  • ▶ 17:41 Rishabh Agarwal So one thing we found consistently, so here what we, we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for… 3 times in the scene
  • ▶ 26:58 Rishabh Agarwal So this kind of thing was used in Jemma DuPo's training.

Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era Feb 28, 2025 · 1 mention

  • ▶ 27:51 unnamed speaker But also, I mean, Google came through with, like, some other presenters, and you also had Kathleen from the Gemma team, and I think people are very excited about open models still.

Why is everyone cloning Deep Research? Feb 18, 2025 · 1 mention

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 1 mention

Best of 2024 in Vision [LS Live @ NeurIPS] Dec 22, 2024 · 2 mentions

The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap) Aug 2, 2024 · 9 mentions

  • ▶ 15:51 unnamed speaker I would say also similar for Gemma, Gemma one and two, uh, Gemma two
  • ▶ 15:51 unnamed speaker I would say also similar for Gemma, Gemma one and two, uh, Gemma two 2 times in the scene
  • ▶ 15:51 unnamed speaker I would say also similar for Gemma, Gemma one and two, uh, Gemma two 3 times in the scene
  • ▶ 49:31 unnamed speaker The last piece I had, which I kind of deleted, was, uh, there's a special mention, honorable mention of Gemma again, with PolyGemma, which is one of the smaller releases from Google I.O.
  • ▶ 50:34 Alessio Fanelli I think maybe a lot of people initially wrote them off, but between, uh, you know, some of the Gemini Nano stuff, like, uh, Gemma II, Pali Gemma, we'll talk about some of the KV cache and context caching. 2 times in the scene

LLM Asia Paper Club Survey Round May 22, 2024 · 3 mentions

  • ▶ 15:06 unnamed speaker So before we go into like the paper itself, like actually, why does this matter?
  • ▶ 23:28 unnamed speaker In the paper, they treat, they do the experiments with Lama seven B and Gemma, I think Lama two seven B and Gemma seven B as black boxes and vice using the other open source model as the white box for the uncertainty estimation.
  • ▶ 47:54 unnamed speaker If you ever play, let's say, Gemma-IIb, it might not really be the same as a Lama-IIb, even if it's able to decode, like, six times as fast.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.