Gemma, every mention
59 scenes, the whole family · ← back to Gemma
tap a year for its mentions
every year anyone Omar Sanseviero 32Alessio Fanelli 17Shawn Wang 12Rishabh Agarwal 11Vibhu (Viboo) 8Emmanuel Ameisen 4Peter Robicheaux 2Jack Morris 2Sean Lie 1Mark Bissell 1
Verbatim, from the transcripts: the passages where Gemma comes up
The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 48:01 Alessio Fanelli How much of this applies to, say I have this MacBook, I want to run Gemma really efficiently, um,
- ▶ 1:08:41 Alessio Fanelli Gemma, no encoder, the latest thinking machines is all from scratch.
- ▶ 1:26:03 Alessio Fanelli There's Gemma as well, right?
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
- ▶ 24:42 unnamed speaker We're very much like, ah, okay, look, it's like, you know, on par with Kimmy, DeepSeek, whatnot, the small ones, Gemma level.
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
- ▶ 33:51 unnamed speaker It's a Gemma level model trained on roughly 40 trillion tokens at this many H 200 over this much time, right?
⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
- ▶ 0:05 Alessio Fanelli Gemma Four, Gemma Three One, Gemma Scope, Med Gemma.
- ▶ 0:05 Alessio Fanelli Gemma Four, Gemma Three One, Gemma Scope, Med Gemma. 5 times in the scene
- ▶ 0:05 Alessio Fanelli Gemma Four, Gemma Three One, Gemma Scope, Med Gemma.
- ▶ 2:13 Omar Sanseviero Yeah, so actually, if you install, like, if you buy a Pixel phone or a high-end Samsung, they come with a Gemini Nano, and Gemini Nano is packed into the operating system, and Gemini Nano is really built on top of Gemma.
- ▶ 2:25 Omar Sanseviero So last year, we released Gemma three end, which was this architecture really designed for phone use cases, and they use a Gemma three end with some additional training, some additional adaptations to make the model good for, like,… 3 times in the scene
- ▶ 2:53 Shawn Wang Is it exactly the same in the Gemma IV stuff? 2 times in the scene
- ▶ 3:21 Omar Sanseviero The Gemma team is actually relatively small. 6 times in the scene
- ▶ 3:49 Omar Sanseviero So we have almost 50 external partners for every, well, for the Gemma for launch, which has been the most complex launch. 4 times in the scene
- ▶ 6:35 Omar Sanseviero Yeah, so Gemma four was built on the same research as Gemini three, which pretty much means that we benefited from all of the improvements that happened with Gemini three. 6 times in the scene
- ▶ 10:44 Omar Sanseviero So, uh, we brought one of the researchers that worked in the Gemma development, in the development of Gemma four.
- ▶ 14:01 Omar Sanseviero So as I was saying, like for Gemma four, we had 50 5 times in the scene
- ▶ 16:29 Alessio Fanelli Yeah, I have a question about the bigger Gemma models. 2 times in the scene
- ▶ 18:59 Omar Sanseviero There are the billion parameters for Gemma two, three, and four, and the intelligence is much higher, right?
- ▶ 18:59 Omar Sanseviero There are the billion parameters for Gemma two, three, and four, and the intelligence is much higher, right?
- ▶ 18:59 Omar Sanseviero There are the billion parameters for Gemma two, three, and four, and the intelligence is much higher, right?
- ▶ 20:09 Alessio Fanelli Gemma Scope.
- ▶ 20:21 Omar Sanseviero And yeah, the team released, I don't know if it was a couple of terabytes, maybe even up to like one petabyte of data that we had to store because we did that for every single layer across all of the Gemma three models. 2 times in the scene
- ▶ 29:18 Omar Sanseviero I mean, the way we are doing Gemma, Gemini, and all of our tools is really like based on the feedback from the startups, the community, the developers, that's why you see like Logan, Paige, everyone in the team talking with the community…
The $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied Intuition
- ▶ 48:58 Shawn Wang So it's like Google just released the Gemma to be model that you effective to be model. 3 times in the scene
The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
- ▶ 50:07 Shawn Wang But, uh, the Gemma models, for example, right?
- ▶ 54:51 Shawn Wang And for listeners, I think I will highlight the Gemma three end paper where they, there was a little bit of that, I think.
Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
- ▶ 46:10 unnamed speaker Um, DeepMind has opened a lot of essays on, um, Gemma.
- ▶ 49:43 Mark Bissell Robotics, I know, like, a lot of the companies just use Gemma as, like, the, the, like, backbone, and then they, like, make it into a VLA that, like, takes these actions.
⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
- ▶ 31:20 Shawn Wang Gemma, Gemma, Gemma.
A Technical History of Generative Media
- ▶ 38:20 unnamed speaker Then you got Gemma, two hundred seventy million.
Better Data is All You Need — Ari Morcos, Datology
- ▶ 1:07:06 Alessio Fanelli And this is a 4.5 B model, which is par with Gemma four B and a little worse than Quan three, but roughly the same. 2 times in the scene
Information Theory for Language Models: Jack Morris
- ▶ 40:57 Shawn Wang Uh, so Gemma three N is like a really good candidate right now because it's like a four B model that is like claimed to be better than Lama four and GPT 4.1, uh, according to, you know, certain arenas that shall not be named.
- ▶ 48:55 Shawn Wang Yeah, I would say, uh, okay, I pulled out something very current, uh, which is Gemma three N, which launched, which, uh, sort of was, uh, generally available yesterday. 2 times in the scene
- ▶ 1:01:58 Jack Morris So like you were just mentioning Gemma three B came out yesterday and you can download it and it takes up a certain amount of space on disk 2 times in the scene
The Utility of Interpretability — Emmanuel Amiesen
- ▶ 1:41 Emmanuel Ameisen So notably maybe the most easy one here is like Gemma two to be.
- ▶ 2:37 Vibhu (Viboo) What, why should we probe Gemma, Lama? 4 times in the scene
- ▶ 7:54 Emmanuel Ameisen And so if you say like, thanks for having me on the whatever, like Gemma seems to have pretty consistently guessed that you're like on a podcast, uh, which makes sense, right? 2 times in the scene
- ▶ 22:35 Vibhu (Viboo) So if I have a base Gemma and I have a chat model, what are differences in their attributions, right? 4 times in the scene
- ▶ 43:10 unnamed speaker He did Neuronpedia and released a bunch of SAEs for, I think, the Llama models and the Gemma models. 2 times in the scene
- ▶ 1:27:29 Emmanuel Ameisen There's some of the JAMA models.
The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
- ▶ 9:59 Rishabh Agarwal And this can, now you can see that this can be easily used for pre-training, and one notable example of this is GemRTool. 3 times in the scene
- ▶ 17:41 Rishabh Agarwal So one thing we found consistently, so here what we, we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for… 4 times in the scene
- ▶ 17:41 Rishabh Agarwal So one thing we found consistently, so here what we, we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for… 3 times in the scene
- ▶ 26:58 Rishabh Agarwal So this kind of thing was used in Jemma DuPo's training.
Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era
- ▶ 27:51 unnamed speaker But also, I mean, Google came through with, like, some other presenters, and you also had Kathleen from the Gemma team, and I think people are very excited about open models still.
Why is everyone cloning Deep Research?
- ▶ 9:48 Arush Sehgal Well, if you use our Gemma open source models, you could fine tune.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 13:25 Shawn Wang And so this includes Gemma.
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 31:35 Peter Robicheaux PolyGemma two, so PolyGemma uses Gemma as the language encoder, and it uses Gemma two B.
- ▶ 31:35 Peter Robicheaux PolyGemma two, so PolyGemma uses Gemma as the language encoder, and it uses Gemma two B.
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 15:51 unnamed speaker I would say also similar for Gemma, Gemma one and two, uh, Gemma two
- ▶ 15:51 unnamed speaker I would say also similar for Gemma, Gemma one and two, uh, Gemma two 2 times in the scene
- ▶ 15:51 unnamed speaker I would say also similar for Gemma, Gemma one and two, uh, Gemma two 3 times in the scene
- ▶ 49:31 unnamed speaker The last piece I had, which I kind of deleted, was, uh, there's a special mention, honorable mention of Gemma again, with PolyGemma, which is one of the smaller releases from Google I.O.
- ▶ 50:34 Alessio Fanelli I think maybe a lot of people initially wrote them off, but between, uh, you know, some of the Gemini Nano stuff, like, uh, Gemma II, Pali Gemma, we'll talk about some of the KV cache and context caching. 2 times in the scene
LLM Asia Paper Club Survey Round
- ▶ 15:06 unnamed speaker So before we go into like the paper itself, like actually, why does this matter?
- ▶ 23:28 unnamed speaker In the paper, they treat, they do the experiments with Lama seven B and Gemma, I think Lama two seven B and Gemma seven B as black boxes and vice using the other open source model as the white box for the uncertainty estimation.
- ▶ 47:54 unnamed speaker If you ever play, let's say, Gemma-IIb, it might not really be the same as a Lama-IIb, even if it's able to decode, like, six times as fast.