Hugging Face Transformers, every mention

11 scenes · ← back to Hugging Face Transformers

tap a year for its mentions
0053106202420252026episodesmentions
036202420252026episodes it came up in
001326202420252026episodesmentions per episode

every year anyone Jeremy Howard 2Ethan He 2Thomas Sohmers 1Omar Sanseviero 1Mitesh Agrawal 1Elie Bakouch 1

Verbatim, from the transcripts: the passages where Hugging Face Transformers comes up

loading…

⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind May 24, 2026 · 1 mention

⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF Oct 20, 2025 · 1 mention

  • ▶ 28:40 Elie Bakouch A GPT OSS coin three doesn't have Shard Expert, but like the coin three next, which is like, uh, really is like, uh, I mean, they didn't release it yet, but they submit the PR to Transformers.

⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI Aug 18, 2025 · 2 mentions

[Paper Club] BERT: Bidirectional Encoder Representations from Transformers Nov 27, 2024 · 1 mention

  • ▶ 35:05 unnamed speaker So a lot of this leans very heavily on the, uh, hugging face transformers library.

[Paper Club] Upcycling Large Language Models into Mixture of Experts Oct 29, 2024 · 2 mentions

  • ▶ 12:37 Ethan He Uh, let's also look at the implementation of Mixtro eight by seven on Hagen-Phys transformer.
  • ▶ 15:13 Ethan He You can combine it with hugging phase transformer without any problem.

Answer.ai & AI Magic with Jeremy Howard Aug 17, 2024 · 2 mentions

  • ▶ 42:17 Jeremy Howard And even as we did so, new regressions were appearing in, like, Transformers and stuff, that Benjamin then had to go away and figure out, like, oh, how come flash attention doesn't work in this version of Transformers anymore with this set… 2 times in the scene

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 1 mention

  • ▶ 1:10:47 unnamed speaker Like, XLM is not deterministic at all, compared to maybe, I think, Transformers is more deterministic.

Breaking down the OG GPT Paper by Alec Radford Apr 23, 2024 · 1 mention

  • ▶ 45:32 unnamed speaker And this is actually how the model looks like if you try to load in the transformers framework.

A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate Feb 28, 2024 · 3 mentions

  • ▶ 52:39 unnamed speaker Mozilla came out with a llama file, and then, um, I don't know if this is in the same category even, but I'm just gonna throw it in there, like Hugging Face has the Transformers and Diffuses library, which is a way of disseminating models… 3 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.