T5, every mention

21 scenes · ← back to T5

tap a year for its mentions
001342582023202420252026episodesmentions
0482023202420252026episodes it came up in
002.54582023202420252026episodesmentions per episode

every year anyone Yi Tay 11Rishabh Agarwal 2Michael Royzen 2Jeremy Howard 1Jack Morris 1Ethan He 1

Verbatim, from the transcripts: the passages where T5 comes up

loading…

Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay Jan 23, 2026 · 5 mentions

  • ▶ 4:16 unnamed speaker Superficially, I see some, you know, in your UL II and T-Five work, um, some,
  • ▶ 1:10:17 Yi Tay So we wanted, at that time LM, we were still using T-Five models at that time.
  • ▶ 1:14:23 unnamed speaker And I think like, it's also somewhat emergent in the sense that when you were using T-Five, you just couldn't actually add that much value on top of a normal BM-T-Five retrieval technique. 3 times in the scene

Information Theory for Language Models: Jack Morris Jul 2, 2025 · 1 mention

  • ▶ 46:39 Jack Morris So I think these are GTR, which is a T five based retrieval model and GTE, which is based on BERT.

The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind Mar 23, 2025 · 2 mentions

[Paper Club] BERT: Bidirectional Encoder Representations from Transformers Nov 27, 2024 · 1 mention

  • ▶ 1:26 unnamed speaker Bart and T five are good, but like, let's look at a host.

[Paper Club] Upcycling Large Language Models into Mixture of Experts Oct 29, 2024 · 1 mention

  • ▶ 7:21 Ethan He Um, we accelerate not only MOE and also all of the LLMs, including, like, GPT, BERT, T-Five, uh, not sure if anyone is still using those now, but primarily GPT models, and in the

Answer.ai & AI Magic with Jeremy Howard Aug 17, 2024 · 1 mention

  • ▶ 33:05 Jeremy Howard T-five pre-trained, ah, encoder backbone as a thing you fine-tune, which I think would be really cool.

The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka Jul 5, 2024 · 10 mentions

  • ▶ 3:21 Yi Tay Like, WMT, like, machine translation, and like, complexity, and stuff like that, it's not really about, you know, there wasn't, like, I think, feel short learning, and feel short in context learning came only about, like, you know, when… 2 times in the scene
  • ▶ 14:47 Yi Tay I mean, I mean, it's not that much larger, but from 11 B, the 11 B, uh, T-Five. 2 times in the scene
  • ▶ 1:08:07 Yi Tay Because every time, like, SuiGlu became popular because of the updated T-Five, the T-Five, uh, 1.1 that uses SuiGlu, right? 3 times in the scene
  • ▶ 1:37:28 Yi Tay And then, people make so much big deal about, about, uh, uh, like, uh, you know, trading past Chinchilla Scaling Law, like, oh, Lamao-Doo's the first mall, like, like, like, T-Five base, right, was one trillion tokens, that was really so… 2 times in the scene
  • ▶ 1:49:23 Yi Tay The sparse upcycling paper was mostly vision focused with a little bit of T-Five experiments.

Breaking down the OG GPT Paper by Alec Radford Apr 23, 2024 · 1 mention

  • ▶ 31:23 unnamed speaker Uh, so they are trying to create a multitask format, and this is similar to like what people have used in feature work like T five and whisper, where basically you're trying to, uh, model different tasks, just using tokens and special…

Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit Apr 11, 2024 · 2 mentions

  • ▶ 12:50 unnamed speaker Um, and then just, just to recap as well, like the models you're using back then were like, I don't know, were they like BERT type stuff, or T-five, or I don't know, I don't know what time frame we're talking about here. 2 times in the scene

A Comprehensive Overview of Large Language Models - Latent Space Paper Club Mar 15, 2024 · 5 mentions

  • ▶ 14:36 unnamed speaker Sampling or your autoregressive sampling of tokens to form your target sequence, which is what we've seen in T five and the decoder models, which I think all of us are familiar with things like GPT two, GP three, three, they are all there.
  • ▶ 16:55 unnamed speaker So that's the first key point, because if we looked at T-V, 2 times in the scene
  • ▶ 49:00 unnamed speaker But that just both seems like the same thing, because my understanding of prefixed language modeling was that, oh, we're gonna specify a specific token, for example, like, uh, like a bracket classify, bracket, like, sentiment, sort of like… 2 times in the scene

Beating GPT-4 with Open Source Models - with Michael Royzen of Phind Nov 3, 2023 · 2 mentions

  • ▶ 14:16 Michael Royzen Um, so T-zero worked well because they took the T-five models, which were, um, closer to Chinchilla Optimal, because I think they were trained on also like 300 something billion tokens similar to GPT-three, but the models were much smaller.
  • ▶ 1:04:18 Michael Royzen Um, they implemented streaming generation for T-Five-based models, which we were running at the time, um, up until we switched to GPT in, um,
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.