T5, every mention
21 scenes · ← back to T5
tap a year for its mentions
every year anyone Yi Tay 11Rishabh Agarwal 2Michael Royzen 2Jeremy Howard 1Jack Morris 1Ethan He 1
Verbatim, from the transcripts: the passages where T5 comes up
Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
- ▶ 4:16 unnamed speaker Superficially, I see some, you know, in your UL II and T-Five work, um, some,
- ▶ 1:10:17 Yi Tay So we wanted, at that time LM, we were still using T-Five models at that time.
- ▶ 1:14:23 unnamed speaker And I think like, it's also somewhat emergent in the sense that when you were using T-Five, you just couldn't actually add that much value on top of a normal BM-T-Five retrieval technique. 3 times in the scene
Information Theory for Language Models: Jack Morris
- ▶ 46:39 Jack Morris So I think these are GTR, which is a T five based retrieval model and GTE, which is based on BERT.
The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
- ▶ 26:31 Rishabh Agarwal So if you look at the T-five base, two-fifty million scenario on GSM-HK,
- ▶ 27:22 Rishabh Agarwal Like, look at the T five base model.
[Paper Club] BERT: Bidirectional Encoder Representations from Transformers
- ▶ 1:26 unnamed speaker Bart and T five are good, but like, let's look at a host.
[Paper Club] Upcycling Large Language Models into Mixture of Experts
Answer.ai & AI Magic with Jeremy Howard
- ▶ 33:05 Jeremy Howard T-five pre-trained, ah, encoder backbone as a thing you fine-tune, which I think would be really cool.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 3:21 Yi Tay Like, WMT, like, machine translation, and like, complexity, and stuff like that, it's not really about, you know, there wasn't, like, I think, feel short learning, and feel short in context learning came only about, like, you know, when… 2 times in the scene
- ▶ 14:47 Yi Tay I mean, I mean, it's not that much larger, but from 11 B, the 11 B, uh, T-Five. 2 times in the scene
- ▶ 1:08:07 Yi Tay Because every time, like, SuiGlu became popular because of the updated T-Five, the T-Five, uh, 1.1 that uses SuiGlu, right? 3 times in the scene
- ▶ 1:37:28 Yi Tay And then, people make so much big deal about, about, uh, uh, like, uh, you know, trading past Chinchilla Scaling Law, like, oh, Lamao-Doo's the first mall, like, like, like, T-Five base, right, was one trillion tokens, that was really so… 2 times in the scene
- ▶ 1:49:23 Yi Tay The sparse upcycling paper was mostly vision focused with a little bit of T-Five experiments.
Breaking down the OG GPT Paper by Alec Radford
- ▶ 31:23 unnamed speaker Uh, so they are trying to create a multitask format, and this is similar to like what people have used in feature work like T five and whisper, where basically you're trying to, uh, model different tasks, just using tokens and special…
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
- ▶ 12:50 unnamed speaker Um, and then just, just to recap as well, like the models you're using back then were like, I don't know, were they like BERT type stuff, or T-five, or I don't know, I don't know what time frame we're talking about here. 2 times in the scene
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 14:36 unnamed speaker Sampling or your autoregressive sampling of tokens to form your target sequence, which is what we've seen in T five and the decoder models, which I think all of us are familiar with things like GPT two, GP three, three, they are all there.
- ▶ 16:55 unnamed speaker So that's the first key point, because if we looked at T-V, 2 times in the scene
- ▶ 49:00 unnamed speaker But that just both seems like the same thing, because my understanding of prefixed language modeling was that, oh, we're gonna specify a specific token, for example, like, uh, like a bracket classify, bracket, like, sentiment, sort of like… 2 times in the scene
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 14:16 Michael Royzen Um, so T-zero worked well because they took the T-five models, which were, um, closer to Chinchilla Optimal, because I think they were trained on also like 300 something billion tokens similar to GPT-three, but the models were much smaller.
- ▶ 1:04:18 Michael Royzen Um, they implemented streaming generation for T-Five-based models, which we were running at the time, um, up until we switched to GPT in, um,