BERT, every mention
25 scenes (2024) · ← back to BERT
tap a year for its mentions
every year 2024 anyone Jack Morris 7Ankur Goyal 7Jeremy Howard 6Varun Mohan 4Swix (Shawn) 4Yi Tay 3Shawn Wang 3William Beauchamp 2RJ Haneke 2Michael Royzen 2
Verbatim, from the transcripts: the passages where BERT comes up
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 12:18 Loubna Ben Allal Uh, it's a classifier, like a birth model.
- ▶ 27:31 Loubna Ben Allal Um, I think that, like, in AI, we just started, like, with fine-tuning, for example, trying to make BERT work on some specific use cases, and really struggling to do that.
[Paper Club] Embeddings in 2024: OpenAI, Nomic Embed, Jina Embed, cde-small-v1 - with swyx
- ▶ 11:04 unnamed speaker Um, and, uh, I, it's not something that I'm familiar with, and neither, neither am I familiar with, like, a lot of these other, uh, types of, uh, sort of, uh, BERT-based models. 4 times in the scene
[Paper Club] BERT: Bidirectional Encoder Representations from Transformers
- ▶ 0:13 unnamed speaker Um, so why, why did we pick birth today? 4 times in the scene
- ▶ 1:55 unnamed speaker Going through the paper, and then, um, there's a ton of additional material out there on BERT. 9 times in the scene
- ▶ 8:28 unnamed speaker So this is the, the, uh, BERT, like architecture, essentially. 2 times in the scene
- ▶ 17:50 unnamed speaker So I, I didn't look at any of the following BERT papers like Roberta or D Berta for this. 3 times in the scene
- ▶ 22:24 unnamed speaker Um, you can see from this in the, at least when it was released, BERT-Large was state-of-the-art, even beating out, uh, GPT-ONE. 3 times in the scene
- ▶ 25:55 unnamed speaker You basically stick a classifier after the BERT. 14 times in the scene
- ▶ 34:24 unnamed speaker So distilbert is a hugging face, uh, like recreation of Bert that, like, has very comparable performance on, uh, many fewer parameters. 5 times in the scene
- ▶ 41:00 unnamed speaker Uh, someone linked a paper on that economic budget, the uh, 24 hour bird, and then I was also trying to find the paper, apparently Mosaic ML showed how you can do it for 20 dollars, pre-trained bird from scratch now, so, um, 10 times in the scene
[Paper Club] Upcycling Large Language Models into Mixture of Experts
Production AI Engineering starts with Evals
- ▶ 17:47 Ankur Goyal And what happened is text, starting with BERT, and then accelerating through and including ChatGPT, just totally cannibalized that. 4 times in the scene
- ▶ 32:14 Ankur Goyal I think the fundamental thing is, prior to BERT, I was, as a traditional software engineer, incapable of participating in the, sort of, what happens behind the scenes in ML development. 2 times in the scene
- ▶ 1:31:58 Ankur Goyal And I got screwed, not necessarily in a bad way, but I sort of felt that by Burt.
Is finetuning GPT4o worth it?
- ▶ 4:29 Alistair Pullen I knew what BERT was.
Answer.ai & AI Magic with Jeremy Howard
- ▶ 20:13 Jeremy Howard Like, these people started appearing, as in our collab sections, we have a collab section for, like, collaborating with outsiders, and these people started appearing, there are all these names that I recognize, like, Burt-twenty-four, and… 5 times in the scene
- ▶ 32:11 Swix (Shawn) Uh, just a little bit more on BERT. 4 times in the scene
- ▶ 1:05:08 Jeremy Howard Hopefully we'll be talking about, like, the whole re-interest in BERT that BERT-twenty-four stimulated.
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 2:00 Thomas Scialom And it was something like I started two weeks before BERT.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 3:21 Yi Tay Like, WMT, like, machine translation, and like, complexity, and stuff like that, it's not really about, you know, there wasn't, like, I think, feel short learning, and feel short in context learning came only about, like, you know, when…
- ▶ 1:13:16 Yi Tay People, like, kind of, always associate, like, encoder decoders with, like, like, BERT, or, like, something, like, like, you know, people get confused about these things, right?
Breaking down the OG GPT Paper by Alec Radford
- ▶ 8:13 unnamed speaker Like for example, this paper came out before birth, and even though the original machine learning, the original transformer paper, attention is all you need used machine translation.
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
- ▶ 12:50 unnamed speaker Um, and then just, just to recap as well, like the models you're using back then were like, I don't know, were they like BERT type stuff, or T-five, or I don't know, I don't know what time frame we're talking about here.
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 14:14 unnamed speaker Uh, we've got the encoder models, and examples of this will be things like BERT, where you learn via mass language modeling, which has been covered before.