BERT, every mention
45 scenes · ← back to BERT
tap a year for its mentions
every year anyone Jack Morris 7Ankur Goyal 7Jeremy Howard 6Varun Mohan 4Swix (Shawn) 4Yi Tay 3Shawn Wang 3William Beauchamp 2RJ Haneke 2Michael Royzen 2
Verbatim, from the transcripts: the passages where BERT comes up
🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
- ▶ 54:43 unnamed speaker Like Burt. 3 times in the scene
The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
- ▶ 22:48 Shawn Wang I often tell a lot of people, uh, that are not steeped in like Google search history that, uh, well, you know, like BERT was like, used like basically immediately inside of Google search, uh, and that improves results a lot, right?
Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
- ▶ 44:21 Pengchuan Zhang So you can see that language models, kind of, before, kind of, in the birth age, kind of, the language model are not human performance, kind of, SFT, kind of, really imitation learning, really do their job, get to very good performance.
The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
- ▶ 30:12 Loïc Houssier Uh, but, uh, we use Base-Ten to run some, uh, I would say some LAMA, some BERT model for classification.
⚡ Inside Google Labs: Building The Gemini Coding Agent — Jed Borovik, Jules + AIE CODE Preview
- ▶ 9:23 unnamed speaker Like, I guess you, you're maybe not too unfamiliar with it because search uses a lot of like machine learned, like black boxy type things, including BERT.
Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
- ▶ 1:18 Barak Lenz We started the company, you know, just before BERT came out.
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 8:26 Varun Mohan If it's a vision, they should be fine tuning resonant or using something like clip, like the less work they can do the better.
- ▶ 34:41 Varun Mohan You compute the BERT embedding of that prompt. 3 times in the scene
Personalized AI Language Education — with Andrew Hsu, Speak
- ▶ 19:17 Shawn Wang So BERT, maybe?
Information Theory for Language Models: Jack Morris
- ▶ 2:13 Jack Morris At that time I was playing a lot with like BERT and BERT based models. 3 times in the scene
- ▶ 7:39 Jack Morris Like there's a huge difference between the bird size models, which are a hundred and twenty-five million parameters to, to 200.
- ▶ 46:25 Jack Morris And like we, we built our own, but like this idea, we took it and we applied it to model embeddings where instead of zebras and horses, we have like BERT embeddings and GPT embeddings, or like two completely different models with different… 2 times in the scene
- ▶ 1:09:50 Jack Morris And then the second thing was Transformers and BERT, and this Attention is All You Need paper, 2017, the first GPT, 2018, which is web scale pre-training.
Bee AI: The Wearable Ambient Agent
- ▶ 3:56 Ethan Sutin This is before Transformers, no BERT even, like, just RNNs, you couldn't really do any convincing dialogue at all.
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 6:50 William Beauchamp So I was able to play around with BERT.
- ▶ 12:34 William Beauchamp Then it would prompt whatever, like I just used some external API for like BERT or GPT-II or like it was a very, very small thing.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 1:29:12 Shawn Wang Um, it, it is the, uh, probably the largest scale rollout of transformers yet, um, after Google rolled out BERT for search and, um, and people are using it and it's a three beach, you know, foundation model that's running locally on your…
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 12:18 Loubna Ben Allal Uh, it's a classifier, like a birth model.
- ▶ 27:31 Loubna Ben Allal Um, I think that, like, in AI, we just started, like, with fine-tuning, for example, trying to make BERT work on some specific use cases, and really struggling to do that.
[Paper Club] Embeddings in 2024: OpenAI, Nomic Embed, Jina Embed, cde-small-v1 - with swyx
- ▶ 11:04 unnamed speaker Um, and, uh, I, it's not something that I'm familiar with, and neither, neither am I familiar with, like, a lot of these other, uh, types of, uh, sort of, uh, BERT-based models. 4 times in the scene
[Paper Club] BERT: Bidirectional Encoder Representations from Transformers
- ▶ 0:13 unnamed speaker Um, so why, why did we pick birth today? 4 times in the scene
- ▶ 1:55 unnamed speaker Going through the paper, and then, um, there's a ton of additional material out there on BERT. 9 times in the scene
- ▶ 8:28 unnamed speaker So this is the, the, uh, BERT, like architecture, essentially. 2 times in the scene
- ▶ 17:50 unnamed speaker So I, I didn't look at any of the following BERT papers like Roberta or D Berta for this. 3 times in the scene
- ▶ 22:24 unnamed speaker Um, you can see from this in the, at least when it was released, BERT-Large was state-of-the-art, even beating out, uh, GPT-ONE. 3 times in the scene
- ▶ 25:55 unnamed speaker You basically stick a classifier after the BERT. 14 times in the scene
- ▶ 34:24 unnamed speaker So distilbert is a hugging face, uh, like recreation of Bert that, like, has very comparable performance on, uh, many fewer parameters. 5 times in the scene
- ▶ 41:00 unnamed speaker Uh, someone linked a paper on that economic budget, the uh, 24 hour bird, and then I was also trying to find the paper, apparently Mosaic ML showed how you can do it for 20 dollars, pre-trained bird from scratch now, so, um, 10 times in the scene
[Paper Club] Upcycling Large Language Models into Mixture of Experts
Production AI Engineering starts with Evals
- ▶ 17:47 Ankur Goyal And what happened is text, starting with BERT, and then accelerating through and including ChatGPT, just totally cannibalized that. 4 times in the scene
- ▶ 32:14 Ankur Goyal I think the fundamental thing is, prior to BERT, I was, as a traditional software engineer, incapable of participating in the, sort of, what happens behind the scenes in ML development. 2 times in the scene
- ▶ 1:31:58 Ankur Goyal And I got screwed, not necessarily in a bad way, but I sort of felt that by Burt.
Is finetuning GPT4o worth it?
- ▶ 4:29 Alistair Pullen I knew what BERT was.
Answer.ai & AI Magic with Jeremy Howard
- ▶ 20:13 Jeremy Howard Like, these people started appearing, as in our collab sections, we have a collab section for, like, collaborating with outsiders, and these people started appearing, there are all these names that I recognize, like, Burt-twenty-four, and… 5 times in the scene
- ▶ 32:11 Swix (Shawn) Uh, just a little bit more on BERT. 4 times in the scene
- ▶ 1:05:08 Jeremy Howard Hopefully we'll be talking about, like, the whole re-interest in BERT that BERT-twenty-four stimulated.
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 2:00 Thomas Scialom And it was something like I started two weeks before BERT.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 3:21 Yi Tay Like, WMT, like, machine translation, and like, complexity, and stuff like that, it's not really about, you know, there wasn't, like, I think, feel short learning, and feel short in context learning came only about, like, you know, when…
- ▶ 1:13:16 Yi Tay People, like, kind of, always associate, like, encoder decoders with, like, like, BERT, or, like, something, like, like, you know, people get confused about these things, right?
Breaking down the OG GPT Paper by Alec Radford
- ▶ 8:13 unnamed speaker Like for example, this paper came out before birth, and even though the original machine learning, the original transformer paper, attention is all you need used machine translation.
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
- ▶ 12:50 unnamed speaker Um, and then just, just to recap as well, like the models you're using back then were like, I don't know, were they like BERT type stuff, or T-five, or I don't know, I don't know what time frame we're talking about here.
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 14:14 unnamed speaker Uh, we've got the encoder models, and examples of this will be things like BERT, where you learn via mass language modeling, which has been covered before.
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 6:24 Michael Royzen Um, and so I go to Hugging Face, and with, um, various encoder models, uh, that were around at the time, I think, uh, I used, I used the standard BERT and also Longformer, 2 times in the scene