BERT, every mention
15 scenes (2025) · ← back to BERT
tap a year for its mentions
every year 2025 anyone Jack Morris 7Ankur Goyal 7Jeremy Howard 6Varun Mohan 4Swix (Shawn) 4Yi Tay 3Shawn Wang 3William Beauchamp 2RJ Haneke 2Michael Royzen 2
Verbatim, from the transcripts: the passages where BERT comes up
SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
- ▶ 44:21 Pengchuan Zhang So you can see that language models, kind of, before, kind of, in the birth age, kind of, the language model are not human performance, kind of, SFT, kind of, really imitation learning, really do their job, get to very good performance.
The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
- ▶ 30:12 Loïc Houssier Uh, but, uh, we use Base-Ten to run some, uh, I would say some LAMA, some BERT model for classification.
⚡ Inside Google Labs: Building The Gemini Coding Agent — Jed Borovik, Jules + AIE CODE Preview
- ▶ 9:23 unnamed speaker Like, I guess you, you're maybe not too unfamiliar with it because search uses a lot of like machine learned, like black boxy type things, including BERT.
Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
- ▶ 1:18 Barak Lenz We started the company, you know, just before BERT came out.
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 8:26 Varun Mohan If it's a vision, they should be fine tuning resonant or using something like clip, like the less work they can do the better.
- ▶ 34:41 Varun Mohan You compute the BERT embedding of that prompt. 3 times in the scene
Personalized AI Language Education — with Andrew Hsu, Speak
- ▶ 19:17 Shawn Wang So BERT, maybe?
Information Theory for Language Models: Jack Morris
- ▶ 2:13 Jack Morris At that time I was playing a lot with like BERT and BERT based models. 3 times in the scene
- ▶ 7:39 Jack Morris Like there's a huge difference between the bird size models, which are a hundred and twenty-five million parameters to, to 200.
- ▶ 46:25 Jack Morris And like we, we built our own, but like this idea, we took it and we applied it to model embeddings where instead of zebras and horses, we have like BERT embeddings and GPT embeddings, or like two completely different models with different… 2 times in the scene
- ▶ 1:09:50 Jack Morris And then the second thing was Transformers and BERT, and this Attention is All You Need paper, 2017, the first GPT, 2018, which is web scale pre-training.
Bee AI: The Wearable Ambient Agent
- ▶ 3:56 Ethan Sutin This is before Transformers, no BERT even, like, just RNNs, you couldn't really do any convincing dialogue at all.
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 6:50 William Beauchamp So I was able to play around with BERT.
- ▶ 12:34 William Beauchamp Then it would prompt whatever, like I just used some external API for like BERT or GPT-II or like it was a very, very small thing.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 1:29:12 Shawn Wang Um, it, it is the, uh, probably the largest scale rollout of transformers yet, um, after Google rolled out BERT for search and, um, and people are using it and it's a three beach, you know, foundation model that's running locally on your…