BERT, every mention
123 scenes across 22 shows · ← back to BERT
Latent Space 110
the MAD Podcast 20
No Priors 14
the a16z Podcast 11
20VC 9
Big Technology 8
A Product Market Fit Show 8
the Y Combinator Startup Podcast 614 more shows
every year every show
Latent Space 110
the MAD Podcast 20
No Priors 14
the a16z Podcast 11
20VC 9
A Product Market Fit Show 8
Big Technology 8
Acquired 6
the Y Combinator Startup Podcast 6
All-In 4
TBPN 4
Cheeky Pint 3
the Neon Show 3
Lenny's Podcast 2
How I Built This 2
In Depth 1
BG2 Pod 1
Innovators & Investors 1
David Senra 1
the Official SaaStr Podcast 1
Invest Like the Best 1
Sourcery 1
Verbatim, from the transcripts: passages where BERT comes up on Latent Space, the MAD Podcast, No Priors, the a16z Podcast, 20VC
Model Mayhem, GPT-6 Astra, Why Nvidia Bought Hugging Face | Diet TBPN
- ▶ 15:34 John Coogan This is pre-GPT-III, and it didn't become a durable consumer business, so in 2018, Google released BERT, which was sort of the first language model, very primitive, but, ah, people were really excited about it, but it wasn't delivered just… 2 times in the scene
Building High-Growth Tech Ventures in the AI Era with Kory Jeffrey of Inovia Capital
- ▶ 32:48 Corey Jeffrey So the BERT models for search and the,
ChatGPT Killed My $50M ARR Business Overnight | Swapnil Jain Observe Al
- ▶ 30:54 unnamed speaker Not ChatGPT, but BERT, Google BERT came out.
How Open Source Became AI's Backbone | Inferact with a16z
- ▶ 3:46 Simon Moe Probably BERT. 5 times in the scene
- ▶ 4:47 Matt Bornstein And so, yeah, so, so look, I mean, um, you know, BERT was an early language model, um, that, uh, you know, newer models are much bigger, much more sophisticated, take up a lot more memory, a lot more compute, and, and, and, um, you know,…
The Man Training GPT, Gemini & Claude Reveals What's Coming Next | Vijay Krishnan, Turing
- ▶ 8:15 Vijay Krishnan The transformer came out of Google, not just transformers, but even this BERT and, uh, this, uh, um, the BERT models very much came out of Google also, uh, which were yet another sort of key building block. 2 times in the scene
🔬Causal Models Need Causal Data - Xaira’s X-Cell model (Bo Wang & Ci Chu)
The Five Year Desert to Product Market Fit & a $5.3BN Valuation with Shiv Rao, Founder @ Abridge · 20VC with Harry Stebbings
Roblox’s David Baszucki Built the Biggest Playground on Earth
- ▶ 1:05:27 David Baszucki So at Roblox for many, many years, we were pushing very early AI BERT models, primitive type models to drive safety.
🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik
- ▶ 54:43 unnamed speaker Like Burt. 3 times in the scene
The history and future of AI at Google, with Sundar Pichai
- ▶ 1:17 Sundar Pichai So Bert and Mom, people underestimate how much, because we measure search quality so religiously, some of the biggest jumps in search quality in that period where search went ahead of everyone else was because of Bert and Mom. 2 times in the scene
How this AI founder is on track to hit $50M ARR just 2 years after launch. | Tarek Alaruri, Co-Fo... · PMF Show
- ▶ 4:43 Tarek Alaruri And so that's what we would do using Fuzzy Burt and some algorithms on top of, uh,
Compliance at scale and why TAM is a distraction with Christina Cacioppo of Vanta
- ▶ 32:34 Christina Cacioppo So the questionnaire is actually a great example, because the questionnaire, so that we tried to build this product in 2018, actually before SOC two, because it's just, it seems easier, actually, but the language models were not good…
The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
- ▶ 22:48 Shawn Wang I often tell a lot of people, uh, that are not steeped in like Google search history that, uh, well, you know, like BERT was like, used like basically immediately inside of Google search, uh, and that improves results a lot, right?
How I grew to $3M ARR in 6 months—& to a $1B valuation in 2 years. | Jay, Founder of Eve Legal · PMF Show
- ▶ 4:27 Jay Madheswaran Real use cases in NLP actually pop up largely from, you know, there's a lot of data labeling companies that end up doing a lot of training in NLP, but also some products started actually relying on the fact that transfer architecture and… 2 times in the scene
Why The Laws of Startup Physics Have Changed | Ben Horowitz Interview · Invest Like The Best
- ▶ 19:36 Ben Horowitz You know, we've had AI going, right, ImageNet was, what, 2012, and then natural language stuff, and Burt and all that was, like,
Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
- ▶ 44:21 Pengchuan Zhang So you can see that language models, kind of, before, kind of, in the birth age, kind of, the language model are not human performance, kind of, SFT, kind of, really imitation learning, really do their job, get to very good performance.
The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
- ▶ 30:12 Loïc Houssier Uh, but, uh, we use Base-Ten to run some, uh, I would say some LAMA, some BERT model for classification.
ChatGPT Turns Three, OpenAI x Thrive, David Sacks vs. The New York Times | Diet TBPN
- ▶ 23:05 unnamed speaker There's no evidence Google has ever trained Gemini on non-TPU hardware going back to pre-GPT models like BERT.
⚡ Inside Google Labs: Building The Gemini Coding Agent — Jed Borovik, Jules + AIE CODE Preview
- ▶ 9:23 unnamed speaker Like, I guess you, you're maybe not too unfamiliar with it because search uses a lot of like machine learned, like black boxy type things, including BERT.
Amjad Masad & Adam D’Angelo: How Far Are We From AGI?
- ▶ 56:34 Amjad Masad Like I was saying denoising, he would take like a single BERT instance and like try to, you know, mask different words and, uh, and just predict like these different tokens.
From Idea to $650M Exit: Lessons in Building AI Startups · Y Combinator
- ▶ 1:49 Jake Heller One of our AI researchers, who is here today, uh, Javed, saw, woo, saw an early application, um, as soon as the Burt paper came out, attention is all you need, et cetera.
Sequoia Partner, David Cahn on Who Wins in AI, Defence & The New $0–$100M Playbook · 20VC with Harry Stebbings
- ▶ 13:18 David Cahn It was the, you know, he had launched this Transformers library.
Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
- ▶ 1:18 Barak Lenz We started the company, you know, just before BERT came out.
Google Part III: The AI Company. Google is amazingly well-positioned... will they win in AI? (Audio) · Acquired
- ▶ 2:01:40 Ben Gilbert Within a year, they build BERT, the large language model.
- ▶ 2:01:52 Ben Gilbert In fact, BERT was one of the first LLMs.
- ▶ 2:02:06 Ben Gilbert They were doing things like BERT and, uh, MUM, this other model,
- ▶ 2:08:14 Ben Gilbert Which we should say is right around the same time as BERT and right around the same time as another large language model based on the Transformer out of here in Seattle, the Allen Institute.
Anthropic Head of Pretraining on Scaling Laws, Compute, and the Future of AI · Y Combinator
- ▶ 5:28 Ankit Gupta There was these, uh, BERT and BERT models that were doing mass language modeling.
Inside Google's Generative AI Reinvention — With Nick Fox and Liz Reid
How This 25-Year-Old Built A $675M Legal AI Startup (With No Legal Experience) · Y Combinator
- ▶ 2:22 Max Junestrand We were playing around in AI and legal way before ChatGPT, and we were using these early models from BERT, coming from Google.
20VC: 15 Term Sheets in 7 Days and Choosing Benchmark | Harvey vs Legora: Who Wins Legal and How to Play When You Have $600M Less Funding | Are AI Models Plateauing Today | Building a 9-9-6 Culture From Stockholm with Max Junestrand
- ▶ 7:45 Max Junestrand And they had been playing around with the early BERT models that came up from Google, and also a, a version of them called SWE BERT,
- ▶ 8:32 Max Junestrand The gap from the BERT models to GPT was the biggest gap ever, and I think the gaps since then have been much more incremental.
- ▶ 1:13:31 Max Junestrand There's a huge difference in Legora's UX and UI since our head of design Bert joined from V seven. 2 times in the scene
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 8:26 Varun Mohan If it's a vision, they should be fine tuning resonant or using something like clip, like the less work they can do the better.
- ▶ 34:41 Varun Mohan You compute the BERT embedding of that prompt. 3 times in the scene
Personalized AI Language Education — with Andrew Hsu, Speak
- ▶ 19:17 Shawn Wang So BERT, maybe?
Information Theory for Language Models: Jack Morris
- ▶ 2:13 Jack Morris At that time I was playing a lot with like BERT and BERT based models. 3 times in the scene
- ▶ 7:39 Jack Morris Like there's a huge difference between the bird size models, which are a hundred and twenty-five million parameters to, to 200.
- ▶ 46:25 Jack Morris And like we, we built our own, but like this idea, we took it and we applied it to model embeddings where instead of zebras and horses, we have like BERT embeddings and GPT embeddings, or like two completely different models with different… 2 times in the scene
- ▶ 1:09:50 Jack Morris And then the second thing was Transformers and BERT, and this Attention is All You Need paper, 2017, the first GPT, 2018, which is web scale pre-training.
Giving New Life to Unstructured Data with LLMs and Agents
- ▶ 4:40 Anant Bhardwaj We, and, and I think the transformer paper, they also released a model called BERT at that time. 4 times in the scene
Inside the Paper That Changed AI Forever - Cohere CEO Aidan Gomez on 2025 Agents
- ▶ 12:09 Aidan Gomez Um, BERT and like the search folks, they figured out how to make use of the transformer super, super fast.
Sundar Pichai, CEO of Alphabet | The All-In Interview
- ▶ 5:08 Sundar Pichai Ah, you know, Transformers drove some of the biggest innovations in search with Bert and Mum.
No Priors Ep. 115 | With Glean Founder and CEO Arvind Jain
- ▶ 3:30 Arvind Jain Like, you know, we, you know, we started with this, uh, BERT, uh, model that Google had put in open domain, which was trained on all of the, all of the internet's, you know, data and knowledge.
20VC: Foundation Models: Who Wins & Who Loses | How Economies and Labour Markets Need to Change in a World of AI | China vs the US in an AI Race: What You Need to Know | Rich Socher, Founder @ You.com
- ▶ 4:35 Richard Socher And that then became Elmo, which became BERT, which is one of the most cited papers still, uh, in the world.
No Priors Ep. 108 | With Abridge Founder and CEO Shiv Rao, MD
- ▶ 9:46 Dr. Shiv Rao When we started with Bird or BioBird, a longformer, a Pegasus, all these other pre-trained models, we, we got to the point where we had a product that worked.
- ▶ 14:03 Dr. Shiv Rao And again, like these weren't even LLMs yet in 2021, we were using like Burt and BioBurt and like,
Amjad Masad on Replit, AI Agents, and the Death of Traditional Software.
- ▶ 27:36 Amjad Masad Thank you, Bert, for doing this show.
He got rejected by 40 VCs & had 6 months of runway—2 years later, he raised $100M from a16z. | Ed... · PMF Show
- ▶ 14:20 Edo Liberty And then, 2017, I think, late 17, a BERT comes out, which is probably the, I think it's, I'm pretty sure this is considered the first true public actual LLM, the transformer models, and so on. 3 times in the scene
- ▶ 57:11 Edo Liberty I mean, we, we, we, in the sense that in the same way that BERT was a sort of like a spark and like, we know something is
Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
- ▶ 18:27 Douwe Kiela Uh, so that was more like a bird style, like mass language modeling.
How To Build The Future: Aravind Srinivas · Y Combinator
- ▶ 33:29 Aravind Srinivas They're not just like, oh, a page rank or like, uh, MapReduce or, uh, you know, all these advances that they made in like visual, like deep learning and, and, and like BERT, Transformers.
Bee AI: The Wearable Ambient Agent
- ▶ 3:56 Ethan Sutin This is before Transformers, no BERT even, like, just RNNs, you couldn't really do any convincing dialogue at all.
Farewell, Chatbots: AI Agents Are Taking Over Customer Service | Mike Murchison, CEO, Ada
- ▶ 9:01 Mike Murchison You know, first with the BERT-based models, and then the true, like, GPTs.
How Emergence Capital Bet Early on Zoom, Salesforce, Veeva—& Now Together.ai · Sourcery with Molly O'Shea
- ▶ 44:21 Gordon Ritter Then ChatGPT comes along November, 22, opens our eyes to what we, you know, scientists knew this technology has been around a while, Bert and others, so it's just, just, and very importantly, they took it all together and brought it to the…
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 6:50 William Beauchamp So I was able to play around with BERT.
- ▶ 12:34 William Beauchamp Then it would prompt whatever, like I just used some external API for like BERT or GPT-II or like it was a very, very small thing.
Going Multi-Product in the Age of AI with CPOs of Webflow, Rubrik, Zoom, and ProductBoard
- ▶ 6:16 Mahesh Ram I think we built the first large language model on BERT that was in, in our space ever in 2017, 20 18.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 1:29:12 Shawn Wang Um, it, it is the, uh, probably the largest scale rollout of transformers yet, um, after Google rolled out BERT for search and, um, and people are using it and it's a three beach, you know, foundation model that's running locally on your…
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 12:18 Loubna Ben Allal Uh, it's a classifier, like a birth model.
- ▶ 27:31 Loubna Ben Allal Um, I think that, like, in AI, we just started, like, with fine-tuning, for example, trying to make BERT work on some specific use cases, and really struggling to do that.
AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
- ▶ 5:43 Dylan Patel Uh, the advent of BERT, which was one of the most, uh, most well-known, most popular transformers before we got to the GPT madness, um, is, has been their, in their production search workloads for years.
No Priors Ep. 94 | With CEO and Founder of Agency Elias Torres
- ▶ 7:41 Elias Torres I knew BERT.
[Paper Club] Embeddings in 2024: OpenAI, Nomic Embed, Jina Embed, cde-small-v1 - with swyx
- ▶ 11:04 unnamed speaker Um, and, uh, I, it's not something that I'm familiar with, and neither, neither am I familiar with, like, a lot of these other, uh, types of, uh, sort of, uh, BERT-based models. 4 times in the scene
[Paper Club] BERT: Bidirectional Encoder Representations from Transformers
- ▶ 0:13 unnamed speaker Um, so why, why did we pick birth today? 4 times in the scene
- ▶ 1:55 unnamed speaker Going through the paper, and then, um, there's a ton of additional material out there on BERT. 9 times in the scene
- ▶ 8:28 unnamed speaker So this is the, the, uh, BERT, like architecture, essentially. 2 times in the scene
- ▶ 17:50 unnamed speaker So I, I didn't look at any of the following BERT papers like Roberta or D Berta for this. 3 times in the scene
- ▶ 22:24 unnamed speaker Um, you can see from this in the, at least when it was released, BERT-Large was state-of-the-art, even beating out, uh, GPT-ONE. 3 times in the scene
- ▶ 25:55 unnamed speaker You basically stick a classifier after the BERT. 14 times in the scene
- ▶ 34:24 unnamed speaker So distilbert is a hugging face, uh, like recreation of Bert that, like, has very comparable performance on, uh, many fewer parameters. 5 times in the scene
- ▶ 41:00 unnamed speaker Uh, someone linked a paper on that economic budget, the uh, 24 hour bird, and then I was also trying to find the paper, apparently Mosaic ML showed how you can do it for 20 dollars, pre-trained bird from scratch now, so, um, 10 times in the scene
[Paper Club] Upcycling Large Language Models into Mixture of Experts
The $4.5B Platform Driving the Open Source AI Revolution | Clem Delangue, CEO, Hugging Face
- ▶ 3:47 Clement Delangue It was GPT, GPT-II, there was, there was BERT, there was ExcelNet, um, obviously Google, uh, kicked off this wave with attention is all you need, which is like the seminal paper for, for Transformers, which is the T in, in ChatGPT.
- ▶ 29:48 Clement Delangue Oh, uh, there's this, uh, this thing that, uh, came up, uh, which is called BERT. 5 times in the scene
- ▶ 32:46 Clement Delangue Even in, kind of, like, the first days of, of Thomas releasing, kind of, like, the, uh, first port of, uh, BERT, I think, uh, contributors, open source contributors started to solve bugs, right, and kind of, like, improve some small things.
Production AI Engineering starts with Evals
- ▶ 17:47 Ankur Goyal And what happened is text, starting with BERT, and then accelerating through and including ChatGPT, just totally cannibalized that. 4 times in the scene
- ▶ 32:14 Ankur Goyal I think the fundamental thing is, prior to BERT, I was, as a traditional software engineer, incapable of participating in the, sort of, what happens behind the scenes in ML development. 2 times in the scene
- ▶ 1:31:58 Ankur Goyal And I got screwed, not necessarily in a bad way, but I sort of felt that by Burt.
Lessons in product leadership and AI strategy from Glean, Google, Amazon, and Slack | Tamar Yehoshua
- ▶ 1:03:08 Tamar Yehoshua Then when, and using AI, so it was an AI search using BERT models and using, uh, vector embeddings in 2019. 2 times in the scene
Is finetuning GPT4o worth it?
- ▶ 4:29 Alistair Pullen I knew what BERT was.
Answer.ai & AI Magic with Jeremy Howard
- ▶ 20:13 Jeremy Howard Like, these people started appearing, as in our collab sections, we have a collab section for, like, collaborating with outsiders, and these people started appearing, there are all these names that I recognize, like, Burt-twenty-four, and… 5 times in the scene
- ▶ 32:11 Swix (Shawn) Uh, just a little bit more on BERT. 4 times in the scene
- ▶ 1:05:08 Jeremy Howard Hopefully we'll be talking about, like, the whole re-interest in BERT that BERT-twenty-four stimulated.
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 2:00 Thomas Scialom And it was something like I started two weeks before BERT.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 3:21 Yi Tay Like, WMT, like, machine translation, and like, complexity, and stuff like that, it's not really about, you know, there wasn't, like, I think, feel short learning, and feel short in context learning came only about, like, you know, when…
- ▶ 1:13:16 Yi Tay People, like, kind of, always associate, like, encoder decoders with, like, like, BERT, or, like, something, like, like, you know, people get confused about these things, right?
He quit his cozy Google job & ignored lean startup advice— then grew to $3M in 1 year. | Arvind J... · PMF Show
- ▶ 23:42 Arvind Jain So when we, when we started, we were actually able to use these BERT family of, uh, language models that Google had actually published in OpenDomain.
Breaking down the OG GPT Paper by Alec Radford
- ▶ 8:13 unnamed speaker Like for example, this paper came out before birth, and even though the original machine learning, the original transformer paper, attention is all you need used machine translation.
AI at Roblox: Revolutionizing Game Creation | Morgan McGuire, Chief Scientist of Roblox
- ▶ 18:46 Morgan McGuire So we adopted a technology called BERT and distill BERT.
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
- ▶ 12:50 unnamed speaker Um, and then just, just to recap as well, like the models you're using back then were like, I don't know, were they like BERT type stuff, or T-five, or I don't know, I don't know what time frame we're talking about here.
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
- ▶ 14:14 unnamed speaker Uh, we've got the encoder models, and examples of this will be things like BERT, where you learn via mass language modeling, which has been covered before.
No Priors Ep. 52 | With Pinecone CEO Edo Liberty
- ▶ 2:32 Edo Liberty Large language models and transformer models like BERT and others started being used by the more mainstream engineering cohorts.
AI and the Future of Law: The 10 Year "Overnight" Success Story · Y Combinator
- ▶ 6:58 Jake Heller As soon as the BERT paper came out, we saw immediate applications to lock.
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 6:24 Michael Royzen Um, and so I go to Hugging Face, and with, um, various encoder models, uh, that were around at the time, I think, uh, I used, I used the standard BERT and also Longformer, 2 times in the scene
Why Google Never Shipped LaMDA Its ChatGPT Predecessor
- ▶ 9:29 Gaurav Nemade I think one, one thing that I remember specifically was the Transformers paper paper came out, but I think it started making a lot of noise when BERT was out. 6 times in the scene
NVIDIA CEO Jensen Huang · Acquired
- ▶ 23:10 Jensen Huang And so my first impression of BERT was really how clever it was.
- ▶ 23:26 Jensen Huang And so obviously, we knew that BERT was going to be a lot larger.