BERT

product on 22 shows · 17 statements across 17 episodes · said 217 times in 93 episodes since 2017

Latent Space 110 the MAD Podcast 20 No Priors 14 the a16z Podcast 11 20VC 9 A Product Market Fit Show 8 Big Technology 8 Acquired 6 the Y Combinator Startup Podcast 6 All-In 4 TBPN 4 Cheeky Pint 3 the Neon Show 3 Lenny's Podcast 2 How I Built This 2 In Depth 1 BG2 Pod 1 Innovators & Investors 1 David Senra 1 the Official SaaStr Podcast 1 Invest Like the Best 1 Sourcery 1

Mentions by year, every show

tap a year for its mentions
005020100402017201820192020202120222023202420252026episodesmentions
020402017201820192020202120222023202420252026episodes it came up in
002.5205402017201820192020202120222023202420252026episodesmentions per episode

Latent Space 110the MAD Podcast 20No Priors 14the a16z Podcast 1120VC 9Big Technology 8A Product Market Fit Show 8the Y Combinator Startup Podcast 614 more shows

2026 28 mentions in 16 episodes 2 per episode
2025 55 mentions in 32 episodes 2 per episode
2024 95 mentions in 19 episodes 5 per episode
2023 29 mentions in 17 episodes 2 per episode
2022 1 mention in 1 episode
2021 4 mentions in 3 episodes 1 per episode
2020 1 mention in 1 episode
2019 3 mentions in 3 episodes 1 per episode
2017 1 mention in 1 episode

every mention on every show, scene by scene, with the transcript →

17 statements about BERT, every show

a16z Assertion Not checkable as stated
Simon Moe: BERT was the first model requiring GPUs for efficient inference
“Probably BERT. And before that, it was like ResNet for computation, like images, computer vision classification. So ResNet already need to run on NVIDIA K-eighty, which is kind of one of the first SEU on AWS and other places. And, but way over, but even at thi…”
Simon Moe Aug 5, 2026 ▶ 3:46 How Open Source Became AI's Backbone | Inferact with a16z
NEON SHOW Assertion Not checkable as stated
Krishnan: Nobody expected Transformers and BERT to unlock emergent reasoning or AGI
“At least at the time Transformers came out, BERT came out and all that, nobody really thought this was a this was a path to emergent reasoning capabilities, a path to AGI itself.”
Vijay Krishnan Jul 31, 2026 ▶ 9:37 The Man Training GPT, Gemini & Claude Reveals What's Coming Next | Vijay Krishnan, Turing
DAVID SENRA Assertion Supported
Baszucki: Roblox used early AI BERT models for safety for years
“So at Roblox for many, many years, we were pushing very early AI BERT models, primitive type models to drive safety.”
David Baszucki Apr 26, 2026 ▶ 1:05:27 Roblox’s David Baszucki Built the Biggest Playground on Earth
CHEEKY PINT Assertion Supported
Pichai: BERT and MUM drove Google Search's largest quality leaps
“BERT and MUM, people underestimate how much, because we measure search quality so religiously, some of the biggest jumps in search quality in that period where search went ahead of everyone else was because of BERT and MUM. We built transformers and used it im…”
Sundar Pichai Apr 7, 2026 ▶ 1:17 The history and future of AI at Google, with Sundar Pichai
PRODUCT MARKET FIT Assertion Partly supported
Madheswaran: BERT cut entity extraction training data requirements to ten documents
“So, right, so you went from needing to train thousands about thousands of documents to 10 documents, maybe, to extract things like dollar amounts from invoice totals simple things like that.”
Jay Madheswaran Feb 10, 2026 ▶ 6:02 How I grew to $3M ARR in 6 months—& to a $1B valuation in 2 years. | Jay, Founder of Eve Legal · PMF Show
LATENT SPACE Disclosure
Superhuman uses Baseten to run LLaMA and BERT classification models
“We use Base-Ten to run some I would say some LAMA, some BERT model for classification.”
Loïc Houssier Dec 11, 2025 ▶ 30:12 The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
BIG TECHNOLOGY Assertion Supported
Google adopted BERT and MUM early to build AI search capabilities
“Generative AI is an example we were doing, you know, we had Bert, we had mum as ranking improvements way back in the early days, really as the technology was new, that enables us to get our feet wet with it.”
Nick Fox Aug 27, 2025 ▶ 2:33 Inside Google's Generative AI Reinvention — With Nick Fox and Liz Reid
Mohan: Enterprises should fine-tune off-the-shelf models over custom architectures
“For a vast majority of enterprises, they should probably be using something off the shelf, fine tuning BERT models. If it's a vision, they should be fine tuning resonant or using something like clip, like the less work they can do the better.”
Varun Mohan Jul 28, 2025 ▶ 8:16 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
LATENT SPACE Assertion Supported
Morris: CycleGAN mapping aligns disparate model embeddings without paired data
“We took it and we applied it to model embeddings where instead of zebras and horses, we have like BERT embeddings and GPT embeddings, or like two completely different models with different architectures. So I think these are GTR, which is a T five based retrie…”
Jack Morris Jul 2, 2025 ▶ 46:25 Information Theory for Language Models: Jack Morris
NO PRIORS Disclosure
Glean Built Custom Enterprise Embeddings on Top of BERT
“We started with this BERT model that Google had put in open domain, which was trained on all of the internet's, you know, data and knowledge. And we would then take those models and then for every customer of ours, we'd actually build custom embeddings, you kn…”
Arvind Jain May 15, 2025 ▶ 3:30 No Priors Ep. 115 | With Glean Founder and CEO Arvind Jain
SAASTR Assertion Not checkable as stated
Ram: Solvvy built conversational AI's first BERT model in 2017–2018
“I think we built the first large language model on BERT that was in, in our space ever in 2017, 20 18.”
Mahesh Ram Jan 17, 2025 ▶ 6:16 Going Multi-Product in the Age of AI with CPOs of Webflow, Rubrik, Zoom, and ProductBoard
LATENT SPACE Assertion Not checkable as stated
Swyx: Apple Intelligence is the largest transformer rollout since Google's BERT
“It is the probably the largest scale rollout of transformers yet after Google rolled out BERT for search”
Shawn Wang Jan 1, 2025 ▶ 1:29:12 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
BG2 Assertion Supported
Patel: Google has run transformers in Search workloads since 2018
“Google was running transformers even in their search workload since 2018, 2019. The advent of BERT, which was one of the most most well-known, most popular transformers before we got to the GPT madness is, has been their, in their production search workloads f…”
Dylan Patel Dec 23, 2024 ▶ 5:36 AI Semiconductor Landscape feat. Dylan Patel | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
NO PRIORS Assertion Not checkable as stated
Liberty: Mainstream Engineers Were Already Adopting BERT by 2019
“In 2019, the earthquake had already happened. Deep learning models and so on have already been grappled with. Large language models and transformer models like BERT and others started being used by the more mainstream engineering cohorts.”
Edo Liberty Feb 22, 2024 ▶ 2:28 No Priors Ep. 52 | With Pinecone CEO Edo Liberty
BIG TECHNOLOGY Assertion Not checkable as stated
Nemade: BERT release triggered Google-wide rush to adopt transformers
“I think one thing that I remember specifically was the Transformers paper paper came out, but I think it started making a lot of noise when BERT was out. Like when people started seeing what BERT could essentially do. That was kind of Like a switch that went o…”
Gaurav Nemade Oct 30, 2023 ▶ 9:29 Why Google Never Shipped LaMDA Its ChatGPT Predecessor
NO PRIORS Assertion Supported
Guu: BERT Demonstrated Vast Unprogrammed World Knowledge Purely From Pre-Training
“I think one of the things that became very apparent early on when playing with BERT was unlike all the prior generations of models, it had a large amount of world knowledge that we didn't deliberately encode into it. It wasn't in, you know, the fine tuning dat…”
Kelvin Guu May 4, 2023 ▶ 1:23 No Priors Ep. 15 | With Kelvin Guu, Staff Research Scientist, Google Brain
MAD Assertion Supported
Domingos: State-of-the-art AI models have more connections than many animals
“Having said that, if you look at the number of connections that the state-of-the-art machine learning systems for some of these problems have, they're more than many animals. So we're actually at the point where the, you know, they have hundreds of millions or…”
Pedro Domingos Oct 22, 2019 ▶ 29:01 Fireside Chat: Pedro Domingos, Head of Machine Learning, DE Shaw (FirstMark's Data Driven NYC)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.