BERT

4 statements across 4 episodes · 3 bullish · 0 bearish · 4 people on the record · first statement Jan 1, 2025 by Shawn Wang · said 110 times in 27 episodes since 2023 · across every show →

Mentions by year

brought up most by Jack Morris (7), Ankur Goyal (7), Jeremy Howard (6), Varun Mohan (4), Swix (Shawn) (4), Yi Tay (3), Shawn Wang (3), William Beauchamp (2)

tap a year for its mentions
00508100152023202420252026episodesmentions
08152023202420252026episodes it came up in
0047.58152023202420252026episodesmentions per episode
2026 7 mentions in 4 episodes 2 per episode
2025 20 mentions in 10 episodes 2 per episode
2024 81 mentions in 12 episodes 7 per episode
2023 2 mentions in 1 episode

every mention, scene by scene, with the transcript →

Everything said about BERT, oldest first

Jan 1, 2025 positive
Assertion Not checkable as stated
Swyx: Apple Intelligence is the largest transformer rollout since Google's BERT
“It is the probably the largest scale rollout of transformers yet after Google rolled out BERT for search”
Shawn Wang Jan 1, 2025 ▶ 1:29:12 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jul 2, 2025 positive
Assertion Supported
Morris: CycleGAN mapping aligns disparate model embeddings without paired data
“We took it and we applied it to model embeddings where instead of zebras and horses, we have like BERT embeddings and GPT embeddings, or like two completely different models with different architectures. So I think these are GTR, which is a T five based retrie…”
Jack Morris Jul 2, 2025 ▶ 46:25 Information Theory for Language Models: Jack Morris
Jul 28, 2025 positive
Insight
Mohan: Enterprises should fine-tune off-the-shelf models over custom architectures
“For a vast majority of enterprises, they should probably be using something off the shelf, fine tuning BERT models. If it's a vision, they should be fine tuning resonant or using something like clip, like the less work they can do the better.”
Varun Mohan Jul 28, 2025 ▶ 8:16 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Dec 11, 2025
Disclosure
Superhuman uses Baseten to run LLaMA and BERT classification models
“We use Base-Ten to run some I would say some LAMA, some BERT model for classification.”
Loïc Houssier Dec 11, 2025 ▶ 30:12 The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.