Oct 25, 2024 · 1h 13m · latent-space

How NotebookLM Was Made

Raiza Martin · 37m spoken Usama Shafqat · 13m spoken Shawn Wang · 10m spoken Alessio Fanelli · 5m spoken NotebookLM Deep Dive · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Latent Space, Google Labs product lead Raiza Martin and AI engineer Usama Shafqat join hosts Swix and Alessio Fanelli to discuss the inception, architectural design, and viral breakout of NotebookLM's 'Deep Dive' conversational audio feature.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.1% of the talking time here. How this is scored →

The hosts as informed peer 4.7 Guest teaching 5.0 Guest disagreement 1.0 The hosts pushing back 0.9
05100:0015:0030:0045:001:00:000:52–3:53 · The hosts as informed peer 3/10 Introductions and Google Labs Origins Hosts warmly welcome the guests and prompt them on their background in Google Labs. Raiza and Usama explain the structure of Labs and Area 120, educating the hosts on their org history in a purely friendly opening exchange.3:54–7:33 · The hosts as informed peer 5/10 From Talk to Small Corpus to Tailwind Swyx brings up early context lengths and RAG techniques when discussing Talk to Small Corpus and Project Tailwind. Raiza details the early user studies with adult learners and the natural emergence of document summarization requests.7:33–12:21 · The hosts as informed peer 4/10 Discord Growth, Expansion, and Unlaunching Features Swyx asks about feature flagging and PM methodologies, prompting Raiza to discuss using Mendel experiments and unlaunching unused complex features like passage transformation. The conversation is collaborative and focused on product development practices.12:21–18:00 · The hosts as informed peer 6/10 Architecture, Gemini 1.5, and Dual Personas Alessio asks technical questions about whether dialogue is generated via two independent system prompts or one multi-turn pass. Usama and Raiza explain how Gemini 1.5 long context pairs with DeepMind TTS and dual personas to prevent robotic readouts.18:00–23:09 · The hosts as informed peer 5/10 Steven Johnson and Expert Thinking Workflows Swyx draws parallels to Andrew Mason at OpenAI, while Raiza details how she silently observed author Steven Johnson's reasoning workflow to model expert thinking patterns directly into the product UX.23:09–27:54 · The hosts as informed peer 5/10 Multimodal Sources, Embeddings, and Audio Tone Alessio queries multimodal ingestion such as handwritten Marie Curie notes and transcribing audio with emotional cues. Raiza and Usama explain the current lossy transcription pipeline and why Gemini requires multimodal embeddings for massive multi-document inputs.27:54–30:23 · The hosts as informed peer 3/10 Viral Deep Dive Use Cases on Social Media The hosts and guests swap stories about viral social media use cases, from Karpathy's Wikipedia podcast to Googlers running Q3 performance reviews through Deep Dive. The dynamic is lighthearted and mutual celebration.30:23–37:30 · The hosts as informed peer 5/10 Evals, Dogfooding, and Measuring Audio Quality Swyx presses on standard engineering advice to establish early quantitative eval benchmarks for human audio quality. Raiza and Usama explain why they rejected early formal benchmarks in favor of highly opinionated internal dogfooding and team listening sessions.37:30–43:13 · The hosts as informed peer 5/10 Pacing, Dialogue Tension, and Speech Dynamics Swyx asks if professional speech experts or comedy writers were consulted to structure dialogue pacing. Usama refutes the appeal to domain authority by sharing an anecdote about linguistics graduates and breaking down narrative tension techniques.43:13–47:03 · The hosts as informed peer 4/10 Contextual Humor and the Chicken Paper The discussion covers emergent humor illustrated by the viral Doug Zongker Chicken Paper deep dive. Usama and Raiza emphasize that humor is contextual rather than explicitly prompted.47:04–53:35 · The hosts as informed peer 6/10 Hyper-Personalized Media and Product Craft Swyx contrasts compound AI system architecture against monolithic prompts, asking how the team chooses between explicit engineering and LLM delegation. Raiza argues that building AI products is an opinionated craft rather than just an engineering exercise.53:35–1:00:59 · The hosts as informed peer 5/10 Roadmap: APIs, Languages, and Resisting Knobs Swyx remarks on how rare it is for PMs to resist exposing knobs and settings to users. Raiza explains her philosophy of maintaining magical one-click simplicity despite loud developer demands for parameter sliders.1:01:00–1:08:27 · The hosts as informed peer 5/10 NotebookLM Workspace Future and Real-Time Chat Alessio and Swyx explore workspace competition like Claude Artifacts and OpenAI Realtime API. Raiza and Usama discuss balancing UI craft against fast-following and give final principles for AI engineers pushing model boundaries.0:52–3:53 · Guest teaching 4/10 Introductions and Google Labs Origins Hosts warmly welcome the guests and prompt them on their background in Google Labs. Raiza and Usama explain the structure of Labs and Area 120, educating the hosts on their org history in a purely friendly opening exchange.3:54–7:33 · Guest teaching 5/10 From Talk to Small Corpus to Tailwind Swyx brings up early context lengths and RAG techniques when discussing Talk to Small Corpus and Project Tailwind. Raiza details the early user studies with adult learners and the natural emergence of document summarization requests.7:33–12:21 · Guest teaching 4/10 Discord Growth, Expansion, and Unlaunching Features Swyx asks about feature flagging and PM methodologies, prompting Raiza to discuss using Mendel experiments and unlaunching unused complex features like passage transformation. The conversation is collaborative and focused on product development practices.12:21–18:00 · Guest teaching 6/10 Architecture, Gemini 1.5, and Dual Personas Alessio asks technical questions about whether dialogue is generated via two independent system prompts or one multi-turn pass. Usama and Raiza explain how Gemini 1.5 long context pairs with DeepMind TTS and dual personas to prevent robotic readouts.18:00–23:09 · Guest teaching 5/10 Steven Johnson and Expert Thinking Workflows Swyx draws parallels to Andrew Mason at OpenAI, while Raiza details how she silently observed author Steven Johnson's reasoning workflow to model expert thinking patterns directly into the product UX.23:09–27:54 · Guest teaching 5/10 Multimodal Sources, Embeddings, and Audio Tone Alessio queries multimodal ingestion such as handwritten Marie Curie notes and transcribing audio with emotional cues. Raiza and Usama explain the current lossy transcription pipeline and why Gemini requires multimodal embeddings for massive multi-document inputs.27:54–30:23 · Guest teaching 3/10 Viral Deep Dive Use Cases on Social Media The hosts and guests swap stories about viral social media use cases, from Karpathy's Wikipedia podcast to Googlers running Q3 performance reviews through Deep Dive. The dynamic is lighthearted and mutual celebration.30:23–37:30 · Guest teaching 6/10 Evals, Dogfooding, and Measuring Audio Quality Swyx presses on standard engineering advice to establish early quantitative eval benchmarks for human audio quality. Raiza and Usama explain why they rejected early formal benchmarks in favor of highly opinionated internal dogfooding and team listening sessions.37:30–43:13 · Guest teaching 6/10 Pacing, Dialogue Tension, and Speech Dynamics Swyx asks if professional speech experts or comedy writers were consulted to structure dialogue pacing. Usama refutes the appeal to domain authority by sharing an anecdote about linguistics graduates and breaking down narrative tension techniques.43:13–47:03 · Guest teaching 5/10 Contextual Humor and the Chicken Paper The discussion covers emergent humor illustrated by the viral Doug Zongker Chicken Paper deep dive. Usama and Raiza emphasize that humor is contextual rather than explicitly prompted.47:04–53:35 · Guest teaching 5/10 Hyper-Personalized Media and Product Craft Swyx contrasts compound AI system architecture against monolithic prompts, asking how the team chooses between explicit engineering and LLM delegation. Raiza argues that building AI products is an opinionated craft rather than just an engineering exercise.53:35–1:00:59 · Guest teaching 6/10 Roadmap: APIs, Languages, and Resisting Knobs Swyx remarks on how rare it is for PMs to resist exposing knobs and settings to users. Raiza explains her philosophy of maintaining magical one-click simplicity despite loud developer demands for parameter sliders.1:01:00–1:08:27 · Guest teaching 5/10 NotebookLM Workspace Future and Real-Time Chat Alessio and Swyx explore workspace competition like Claude Artifacts and OpenAI Realtime API. Raiza and Usama discuss balancing UI craft against fast-following and give final principles for AI engineers pushing model boundaries.0:52–3:53 · Guest disagreement 0/10 Introductions and Google Labs Origins Hosts warmly welcome the guests and prompt them on their background in Google Labs. Raiza and Usama explain the structure of Labs and Area 120, educating the hosts on their org history in a purely friendly opening exchange.3:54–7:33 · Guest disagreement 1/10 From Talk to Small Corpus to Tailwind Swyx brings up early context lengths and RAG techniques when discussing Talk to Small Corpus and Project Tailwind. Raiza details the early user studies with adult learners and the natural emergence of document summarization requests.7:33–12:21 · Guest disagreement 1/10 Discord Growth, Expansion, and Unlaunching Features Swyx asks about feature flagging and PM methodologies, prompting Raiza to discuss using Mendel experiments and unlaunching unused complex features like passage transformation. The conversation is collaborative and focused on product development practices.12:21–18:00 · Guest disagreement 1/10 Architecture, Gemini 1.5, and Dual Personas Alessio asks technical questions about whether dialogue is generated via two independent system prompts or one multi-turn pass. Usama and Raiza explain how Gemini 1.5 long context pairs with DeepMind TTS and dual personas to prevent robotic readouts.18:00–23:09 · Guest disagreement 0/10 Steven Johnson and Expert Thinking Workflows Swyx draws parallels to Andrew Mason at OpenAI, while Raiza details how she silently observed author Steven Johnson's reasoning workflow to model expert thinking patterns directly into the product UX.23:09–27:54 · Guest disagreement 1/10 Multimodal Sources, Embeddings, and Audio Tone Alessio queries multimodal ingestion such as handwritten Marie Curie notes and transcribing audio with emotional cues. Raiza and Usama explain the current lossy transcription pipeline and why Gemini requires multimodal embeddings for massive multi-document inputs.27:54–30:23 · Guest disagreement 0/10 Viral Deep Dive Use Cases on Social Media The hosts and guests swap stories about viral social media use cases, from Karpathy's Wikipedia podcast to Googlers running Q3 performance reviews through Deep Dive. The dynamic is lighthearted and mutual celebration.30:23–37:30 · Guest disagreement 2/10 Evals, Dogfooding, and Measuring Audio Quality Swyx presses on standard engineering advice to establish early quantitative eval benchmarks for human audio quality. Raiza and Usama explain why they rejected early formal benchmarks in favor of highly opinionated internal dogfooding and team listening sessions.37:30–43:13 · Guest disagreement 2/10 Pacing, Dialogue Tension, and Speech Dynamics Swyx asks if professional speech experts or comedy writers were consulted to structure dialogue pacing. Usama refutes the appeal to domain authority by sharing an anecdote about linguistics graduates and breaking down narrative tension techniques.43:13–47:03 · Guest disagreement 1/10 Contextual Humor and the Chicken Paper The discussion covers emergent humor illustrated by the viral Doug Zongker Chicken Paper deep dive. Usama and Raiza emphasize that humor is contextual rather than explicitly prompted.47:04–53:35 · Guest disagreement 1/10 Hyper-Personalized Media and Product Craft Swyx contrasts compound AI system architecture against monolithic prompts, asking how the team chooses between explicit engineering and LLM delegation. Raiza argues that building AI products is an opinionated craft rather than just an engineering exercise.53:35–1:00:59 · Guest disagreement 2/10 Roadmap: APIs, Languages, and Resisting Knobs Swyx remarks on how rare it is for PMs to resist exposing knobs and settings to users. Raiza explains her philosophy of maintaining magical one-click simplicity despite loud developer demands for parameter sliders.1:01:00–1:08:27 · Guest disagreement 1/10 NotebookLM Workspace Future and Real-Time Chat Alessio and Swyx explore workspace competition like Claude Artifacts and OpenAI Realtime API. Raiza and Usama discuss balancing UI craft against fast-following and give final principles for AI engineers pushing model boundaries.0:52–3:53 · The hosts pushing back 0/10 Introductions and Google Labs Origins Hosts warmly welcome the guests and prompt them on their background in Google Labs. Raiza and Usama explain the structure of Labs and Area 120, educating the hosts on their org history in a purely friendly opening exchange.3:54–7:33 · The hosts pushing back 1/10 From Talk to Small Corpus to Tailwind Swyx brings up early context lengths and RAG techniques when discussing Talk to Small Corpus and Project Tailwind. Raiza details the early user studies with adult learners and the natural emergence of document summarization requests.7:33–12:21 · The hosts pushing back 1/10 Discord Growth, Expansion, and Unlaunching Features Swyx asks about feature flagging and PM methodologies, prompting Raiza to discuss using Mendel experiments and unlaunching unused complex features like passage transformation. The conversation is collaborative and focused on product development practices.12:21–18:00 · The hosts pushing back 1/10 Architecture, Gemini 1.5, and Dual Personas Alessio asks technical questions about whether dialogue is generated via two independent system prompts or one multi-turn pass. Usama and Raiza explain how Gemini 1.5 long context pairs with DeepMind TTS and dual personas to prevent robotic readouts.18:00–23:09 · The hosts pushing back 1/10 Steven Johnson and Expert Thinking Workflows Swyx draws parallels to Andrew Mason at OpenAI, while Raiza details how she silently observed author Steven Johnson's reasoning workflow to model expert thinking patterns directly into the product UX.23:09–27:54 · The hosts pushing back 1/10 Multimodal Sources, Embeddings, and Audio Tone Alessio queries multimodal ingestion such as handwritten Marie Curie notes and transcribing audio with emotional cues. Raiza and Usama explain the current lossy transcription pipeline and why Gemini requires multimodal embeddings for massive multi-document inputs.27:54–30:23 · The hosts pushing back 0/10 Viral Deep Dive Use Cases on Social Media The hosts and guests swap stories about viral social media use cases, from Karpathy's Wikipedia podcast to Googlers running Q3 performance reviews through Deep Dive. The dynamic is lighthearted and mutual celebration.30:23–37:30 · The hosts pushing back 2/10 Evals, Dogfooding, and Measuring Audio Quality Swyx presses on standard engineering advice to establish early quantitative eval benchmarks for human audio quality. Raiza and Usama explain why they rejected early formal benchmarks in favor of highly opinionated internal dogfooding and team listening sessions.37:30–43:13 · The hosts pushing back 1/10 Pacing, Dialogue Tension, and Speech Dynamics Swyx asks if professional speech experts or comedy writers were consulted to structure dialogue pacing. Usama refutes the appeal to domain authority by sharing an anecdote about linguistics graduates and breaking down narrative tension techniques.43:13–47:03 · The hosts pushing back 0/10 Contextual Humor and the Chicken Paper The discussion covers emergent humor illustrated by the viral Doug Zongker Chicken Paper deep dive. Usama and Raiza emphasize that humor is contextual rather than explicitly prompted.47:04–53:35 · The hosts pushing back 2/10 Hyper-Personalized Media and Product Craft Swyx contrasts compound AI system architecture against monolithic prompts, asking how the team chooses between explicit engineering and LLM delegation. Raiza argues that building AI products is an opinionated craft rather than just an engineering exercise.53:35–1:00:59 · The hosts pushing back 1/10 Roadmap: APIs, Languages, and Resisting Knobs Swyx remarks on how rare it is for PMs to resist exposing knobs and settings to users. Raiza explains her philosophy of maintaining magical one-click simplicity despite loud developer demands for parameter sliders.1:01:00–1:08:27 · The hosts pushing back 1/10 NotebookLM Workspace Future and Real-Time Chat Alessio and Swyx explore workspace competition like Claude Artifacts and OpenAI Realtime API. Raiza and Usama discuss balancing UI craft against fast-following and give final principles for AI engineers pushing model boundaries.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 20.4% · guest 79.6%0:00 · the hosts 20.4% · guest 79.6%3:00 · the hosts 13.6% · guest 86.4%3:00 · the hosts 13.6% · guest 86.4%6:00 · the hosts 15.4% · guest 84.6%6:00 · the hosts 15.4% · guest 84.6%9:00 · the hosts 18.7% · guest 81.3%9:00 · the hosts 18.7% · guest 81.3%12:00 · the hosts 15.1% · guest 84.9%12:00 · the hosts 15.1% · guest 84.9%15:00 · the hosts 11.2% · guest 88.8%15:00 · the hosts 11.2% · guest 88.8%18:00 · the hosts 8.9% · guest 91.1%18:00 · the hosts 8.9% · guest 91.1%21:00 · the hosts 29.2% · guest 70.8%21:00 · the hosts 29.2% · guest 70.8%24:00 · the hosts 16.1% · guest 83.9%24:00 · the hosts 16.1% · guest 83.9%27:00 · the hosts 18.7% · guest 81.3%27:00 · the hosts 18.7% · guest 81.3%30:00 · the hosts 29.1% · guest 70.9%30:00 · the hosts 29.1% · guest 70.9%33:00 · the hosts 18.9% · guest 81.1%33:00 · the hosts 18.9% · guest 81.1%36:00 · the hosts 30.6% · guest 69.4%36:00 · the hosts 30.6% · guest 69.4%39:00 · the hosts 22.9% · guest 77.1%39:00 · the hosts 22.9% · guest 77.1%42:00 · the hosts 17.2% · guest 82.8%42:00 · the hosts 17.2% · guest 82.8%45:00 · the hosts 30.9% · guest 69.1%45:00 · the hosts 30.9% · guest 69.1%48:00 · the hosts 24.7% · guest 75.3%48:00 · the hosts 24.7% · guest 75.3%51:00 · the hosts 50% · guest 50%51:00 · the hosts 50% · guest 50%54:00 · the hosts 34% · guest 66%54:00 · the hosts 34% · guest 66%57:00 · the hosts 28.3% · guest 71.7%57:00 · the hosts 28.3% · guest 71.7%1:00:00 · the hosts 26.6% · guest 73.4%1:00:00 · the hosts 26.6% · guest 73.4%1:03:00 · the hosts 9.2% · guest 90.8%1:03:00 · the hosts 9.2% · guest 90.8%1:06:00 · the hosts 26.4% · guest 73.6%1:06:00 · the hosts 26.4% · guest 73.6%1:09:00 · the hosts 13.1% · guest 86.9%1:09:00 · the hosts 13.1% · guest 86.9%1:12:00 · the hosts 65.8% · guest 34.2%1:12:00 · the hosts 65.8% · guest 34.2%
Sharpest disagreement ▶ 42:35 Dismissing Linguistic Authority

Usama counters Swyx's suggestion to hire domain experts by asserting that linguistics graduates are not necessarily good at generating eloquent spoken dialogue.

Hardest push from the hosts ▶ 31:05 Challenging Missing Eval Benchmarks

Swyx presses the guests on traditional engineering rigour, insisting that audio products need measurable baseline test benchmarks rather than subjective listening.

Biggest teaching moment ▶ 32:02 Strong Taste Over Slow Evals

Raiza educates the hosts on why holding an aggressive internal taste bar through dogfooding is faster and more effective than waiting months for formal rater evaluations.

The host holds their own ▶ 51:45 Framing Compound AI vs Monolithic Models

Swyx demonstrates technical industry breadth by contrasting Databricks compound deterministic pipelines against OpenAI end-to-end prompt architectures.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions and Google Labs Origins 3400 Hosts warmly welcome the guests and prompt them on their background in Google Labs. Raiza and Usama explain the structure of Labs and Area 120, educating the hosts on their org history in a purely friendly opening exchange.
From Talk to Small Corpus to Tailwind 5511 Swyx brings up early context lengths and RAG techniques when discussing Talk to Small Corpus and Project Tailwind. Raiza details the early user studies with adult learners and the natural emergence of document summarization requests.
Discord Growth, Expansion, and Unlaunching Features 4411 Swyx asks about feature flagging and PM methodologies, prompting Raiza to discuss using Mendel experiments and unlaunching unused complex features like passage transformation. The conversation is collaborative and focused on product development practices.
Architecture, Gemini 1.5, and Dual Personas 6611 Alessio asks technical questions about whether dialogue is generated via two independent system prompts or one multi-turn pass. Usama and Raiza explain how Gemini 1.5 long context pairs with DeepMind TTS and dual personas to prevent robotic readouts.
Steven Johnson and Expert Thinking Workflows 5501 Swyx draws parallels to Andrew Mason at OpenAI, while Raiza details how she silently observed author Steven Johnson's reasoning workflow to model expert thinking patterns directly into the product UX.
Multimodal Sources, Embeddings, and Audio Tone 5511 Alessio queries multimodal ingestion such as handwritten Marie Curie notes and transcribing audio with emotional cues. Raiza and Usama explain the current lossy transcription pipeline and why Gemini requires multimodal embeddings for massive multi-document inputs.
Viral Deep Dive Use Cases on Social Media 3300 The hosts and guests swap stories about viral social media use cases, from Karpathy's Wikipedia podcast to Googlers running Q3 performance reviews through Deep Dive. The dynamic is lighthearted and mutual celebration.
Evals, Dogfooding, and Measuring Audio Quality 5622 Swyx presses on standard engineering advice to establish early quantitative eval benchmarks for human audio quality. Raiza and Usama explain why they rejected early formal benchmarks in favor of highly opinionated internal dogfooding and team listening sessions.
Pacing, Dialogue Tension, and Speech Dynamics 5621 Swyx asks if professional speech experts or comedy writers were consulted to structure dialogue pacing. Usama refutes the appeal to domain authority by sharing an anecdote about linguistics graduates and breaking down narrative tension techniques.
Contextual Humor and the Chicken Paper 4510 The discussion covers emergent humor illustrated by the viral Doug Zongker Chicken Paper deep dive. Usama and Raiza emphasize that humor is contextual rather than explicitly prompted.
Hyper-Personalized Media and Product Craft 6512 Swyx contrasts compound AI system architecture against monolithic prompts, asking how the team chooses between explicit engineering and LLM delegation. Raiza argues that building AI products is an opinionated craft rather than just an engineering exercise.
Roadmap: APIs, Languages, and Resisting Knobs 5621 Swyx remarks on how rare it is for PMs to resist exposing knobs and settings to users. Raiza explains her philosophy of maintaining magical one-click simplicity despite loud developer demands for parameter sliders.
NotebookLM Workspace Future and Real-Time Chat 5511 Alessio and Swyx explore workspace competition like Claude Artifacts and OpenAI Realtime API. Raiza and Usama discuss balancing UI craft against fast-following and give final principles for AI engineers pushing model boundaries.

Statements from this episode (26)

Insight
Martin: Users' first prompt on document LLMs is almost always 'summarize'
“Before we did the IO announcement in 23, we'd already done a lot of studies. And one of the first things that I realized was the first thing anybody ever typed was summarize the thing, right? Summarize the document. And it was like half like a test and half ju…”
Raiza Martin Oct 25, 2024 ▶ 6:41
Assertion Supported
Swix: NotebookLM's dedicated Discord server reached 65,000 members
“I just checked in on a Notebook LM Discord. 65,000 people.”
Shawn Wang Oct 25, 2024 ▶ 9:26
Assertion Not checkable as stated
Martin: NotebookLM Discord reports detect server outages faster than internal monitoring
“I think, honestly, like our fastest way that we've been able to find out if like the servers are down or there's just an influx of people being like, it says system unable to answer, anybody else getting this? And I'm like, all right, let's go. And it actually…”
Raiza Martin Oct 25, 2024 ▶ 9:44
Disclosure
Martin: NotebookLM is sunsetting source text transformation due to low usage
“We had this idea that you could highlight the text in your source passage, and then you could transform it. And nobody was really using it. And it was like a very complicated piece of our architecture, and it's very hard to continue supporting it in the contex…”
Raiza Martin Oct 25, 2024 ▶ 11:08
Disclosure
Shafqat: NotebookLM Audio Feature Began as Independent Content Transformation Demo
“We didn't actually start out thinking this would live in notebook, right? Like notebook was sort of, we, Built this demo out independently, tried out like a few different sort of sources that the main idea was like go from some sort of sources and transform it…”
Usama Shafqat Oct 25, 2024 ▶ 13:28
Assertion Supported
Shafqat: NotebookLM Audio Combines DeepMind Voice Tech With Gemini 1.5
“Like we work with the DeepMind audio folks pretty closely. So they're always cooking up new techniques to like get better, more human-like audio. And then Gemini, 1.5 is really, really good at absorbing long context. So we sort of like generally put those thin…”
Usama Shafqat Oct 25, 2024 ▶ 14:28
Insight
Martin: Dual Personas With Editorial Takes Make AI Audio Engaging
“There is a transform that needs to happen. That is inherently editorial. And I think this is where like that two person persona, right? Dialogue model, they have takes on the material that you've presented. That's where it really sort of like brings the conten…”
Raiza Martin Oct 25, 2024 ▶ 15:43
Insight
Martin: NotebookLM's mission is productizing Steven Johnson's research workflow
“And then I had this realization of like, maybe Steven is the product. Maybe the work is to take Steven's expertise and bring it to like everyday people that could really benefit from this.”
Raiza Martin Oct 25, 2024 ▶ 19:09
Assertion Supported
Martin: NotebookLM recognizes and analyzes purely image-based PDFs
“So if you have a PDF that's purely images, it will recognize it”
Raiza Martin Oct 25, 2024 ▶ 22:31
Assertion Supported
Martin: NotebookLM supports up to 50 sources of 500,000 words each
“Notebook.ln can handle up to 50 sources, 500,000 words each, like you're not going to be able to jam all of that into like the context window.”
Raiza Martin Oct 25, 2024 ▶ 25:58
Disclosure
Shafqat: Deep Dive audio pairs text signals with intonation modeling
“The audio model is definitely trying to mimic like certain human intonations and like sort of natural, like, you know, breathing and pauses and like laughter and things like that. But yeah, in generating like the text, we also have to sort of give signals on l…”
Usama Shafqat Oct 25, 2024 ▶ 26:28
Assertion Supported
Martin: NotebookLM transcribes audio inputs without capturing emotion
“So when you upload audio today, we just transcribe it. So it is quite lossy in the sense that like, we don't transcribe like the emotion from that as a source.”
Raiza Martin Oct 25, 2024 ▶ 27:09
Assertion Supported
Martin: Andrej Karpathy launched Spotify podcast using NotebookLM Wikipedia summaries
“I think that's what Karpathy did, right? Like he has now a Spotify channel called histories of mysteries, which is basically like, he just took like interesting stuff from Wikipedia and made audio overviews out of it.”
Raiza Martin Oct 25, 2024 ▶ 29:38
Insight
Martin: Strong internal product taste accelerates AI iteration over external raters
“I think you just have to be really opinionated. I think that sometimes if you are, your intuition is just sharper and you can move a lot faster on the product because it's like, if you hold that bar high, right? Like if you think about like the iterative cycle…”
Raiza Martin Oct 25, 2024 ▶ 32:03
Assertion Not checkable as stated
Shafqat: Steven Johnson decided NotebookLM AI hosts should not have names
“That was a Steven catch. Like not give them names.”
Usama Shafqat Oct 25, 2024 ▶ 37:06
Disclosure
Martin: NotebookLM team includes a dedicated character designer for host personas
“I mean, we have like a really, really good like character designer on our team.”
Raiza Martin Oct 25, 2024 ▶ 37:58
Insight
Shafqat: Engaging audio dialogues require tension and gradual information reveal
“I think one thing we find often is if there's just too much agreement between people, like that's not Fun to listen to. So there needs to be some sort of tension and build up, you know, withholding information, for example, like as you listen to a story unfold…”
Usama Shafqat Oct 25, 2024 ▶ 39:52
Insight
Martin: Chatbot novelty relies on human trickery, not interesting models
“Interacting with a chatbot is sort of novel at first, but it's not interesting. Right. And it's like humans are what makes interacting with chatbots interesting. It's like, ha ha ha, I'm going to try to trick it. It's like, that's interesting. Spell strawberry…”
Raiza Martin Oct 25, 2024 ▶ 43:48
Disclosure
Shafqat: NotebookLM does not directly prompt AI hosts for humor
“Humor is contextual also, like super contextual is what we're realizing. So we're not prompting for humor, but we're prompting for maybe a lot of other things that are bringing out that humor.”
Usama Shafqat Oct 25, 2024 ▶ 46:55
Disclosure
Martin: Google is exploring developer API access for NotebookLM technology
“I think at the same time, right, there are a lot of developers that are interested in using the same technology to build their own thing. We're going to look into that. How soon that's going to be ready, I can't really comment, but these are the things that li…”
Raiza Martin Oct 25, 2024 ▶ 54:34
Prediction Not checkable as stated
Shafqat: NotebookLM will theoretically cover most languages soon
“So I guess high level, like we're definitely working on adding more languages. That's like top priority. We're going to start small, but like theoretically we should be able to cover like most languages pretty soon.”
Usama Shafqat Oct 25, 2024 ▶ 55:36
Insight
Shafqat: Adding user parameters to AI products requires redoing quality evaluations
“The knobs are not as easy to add as simply like, I'm going to add a parameter to this and it's going to make it happen. It's like, you kind of have to redo the quality process for everything.”
Usama Shafqat Oct 25, 2024 ▶ 1:00:23
Assertion Not checkable as stated
Martin: Audio hooks NotebookLM users, but core features drive retention
“We have some early signal that says it's a really good hook. But people stay for the other features.”
Raiza Martin Oct 25, 2024 ▶ 1:01:44
Disclosure
Martin: NotebookLM is prioritizing real-time interactive audio chat
“We're actively, that's one of the things we're actively prioritizing. Actually, one of the interesting things is now we're like, why would anyone want to do that? Right? Like, what are the actual, like kind of going back to sort of having a strong POV about th…”
Raiza Martin Oct 25, 2024 ▶ 1:07:27
Opinion
Swix: OpenAI's Realtime API does not handle interruptions well
“OpenAI has just launched a real-time chat. It's a very hot topic. I would say one of the toughest AI engineering disciplines out there, because even their API doesn't do interruptions that well, to be honest.”
Shawn Wang Oct 25, 2024 ▶ 1:08:09
Insight
Shafqat: Longer thinking time consistently yields better AI results
“More thinking time equals just better results consistently. And that holds true for probably every single time that I've tried to build something.”
Usama Shafqat Oct 25, 2024 ▶ 1:11:49
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.