Feb 18, 2025 · 1h 0m · latent-space

Why is everyone cloning Deep Research?

Mukund Sridhar · 21m spoken Arush Sehgal · 19m spoken Shawn Wang · 10m spoken Alessio Fanelli · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Google Deep Research creators Arush Sehgal and Mukund Sridhar join the Latent Space podcast to unpack the architecture, user experience design, and engineering innovations behind autonomous multi-minute web research agents.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 25.8% of the talking time here. How this is scored →

The hosts as informed peer 5.6 Guest teaching 4.6 Guest disagreement 1.4 The hosts pushing back 3.4
05100:0015:0030:0045:001:00:002:53–6:13 · The hosts as informed peer 4/10 Demonstrating the Query Workflow and Editable Research Plans Arush and Mukund demonstrate the research plan UI using food regulations. Swyx connects the UX feature to an editable chain of thought.6:13–9:14 · The hosts as informed peer 4/10 Under the Hood: Parallel Tool Execution and Iterative Reasoning Mukund breaks down parallel tool execution, breadth-first search, and iterative grounding. Alessio probes how initial ranking and double-clicking work.9:15–11:46 · The hosts as informed peer 5/10 Post-Training Customization and Live Report Analysis Swyx asks if Deep Research can be built over the vanilla Gemini API. Mukund confirms custom post-training was applied to Gemini 1.5 Pro.11:47–16:42 · The hosts as informed peer 5/10 Follow-Up Workflows, Artifact UI, and Ecosystem Integration The conversation covers the side-by-side artifact UI and ecosystem extensions. Swyx playfully challenges who actually wants Spotify integrated into a research tool.16:42–22:40 · The hosts as informed peer 7/10 Technical Architecture: Context Windows versus RAG Systems The discussion covers token context versus RAG systems. Mukund explains vector dot-product limitations on multi-attribute queries, while Swyx and Alessio drill into parsing and vision tradeoffs.22:40–29:11 · The hosts as informed peer 5/10 Evaluation Methodology and User Research Ontologies Arush presents the internal ontology of user research behavior. Swyx pushes back, observing that the current UI visually signals a stopping point rather than an ongoing conversation.29:11–37:17 · The hosts as informed peer 7/10 User Perception of Latency and Execution Trade-offs The hosts and guests explore the paradox where users value slower AI execution. Swyx strongly advocates for Devin-style interactive mid-flight planning rather than locking the chat.37:17–43:03 · The hosts as informed peer 5/10 Product Philosophy, Competitor Clones, and NotebookLM Inspiration Alessio asks about OpenAI copying the exact Deep Research name. Arush explains their foundational design bets around transparent cards and publisher-forward citations.43:04–49:05 · The hosts as informed peer 8/10 Asynchronous Infrastructure and Reasoning Model Trade-offs Swyx explores the mechanics of thinking models versus search verification and maps Google's async infrastructure to durable workflow engines like Temporal and Step Functions.49:06–53:42 · The hosts as informed peer 6/10 Benchmark Limitations and Autonomous Scientific Discovery Swyx presses the team on OpenAI beating Google in marketing via benchmark releases. Mukund and Arush defend prioritizing real user utility over contrived benchmark optimization.2:53–6:13 · Guest teaching 3/10 Demonstrating the Query Workflow and Editable Research Plans Arush and Mukund demonstrate the research plan UI using food regulations. Swyx connects the UX feature to an editable chain of thought.6:13–9:14 · Guest teaching 5/10 Under the Hood: Parallel Tool Execution and Iterative Reasoning Mukund breaks down parallel tool execution, breadth-first search, and iterative grounding. Alessio probes how initial ranking and double-clicking work.9:15–11:46 · Guest teaching 4/10 Post-Training Customization and Live Report Analysis Swyx asks if Deep Research can be built over the vanilla Gemini API. Mukund confirms custom post-training was applied to Gemini 1.5 Pro.11:47–16:42 · Guest teaching 3/10 Follow-Up Workflows, Artifact UI, and Ecosystem Integration The conversation covers the side-by-side artifact UI and ecosystem extensions. Swyx playfully challenges who actually wants Spotify integrated into a research tool.16:42–22:40 · Guest teaching 6/10 Technical Architecture: Context Windows versus RAG Systems The discussion covers token context versus RAG systems. Mukund explains vector dot-product limitations on multi-attribute queries, while Swyx and Alessio drill into parsing and vision tradeoffs.22:40–29:11 · Guest teaching 6/10 Evaluation Methodology and User Research Ontologies Arush presents the internal ontology of user research behavior. Swyx pushes back, observing that the current UI visually signals a stopping point rather than an ongoing conversation.29:11–37:17 · Guest teaching 4/10 User Perception of Latency and Execution Trade-offs The hosts and guests explore the paradox where users value slower AI execution. Swyx strongly advocates for Devin-style interactive mid-flight planning rather than locking the chat.37:17–43:03 · Guest teaching 4/10 Product Philosophy, Competitor Clones, and NotebookLM Inspiration Alessio asks about OpenAI copying the exact Deep Research name. Arush explains their foundational design bets around transparent cards and publisher-forward citations.43:04–49:05 · Guest teaching 5/10 Asynchronous Infrastructure and Reasoning Model Trade-offs Swyx explores the mechanics of thinking models versus search verification and maps Google's async infrastructure to durable workflow engines like Temporal and Step Functions.49:06–53:42 · Guest teaching 6/10 Benchmark Limitations and Autonomous Scientific Discovery Swyx presses the team on OpenAI beating Google in marketing via benchmark releases. Mukund and Arush defend prioritizing real user utility over contrived benchmark optimization.2:53–6:13 · Guest disagreement 1/10 Demonstrating the Query Workflow and Editable Research Plans Arush and Mukund demonstrate the research plan UI using food regulations. Swyx connects the UX feature to an editable chain of thought.6:13–9:14 · Guest disagreement 1/10 Under the Hood: Parallel Tool Execution and Iterative Reasoning Mukund breaks down parallel tool execution, breadth-first search, and iterative grounding. Alessio probes how initial ranking and double-clicking work.9:15–11:46 · Guest disagreement 1/10 Post-Training Customization and Live Report Analysis Swyx asks if Deep Research can be built over the vanilla Gemini API. Mukund confirms custom post-training was applied to Gemini 1.5 Pro.11:47–16:42 · Guest disagreement 2/10 Follow-Up Workflows, Artifact UI, and Ecosystem Integration The conversation covers the side-by-side artifact UI and ecosystem extensions. Swyx playfully challenges who actually wants Spotify integrated into a research tool.16:42–22:40 · Guest disagreement 1/10 Technical Architecture: Context Windows versus RAG Systems The discussion covers token context versus RAG systems. Mukund explains vector dot-product limitations on multi-attribute queries, while Swyx and Alessio drill into parsing and vision tradeoffs.22:40–29:11 · Guest disagreement 1/10 Evaluation Methodology and User Research Ontologies Arush presents the internal ontology of user research behavior. Swyx pushes back, observing that the current UI visually signals a stopping point rather than an ongoing conversation.29:11–37:17 · Guest disagreement 2/10 User Perception of Latency and Execution Trade-offs The hosts and guests explore the paradox where users value slower AI execution. Swyx strongly advocates for Devin-style interactive mid-flight planning rather than locking the chat.37:17–43:03 · Guest disagreement 1/10 Product Philosophy, Competitor Clones, and NotebookLM Inspiration Alessio asks about OpenAI copying the exact Deep Research name. Arush explains their foundational design bets around transparent cards and publisher-forward citations.43:04–49:05 · Guest disagreement 2/10 Asynchronous Infrastructure and Reasoning Model Trade-offs Swyx explores the mechanics of thinking models versus search verification and maps Google's async infrastructure to durable workflow engines like Temporal and Step Functions.49:06–53:42 · Guest disagreement 2/10 Benchmark Limitations and Autonomous Scientific Discovery Swyx presses the team on OpenAI beating Google in marketing via benchmark releases. Mukund and Arush defend prioritizing real user utility over contrived benchmark optimization.2:53–6:13 · The hosts pushing back 1/10 Demonstrating the Query Workflow and Editable Research Plans Arush and Mukund demonstrate the research plan UI using food regulations. Swyx connects the UX feature to an editable chain of thought.6:13–9:14 · The hosts pushing back 2/10 Under the Hood: Parallel Tool Execution and Iterative Reasoning Mukund breaks down parallel tool execution, breadth-first search, and iterative grounding. Alessio probes how initial ranking and double-clicking work.9:15–11:46 · The hosts pushing back 2/10 Post-Training Customization and Live Report Analysis Swyx asks if Deep Research can be built over the vanilla Gemini API. Mukund confirms custom post-training was applied to Gemini 1.5 Pro.11:47–16:42 · The hosts pushing back 3/10 Follow-Up Workflows, Artifact UI, and Ecosystem Integration The conversation covers the side-by-side artifact UI and ecosystem extensions. Swyx playfully challenges who actually wants Spotify integrated into a research tool.16:42–22:40 · The hosts pushing back 4/10 Technical Architecture: Context Windows versus RAG Systems The discussion covers token context versus RAG systems. Mukund explains vector dot-product limitations on multi-attribute queries, while Swyx and Alessio drill into parsing and vision tradeoffs.22:40–29:11 · The hosts pushing back 4/10 Evaluation Methodology and User Research Ontologies Arush presents the internal ontology of user research behavior. Swyx pushes back, observing that the current UI visually signals a stopping point rather than an ongoing conversation.29:11–37:17 · The hosts pushing back 6/10 User Perception of Latency and Execution Trade-offs The hosts and guests explore the paradox where users value slower AI execution. Swyx strongly advocates for Devin-style interactive mid-flight planning rather than locking the chat.37:17–43:03 · The hosts pushing back 3/10 Product Philosophy, Competitor Clones, and NotebookLM Inspiration Alessio asks about OpenAI copying the exact Deep Research name. Arush explains their foundational design bets around transparent cards and publisher-forward citations.43:04–49:05 · The hosts pushing back 4/10 Asynchronous Infrastructure and Reasoning Model Trade-offs Swyx explores the mechanics of thinking models versus search verification and maps Google's async infrastructure to durable workflow engines like Temporal and Step Functions.49:06–53:42 · The hosts pushing back 5/10 Benchmark Limitations and Autonomous Scientific Discovery Swyx presses the team on OpenAI beating Google in marketing via benchmark releases. Mukund and Arush defend prioritizing real user utility over contrived benchmark optimization.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 47% · guest 53%0:00 · the hosts 47% · guest 53%3:00 · the hosts 6.9% · guest 93.1%3:00 · the hosts 6.9% · guest 93.1%6:00 · the hosts 12.3% · guest 87.7%6:00 · the hosts 12.3% · guest 87.7%9:00 · the hosts 28.4% · guest 71.6%9:00 · the hosts 28.4% · guest 71.6%12:00 · the hosts 14.1% · guest 85.9%12:00 · the hosts 14.1% · guest 85.9%15:00 · the hosts 37.8% · guest 62.2%15:00 · the hosts 37.8% · guest 62.2%18:00 · the hosts 34.8% · guest 65.2%18:00 · the hosts 34.8% · guest 65.2%21:00 · the hosts 8.1% · guest 91.9%21:00 · the hosts 8.1% · guest 91.9%24:00 · the hosts 4.7% · guest 95.3%24:00 · the hosts 4.7% · guest 95.3%27:00 · the hosts 20.1% · guest 79.9%27:00 · the hosts 20.1% · guest 79.9%30:00 · the hosts 23.9% · guest 76.1%30:00 · the hosts 23.9% · guest 76.1%33:00 · the hosts 16% · guest 84%33:00 · the hosts 16% · guest 84%36:00 · the hosts 49.1% · guest 50.9%36:00 · the hosts 49.1% · guest 50.9%39:00 · the hosts 18.8% · guest 81.2%39:00 · the hosts 18.8% · guest 81.2%42:00 · the hosts 29.8% · guest 70.2%42:00 · the hosts 29.8% · guest 70.2%45:00 · the hosts 30.5% · guest 69.5%45:00 · the hosts 30.5% · guest 69.5%48:00 · the hosts 34.1% · guest 65.9%48:00 · the hosts 34.1% · guest 65.9%51:00 · the hosts 24.3% · guest 75.7%51:00 · the hosts 24.3% · guest 75.7%54:00 · the hosts 6.8% · guest 93.2%54:00 · the hosts 6.8% · guest 93.2%57:00 · the hosts 71.8% · guest 28.2%57:00 · the hosts 71.8% · guest 28.2%1:00:00 · the hosts 0% · guest 0%1:00:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 49:21 Dismissing contrived industry benchmarks

Arush forcefully rejects the premise of standard benchmarks by mocking esoteric questions like asking who a president's nephew was on the day Kobe entered the league.

Hardest push from the hosts ▶ 36:37 Demanding unlocked mid-execution chat

Swyx refuses the current locked UX paradigm of Deep Research, insisting that systems should allow live conversational steering while executing plans like Devin.

Biggest teaching moment ▶ 20:19 Why RAG dot-products break down on complex queries

Mukund educates the hosts on the mathematical limitations of cosine similarity in vector search when queries have multiple attributes compared to full long-context processing.

The host holds their own ▶ 48:08 Architectural breakdown of async durable execution

Swyx demonstrates deep domain expertise in distributed orchestration, correctly categorizing Google's unannounced async system alongside Temporal, Apache Airflow, and AWS Step Functions.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Demonstrating the Query Workflow and Editable Research Plans 4311 Arush and Mukund demonstrate the research plan UI using food regulations. Swyx connects the UX feature to an editable chain of thought.
Under the Hood: Parallel Tool Execution and Iterative Reasoning 4512 Mukund breaks down parallel tool execution, breadth-first search, and iterative grounding. Alessio probes how initial ranking and double-clicking work.
Post-Training Customization and Live Report Analysis 5412 Swyx asks if Deep Research can be built over the vanilla Gemini API. Mukund confirms custom post-training was applied to Gemini 1.5 Pro.
Follow-Up Workflows, Artifact UI, and Ecosystem Integration 5323 The conversation covers the side-by-side artifact UI and ecosystem extensions. Swyx playfully challenges who actually wants Spotify integrated into a research tool.
Technical Architecture: Context Windows versus RAG Systems 7614 The discussion covers token context versus RAG systems. Mukund explains vector dot-product limitations on multi-attribute queries, while Swyx and Alessio drill into parsing and vision tradeoffs.
Evaluation Methodology and User Research Ontologies 5614 Arush presents the internal ontology of user research behavior. Swyx pushes back, observing that the current UI visually signals a stopping point rather than an ongoing conversation.
User Perception of Latency and Execution Trade-offs 7426 The hosts and guests explore the paradox where users value slower AI execution. Swyx strongly advocates for Devin-style interactive mid-flight planning rather than locking the chat.
Product Philosophy, Competitor Clones, and NotebookLM Inspiration 5413 Alessio asks about OpenAI copying the exact Deep Research name. Arush explains their foundational design bets around transparent cards and publisher-forward citations.
Asynchronous Infrastructure and Reasoning Model Trade-offs 8524 Swyx explores the mechanics of thinking models versus search verification and maps Google's async infrastructure to durable workflow engines like Temporal and Step Functions.
Benchmark Limitations and Autonomous Scientific Discovery 6625 Swyx presses the team on OpenAI beating Google in marketing via benchmark releases. Mukund and Arush defend prioritizing real user utility over contrived benchmark optimization.

Statements from this episode (27)

Insight
Sridhar: Multi-minute AI research creates unique UX alignment and web navigation challenges
“This is one of the first times, you know, something takes about five, six minutes trying to perform your research, so there's a few challenges that brings, like, you want to make sure you're spending that time in the computer doing what the user wants, so ther…”
Mukund Sridhar Feb 18, 2025 ▶ 1:43
Insight
Sridhar: Deep Research targets multi-tab exploratory queries rather than direct searches
“There are things that, you know exactly what you're looking for and their search is still probably, you know, a very, you know, probably one of the best places to go. I think where deep research really shines is that, like, Multiple facets to your question, an…”
Mukund Sridhar Feb 18, 2025 ▶ 2:27
Insight
Sehgal: AI deep research provides most lift on niche, non-Wikipedia topics
“We love to test, like, super niche random things, like, things where there's, like, No Wikipedia page already about this topic or something like that, right? Because that's where you'll see the most lift from a feature like this.”
Arush Sehgal Feb 18, 2025 ▶ 2:59
Disclosure
Sehgal: Early Gemini Deep Research testers never edited initial research plans
“Actually like in early rounds of testing, we saw no one was editing, and so we were just like, if we just put a button here, Maybe people will, like, engage more.”
Arush Sehgal Feb 18, 2025 ▶ 5:30
Assertion Supported
Sridhar: Gemini Deep Research operates primarily via search and page-deepening tools
“What's happening behind the scenes actually is we kind of give this research plan that is a contract and that you know, has been accepted. But then if you look at the plan, there are things that are obviously parallelizable. So the model figures out which of t…”
Mukund Sridhar Feb 18, 2025 ▶ 6:41
Insight
Sridhar: Sequential grounding on prior search turns is key for deep research
“This notion of being able to read outputs from the previous turn ground on that to decide what to do next, I think was key. Otherwise, you have, like, incomplete information, and your report becomes a little bit of a, like, a high-level bullet point.”
Mukund Sridhar Feb 18, 2025 ▶ 7:29
Assertion Supported
Sridhar: Deep Research uses self-critique to resolve source inconsistencies in reports
“So this happens iteratively until the model thinks it's finished all its steps, and then we kind of enter this analysis mode, and here there can be inconsistencies across sources. You kind of come up with an outline for the report, start generating a draft. Th…”
Mukund Sridhar Feb 18, 2025 ▶ 7:52
Disclosure
Sridhar: Deep Research Uses Base Gemini With Custom Post-Training
“Yeah, I don't think we have special access, per se. It's pretty much the same model. We, of course, have our own post-training work that we do, and Y'all can also, like, you know, you can fine tune from the base model and so on.”
Mukund Sridhar Feb 18, 2025 ▶ 9:32
Assertion Supported
Sehgal: Gemini Deep Research retains all browsed websites in context
“We actually keep everything in context, like all the sites that it's read remain in context. So if there's a piece of missing information, it can just fetch that.”
Arush Sehgal Feb 18, 2025 ▶ 12:35
Disclosure
Sridhar: Model decides whether follow-ups trigger new Deep Research runs
“One of the challenges is currently we kind of let the model decide based on your query, like amongst the three categories. So some, there is a boundary there. Like some of these things, depending on how deep you want to go, you might just want a quick answer v…”
Mukund Sridhar Feb 18, 2025 ▶ 15:53
Disclosure
Sridhar: Gemini Deep Research falls back to RAG beyond context limits
“We also have we have retrieval mechanisms, if required. So we natively try to use the context as much as it's available beyond which you know, we have, like, a rag setup to figure out”
Mukund Sridhar Feb 18, 2025 ▶ 19:03
Insight
Sridhar: Vector dot-product RAG breaks down on multi-attribute queries
“The tricky thing for RAG, it really works well because a lot of these things are doing like cosine distance, like a dot product kind of a thing, and that kind of gets challenging when your query side has multiple different attributes. The dot product doesn't r…”
Mukund Sridhar Feb 18, 2025 ▶ 20:19
Insight
Sehgal: Keep recent research tasks in context, relegating older ones to RAG
“Just to add to that, I think like, just like a simple rule of thumb that we use is like, if it's the most recent set of research tasks where the user is likely to ask lots of follow-up questions, that should be in context. But like, as stuff gets 10 tasks ago,…”
Arush Sehgal Feb 18, 2025 ▶ 21:11
Insight
Sehgal: Deep research evals should categorize research behavior, not domain verticals
“And really what we tried to do is like, stay away from like verticals, like travel or shopping and things like that, but really try and go into like, what is the underlying research behavior Type that a person is doing.”
Arush Sehgal Feb 18, 2025 ▶ 24:53
Assertion Not checkable as stated
Sehgal: Gemini Deep Research lacks turn limits, but users rarely go deep
“We don't have any hard limits on the, how many turns you can do. One thing I will say is most users don't go very deep right now.”
Arush Sehgal Feb 18, 2025 ▶ 27:13
Disclosure
Sridhar: Google launched Deep Research horizontally rather than for a single vertical
“Our primary goal was Not to specialize in, in, in a particular vertical or target one type of user. We just want to put this in the hands of like we had like this busy parent persona and like various different user profiles and see like what people try to use …”
Mukund Sridhar Feb 18, 2025 ▶ 28:53
Insight
Swix: AI research agents face perverse incentives favoring inefficiency and latency
“I think there's a perverse incentives for research agents to take longer and it be perceived to be better to people are like, oh, you're searching like so many websites for me, you know, but like 30 of them are irrelevant. You know, like, I feel like right now…”
Shawn Wang Feb 18, 2025 ▶ 31:16
Disclosure
Sehgal: Google shipped a 5-minute Deep Research mode fearing user drop-off
“I remember we actually built two versions of deep research. We had like a hardcore mode that takes like 15 minutes. And then what we actually shipped is a thing that takes five minutes. And I even went to Eng and I was like, there has to be a hard stop, by the…”
Arush Sehgal Feb 18, 2025 ▶ 32:09
Insight
Sehgal: Users always max out AI power toggles if given the option
“If you like give a max power button, users are always just going to hit that button, right? So then the question comes like, why don't you just decide from the product POV where's the right balance?”
Arush Sehgal Feb 18, 2025 ▶ 34:58
Opinion
Swix: AI agent interfaces should never lock chat during execution
“I think you should never lock the chat. You should always be able to chat with the plan and update the plan, and the plan scheduler, whatever orchestration system you have under the hood, should just pick off the next job on the list.”
Shawn Wang Feb 18, 2025 ▶ 37:07
Opinion
Swix: Devin's hourly billing model incentivizes slow execution
“And it's perverse in senses where they charge by hour. So they make more money, the slower they are.”
Shawn Wang Feb 18, 2025 ▶ 38:07
Insight
Sridhar: Thinking models inherently enable self-critiquing of partial steps
“The new generation models, especially with these thinking models, they unlock a few things. So I think one is obviously the, Better capability in, like, analytical thinking, like in math, coding, and these type of things, but also this notion of, you know, as …”
Mukund Sridhar Feb 18, 2025 ▶ 42:31
Insight
Sridhar: Multi-minute agent jobs require persistent state to survive inevitable failures
“If you build, like, five, six minute jobs, they're bound to be, like, failures and you don't want to, like, retry, lose your progress and so on, so this notion of, like, keeping state knowing what to retry and kind of keep the journey going.”
Mukund Sridhar Feb 18, 2025 ▶ 47:30
Opinion
Sridhar: High HLE benchmark scores do not translate to deep research products
“The benchmarks, at least the ones that we are seeing, they don't directly translate to the product. There's definitely some technical challenges that you can benchmark against, but they don't really, like if I do grade on HLE, that doesn't really mean I'm a g…”
Mukund Sridhar Feb 18, 2025 ▶ 50:03
Insight
Sridhar: Autonomous AI discovery requires verifier sandboxes and second-order reasoning
“My personal opinion is the model doesn't, has to do the second order thinking and so on that we're seeing now with these new models, but also be able to play and test that out in an environment where you can, you know, verify and give it feedback so that it ca…”
Mukund Sridhar Feb 18, 2025 ▶ 53:08
Insight
Sridhar: Horizontal plug-and-play AI agent platforms are premature
“I feel like it's still early days for us, like to try to platformatize or like try to build these, oh, there are these five horizontal pieces. And you can plug and play and build your own agent. My personal opinion is we are not there yet. In order to build a …”
Mukund Sridhar Feb 18, 2025 ▶ 55:58
Opinion
Swix: Deep research agents are the first agent category with true PMF
“What are the hard problems in this brand of agent that is like probably the first real product market fit agent. I will say more so than the computer use ones. This is the one where like, yeah, people are like, yeah, easily pays for 200 dollars worth a month w…”
Shawn Wang Feb 18, 2025 ▶ 58:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.