Arush Sehgal (Lead PM for Gemini Deep Research) explains how Google developed their evaluation benchmark ontology across research archetypes instead of specific industries.
Insight
Sehgal: Keep recent research tasks in context, relegating older ones to RAG
“Just to add to that, I think like, just like a simple rule of thumb that we use is like, if it's the most recent set of research tasks where the user is likely to ask lots of follow-up questions, that should be in context. But like, as stuff gets 10 tasks ago,…”
Insight
Sehgal: AI deep research provides most lift on niche, non-Wikipedia topics
“We love to test, like, super niche random things, like, things where there's, like, No Wikipedia page already about this topic or something like that, right? Because that's where you'll see the most lift from a feature like this.”
Disclosure
Sehgal: Google shipped a 5-minute Deep Research mode fearing user drop-off
“I remember we actually built two versions of deep research. We had like a hardcore mode that takes like 15 minutes. And then what we actually shipped is a thing that takes five minutes. And I even went to Eng and I was like, there has to be a hard stop, by the…”
Insight
Sehgal: Users always max out AI power toggles if given the option
“If you like give a max power button, users are always just going to hit that button, right? So then the question comes like, why don't you just decide from the product POV where's the right balance?”
Disclosure
Sehgal: Early Gemini Deep Research testers never edited initial research plans
“Actually like in early rounds of testing, we saw no one was editing, and so we were just like, if we just put a button here, Maybe people will, like, engage more.”
Assertion Supported
Sehgal: Gemini Deep Research retains all browsed websites in context
“We actually keep everything in context, like all the sites that it's read remain in context. So if there's a piece of missing information, it can just fetch that.”