why aren't all 14 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Disclosure
Sridhar: Gemini Deep Research falls back to RAG beyond context limits
“We also have we have retrieval mechanisms, if required. So we natively try to use the context as much as it's available beyond which you know, we have, like, a rag setup to figure out”
Opinion
Sridhar: High HLE benchmark scores do not translate to deep research products
“The benchmarks, at least the ones that we are seeing, they don't directly translate to the product. There's definitely some technical challenges that you can benchmark against, but they don't really, like if I do grade on
HLE, that doesn't really mean I'm a g…”
Insight
Sridhar: Sequential grounding on prior search turns is key for deep research
“This notion of being able to read outputs from the previous turn ground on that to decide what to do next, I think was key. Otherwise, you have, like, incomplete information, and your report becomes a little bit of a, like, a high-level bullet point.”
Insight
Sehgal: Deep research evals should categorize research behavior, not domain verticals
“And really what we tried to do is like, stay away from like verticals, like travel or shopping and things like that, but really try and go into like, what is the underlying research behavior
Type that a person is doing.”
Disclosure
Sridhar: Google launched Deep Research horizontally rather than for a single vertical
“Our primary goal was
Not to specialize in, in, in a particular vertical or target one type of user. We just want to put this in the hands of like we had like this busy parent persona and like various different user profiles and see like what people try to use …”
Insight
Sridhar: Autonomous AI discovery requires verifier sandboxes and second-order reasoning
“My personal opinion is the model doesn't, has to do the second order thinking and so on that we're seeing now with these new models, but also be able to play and test that out in an environment where you can, you know, verify and give it feedback so that it ca…”
Disclosure
Sridhar: Deep Research Uses Base Gemini With Custom Post-Training
“Yeah, I don't think we have special access, per se. It's pretty much the same model. We, of course, have our own post-training work that we do, and Y'all can also, like, you know, you can fine tune from the base model and so on.”
Disclosure
Sehgal: Google shipped a 5-minute Deep Research mode fearing user drop-off
“I remember we actually built two versions of deep research. We had like a hardcore mode that takes like 15 minutes. And then what we actually shipped is a thing that takes five minutes. And I even went to Eng and I was like, there has to be a hard stop, by the…”
Disclosure
Sridhar: Model decides whether follow-ups trigger new Deep Research runs
“One of the challenges is currently we kind of let the model decide based on your query, like amongst the three categories. So some, there is a boundary there. Like some of these things, depending on how deep you want to go, you might just want a quick answer v…”
Assertion Supported
Sridhar: Deep Research uses self-critique to resolve source inconsistencies in reports
“So this happens iteratively until the model thinks it's finished all its steps, and then we kind of enter this analysis mode, and here there can be inconsistencies across sources. You kind of come up with an outline for the report, start generating a draft. Th…”
Assertion Not checkable as stated
Sehgal: Gemini Deep Research lacks turn limits, but users rarely go deep
“We don't have any hard limits on the, how many turns you can do. One thing I will say is most users don't go very deep right now.”
Assertion Supported
Sehgal: Gemini Deep Research retains all browsed websites in context
“We actually keep everything in context, like all the sites that it's read remain in context. So if there's a piece of missing information, it can just fetch that.”
Disclosure
Sehgal: Early Gemini Deep Research testers never edited initial research plans
“Actually like in early rounds of testing, we saw no one was editing, and so we were just like, if we just put a button here, Maybe people will, like, engage more.”
Assertion Supported
Sridhar: Gemini Deep Research operates primarily via search and page-deepening tools
“What's happening behind the scenes actually is we kind of give this research plan that is a contract and that you know, has been accepted. But then if you look at the plan, there are things that are obviously parallelizable. So the model figures out which of t…”