why aren't all 40 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Hong: DeepMind's Formal Math Slowdown Post-AlphaProof Was Non-Technical
“After AlphaProof, kind of like, we didn't see a lot of the formal math you know, results or kind of progress from Google DeepMind, and that's actually because of reasons that are not necessarily technical.”
Opinion
Yi Tay: Today's Models Likely Couldn't Invent the Transformer from Pre-2015 Data
“Even today's models, they might not even be able to invent the transformer. Like, if you freeze the time at a certain time, and even you bring the time, I mean, the model is a transformer, so I just say there's no, assuming there's no leakage.”
Prediction Not checkable as stated
Tay: Most specialized tools will be subsumed directly into model parameters
“Then the most I can see in the future is there'll be a model then that, that is, there's something that really cannot be subsumed by a model. Then you just use a tool or something, right? But my prediction is that I think most things can be subsumed by the mod…”
Disclosure
DeepMind Abandoned AlphaProof to Run Gemini End-to-End for IMO Math
“We wanted to try to, like, use, actually use Gemini as an end-to-end model. Basically, no, no second system with alpha proof. No second system. In, text out.”
Prediction Not checkable as stated
Kilpatrick: Generative UI will be the killer use case for diffusion LLMs
“But I do think that's going to be the killer use case will be like this generative UI experience that doesn't exist today because the models just take too long to generate tokens.”
Assertion Not checkable as stated
Kilpatrick: Gemini's SOTA video performance resulted from reasoning, not video engineering
“With reasoning is a great example of this where like multimodal with video understanding ended up like having this huge, like it's having this beautiful moment.
The model is like soda out of the box because of all the reasoning capabilities that were baked in…”
Disclosure
Kilpatrick: Google's strategy is building Gemini as one single unified model
“Like we're here to make one model and like that model is Gemini.
And like, I think you do need to just trust this point, like to make the capabilities work in some cases, like you do need to have these forks that like go off and make that capability and harde…”
Assertion Supported
Agarwal: Filtered 9B Synthetic Data Outperforms 27B Self-Generated Data
“One thing we found consistently, so here what we had two models, nine Gemma, nine B and Gemma, 27 B, and we found consistently that actually generating data from nine B in a compute match setting is always better, even better for distilling or actually improvi…”
Opinion
Swyx: DeepMind has roughly a four-year advantage over OpenAI in world modeling
“So like they have maybe four years advantage on world modeling that OpenAI does not have. Cause OpenAI basically only started Diffusion Transformers last year when they hired Build Peebles. So DeepMind has a bit of advantage here.”
Assertion Supported
DeepMind, Microsoft, and Meta are building or using physical science labs
“You see people like Google DeepMind, Microsoft, other places like Meta, either building their own lab or running experiments at someone else's lab to get that data back.”
Prediction Not checkable as stated
Sanseviero: Deep architecture research will not be automated within two years
“If you want to do, like, deeper research in the architecture, my hunch is that most likely this will not be, like, automatable, at least in the next one or two years.”
Assertion Not checkable as stated
Gemini's IMO Model Checkpoint Required Only One Week of Training
“The training process of this IMO model itself was, like, maybe a week or so.”
Insight
Yi Tay: On-policy RL is more generalizable than imitation fine-tuning
“So I think on policyness is basically this idea of like model training on its own outputs and letting the model like generate its own trajectories and then letting some reward verify it and then the model train its own outputs. I think this is more generalizab…”
Insight
Yi Tay: AI progress is driven by compounding small incremental changes
“I think that it's true that sometimes a lot of progress on the whole is just a series of small incremental changes that, yeah, that push. I think that's accurate. That's true. There's also, it also feels that there's also a lot of like small, like seemingly mi…”
Assertion Contradicted
Mallick: Gemini recognizes distinct voices as an unsupported emergent behavior
“This is not officially supported yet. The model just does it.”
Insight
Agarwal: Distillation Drives Year-Over-Year AI Capability Cost Reductions
“The capability which we have right now, maybe next year will be much cheaper to have that same thing. And that likely is the result of distillation, right? Like that's, it's not just because we are doing or figured out something magical. It is because distilla…”
Assertion Not checkable as stated
Malde walked away from $2B DeepMind acquisition to found Trajectory.ai
“Obviously the acquisition was for two billion dollars and went over to DeepBind, and then I decided to give up all the acquisition money to start trajectory.”
Opinion
Gemma 4 matches frontier AI state of the art from 18 months ago
“I mean, if you look at Gemma, you compare to how we were one year ago, I would say Gemma four is matching state of the art from one and a half years ago for most. Things.”
Assertion Not checkable as stated
Sanseviero: Text diffusion model quality is still worse than autoregressive models
“I think especially like the model quality is still a bit worse from what you would get from the normal autoregressive model.”
Prediction Not checkable as stated
Sanseviero: Next generation of AI fine-tuners will not write code
“I do think the next generation of fine tuners will not be, I mean, will be people that are not coding at all, right? Like one year ago, we had to write like our own Colab with Transformers or Oncelot or whichever library of your choice. I do think as we like k…”
Disclosure
Sanseviero: Many Gemma 4 Launch Partners Skipped Fine-Tuning Due to Base Performance
“For Gemma four, we had 50 To 60 partners. And some of them were like, oh yeah, we're going to try and fine tune the 27 B model for this vision task. And they were like, oh, actually the model works too well out of the box. We don't need to fine tune it. Yeah. …”
Opinion
Yi Tay: Reinforcement learning is the primary AI modeling toolset today
“So I think RL is basically the main modeling tool set that we play around with these days.”
Disclosure
DeepMind Shipped Full Gemini IMO Config Only to Select Mathematicians
“So the inference time config was like the one serve to most people is different, but And the full IMO, like, inference config was presented, like, shipped to some mathematicians just because of the inference cost, right? But that was good enough to be a genera…”
Insight
Yi Tay: Independent research taste is a stronger hiring signal than execution
“If somebody comes up with something and then does something that you feel that is very tasteful and it aligns with what Like researchers in the labs, like one, and they come up with that independently, you know that the function that is good, right? Like if yo…”
Disclosure
DeepMind Singapore is keeping its Gemini RL team small to maximize compute per capita
“We're hiring, like, my team will work on like RL and reasoning for Gemini and Gemini deep thing. I think we care more about like talent density now. So we're not like also like growing that big, this small team first, just because compute per capita is probabl…”
Assertion Supported
DeepMind and OpenAI eliminated formal Lean translation for 2025 IMO solutions
“What surprised me is this time they don't use formal language, but instead they just use LM. And so last year when they tried to do the IMO, they need like a like a human to kind of translate the natural language. Problems to Lean, and then they use Lean to ki…”
Assertion Supported
Noam Shazeer and Jack Rae co-lead Google DeepMind's reasoning effort
“Jack Ray. Yeah. He's been a long time deep mind research scientist, was previously a pre-training person. We actually overlapped at open AI together a little bit, and then is now back at deep mind with no co-leading the reasoning effort.”
Disclosure
Sanseviero: DeepMind is building agentic tools for research ablations and evaluations
“So for example, within the team, we are building skills to do experiments and ablations and evaluations and how the research team can use all of these agentic tools as part of their research process is also quite interesting.”
Assertion Supported
Sanseviero: Smaller Gemma 4 models process audio and 30-60 second videos
“Multimodal wise, the smaller models can understand audio Images and short videos, so, 30 to 62nd videos and audios.”
Assertion Supported
Sanseviero: Gemma 4 is Google's most capable open model yet
“Gemma four is just out. It's the most capable open model we've released so far. We already tried to compact as much intelligence per parameter as we could, bring all of these multimodal capabilities.”
Disclosure
DeepMind's Gemma team runs with two to three PMs and one marketer
“The Gemma team is actually relatively small. We have, like, two or three PMs. We have one marketing person, and then there are, like, engineers and researchers working on shipping this.”
Assertion Supported
Sanseviero: Gemma 4 cannot yet process video and audio simultaneously
“The other thing we do not support yet is video with audio, so we can understand, like, video input or audio input separately, but if you want to pass, like, in the same, from both the visual part and the audio part, we still need to do some improvements around…”
Disclosure
Tay: DeepMind did not optimize Gemini specifically for Pokémon
“There's actually nothing specifically done for Pokemon.”
Disclosure
Yi Tay: Four captains across three locations trained DeepMind's IMO model
“So I think there were four captains for the IMO, two from London. Jonathan was from Mountain View. I was from Singapore. So I think four of us basically trained this model together.”
Disclosure
Yi Tay: I had almost no RL background before returning to DeepMind
“I spent a lot of my past life, I call it the past art, working on like architectures and pre-training, but I think now I more, I have like transitioned more into RL. I'm not like old school RL, but the games RL and the old school RL, and to be honest, I had al…”
Prediction Not checkable as stated
Mallick: Google's URL Context tool will enable developer research agents
“And I think that this will unlock new use cases. Like if people want to build their own version of a research agent, which is something developers ask us for a lot.”
Disclosure
Mallick: Google launched async function calling for cascaded voice models
“One thing that we launched on the cascaded architecture that we hope to eventually bring to the native audio as well is asynchronous function calling.”
Assertion Supported
Mallick: Gemini Live API allows custom VAD tuning and third-party integration
“Now developers can actually tune the sensitivity on our voice activity detection model as well as, you know, how much of the prefix, like how much of a time duration at the beginning, at the start or stop of saying things. And we also have a mode where now you…”
Disclosure
Kilpatrick: Gemini 2.5 Pro will allow disabling thinking in early June
“Thinking budgets coming to 2.5 pro. So, and you can also, you'll be able to disable thinking as well. So if you just want 2.5 pro is like a raw non reasoning model, we'll have that hopefully in early June”
Disclosure
DeepMind Singapore Explicitly Adds AGI to Job Postings
“I think that, like, one reason why we work on these models is that we want to get to AGI, and, like, this was a right thing that we added AGI to the job posting, yeah. There is no, like, formal name of the team yet, but it's basically the Gemini theme Singapor…”