why aren't all 27 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Prediction Not checkable as stated
Tay: Most specialized tools will be subsumed directly into model parameters
“Then the most I can see in the future is there'll be a model then that, that is, there's something that really cannot be subsumed by a model. Then you just use a tool or something, right? But my prediction is that I think most things can be subsumed by the mod…”
Disclosure
DeepMind Abandoned AlphaProof to Run Gemini End-to-End for IMO Math
“We wanted to try to, like, use, actually use Gemini as an end-to-end model. Basically, no, no second system with alpha proof. No second system. In, text out.”
Assertion Not checkable as stated
Kilpatrick: Gemini's SOTA video performance resulted from reasoning, not video engineering
“With reasoning is a great example of this where like multimodal with video understanding ended up like having this huge, like it's having this beautiful moment.
The model is like soda out of the box because of all the reasoning capabilities that were baked in…”
Disclosure
Kilpatrick: Google's strategy is building Gemini as one single unified model
“Like we're here to make one model and like that model is Gemini.
And like, I think you do need to just trust this point, like to make the capabilities work in some cases, like you do need to have these forks that like go off and make that capability and harde…”
Prediction Not checkable as stated
Howard: Reka's model is probably superior to GPT and Claude for certain tasks
“There's a whole model that's been trained in a different way. So there's probably a whole lot of tasks it's probably better at than you know, GPT and Gemini and Claude.”
Opinion
Feinberg: Gemini is obviously worse at coding despite benchmark wins
“So Gemini does pretty well on Sweebench. Sometimes Gemini publishes models that win on some of those software benchmarks. Raise your hand if you're using Gemini to write code right now instead of, you know, the obvious other name competitors. No one. Like, why…”
Insight
Dean: Adding training data for hundreds of languages displaces other model capabilities
“We're always making these kind of you know, trade-offs in the data mix that we train the base Gemini models on. You know, we'd love to include Data from 200 more languages and as much data as we have for those languages. But that's going to displace some other…”
Disclosure
Dean: A one-page internal memo sparked the Gemini unification effort
“I actually wrote a one-page memo saying we were being stupid by fragmenting our resources. So in particular at the time we had you know efforts within Google research on and in the brain team in particular on large language models. We also had efforts on multi…”
Disclosure
Jeff Dean: Gemini was designed to ingest Waymo LIDAR and robotics telemetry
“I think one of the things about Gemini's multimodal aspects is we've always wanted it to be multimodal from the start. And so, you know, that sometimes to people means text and images and video sort of human-like and audio, audio, human-like modalities, but I …”
Assertion Not checkable as stated
Gemini's IMO Model Checkpoint Required Only One Week of Training
“The training process of this IMO model itself was, like, maybe a week or so.”
Opinion
Davis: Claude Deep Research Outperforms OpenAI, Perplexity, and Gemini
“And time and time again, over the last couple of weeks, I found that Claude has by far outperformed the others. And I guess the definition of good for me right now is not just length, but also the number of sources and diversity of response.”
Disclosure
Gemini 1 Proved GPT-4-Level Models Can Bootstrap Reinforcement Learning
“Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were basically at GPT-IV level or above, were capable en…”
Assertion Partly supported
Swyx: Claude Sonnet and Gemini Outperform o1-Preview in Coding
“Claude Sonnet so far is beating O-one on coding tasks without At least one preview without being a reasoning model and same for Gemini pro or Gemini two point O.”
Assertion Supported
Molmo-72B Ranks Second Behind Only GPT-4o in Human Preference Elo
“Most preference ranked was GPT-Four-O, then Momo-Semety-Two-B, then Gemini, then Sonnet, then the Seven-B.”
Insight
Yi Tay: Meta's Llama is corporate open weights, not grassroots open source
“To me, Lama Tree is like... Meta has an org that is hypothetically very similar to Gemini or something but they just decide to release the weights It's open weights It's open weights and everything”
Opinion
Huang: RULER benchmark is more comprehensive than Gemini's multi-needle test
“I would even argue is more comprehensive than the benchmark that, that Gemini released for their, like, multi-needle in the haystack.”
Assertion Not checkable as stated
Sanseviero: Gemma 3 outperforms stronger general models when fine-tuned on non-English languages
“If you compare Gemma III to other models from back then, maybe the other models were better than Gemma III like as general model, But if you train all of these models for, I don't know, a specific Southeast Asian language, I don't know, Vietnamese, let's say, …”
Assertion Supported
Non-reasoning Grok, GPT, and Gemini models exhibit recursive self-correction loops
“And so then I tried this across models, and I saw consistently across Grok, and GPT, and Gemini, that you were seeing this phenomena where models will, like, self-correct themselves quite a bit, and these were, like, non-thinking models. They were, like, the s…”
Disclosure
DeepMind Singapore is keeping its Gemini RL team small to maximize compute per capita
“We're hiring, like, my team will work on like RL and reasoning for Gemini and Gemini deep thing. I think we care more about like talent density now. So we're not like also like growing that big, this small team first, just because compute per capita is probabl…”
Insight
Swix: ChatGPT Canvas Inverts Google Docs and Gemini's Interface Architecture
“It's basically an inversion of what Google Docs is, wants to do with Gemini. It's like Google Docs on the main screen and then Gemini on the side. And right, whatnot, what ChatGPT has done is Do the chat thing first, and then the docs on the side. But it's kin…”
Prediction Not checkable as stated
Swyx: Always-on vision AI assistants will dominate desktop software by late 2025
“And like this time next year, I would be willing to bet that I would just have this running on my machine. And you know, I think That assistance always on that you can talk to with vision that sees what you're seeing. I think that is where at least one hour so…”
Disclosure
Tay: DeepMind did not optimize Gemini specifically for Pokémon
“There's actually nothing specifically done for Pokemon.”
Assertion Supported
Mallick: Gemini officially supports 24 languages but responds in Klingon
“We officially support 24 languages, but you can try talking to the model and cling on and it'll respond to you.”
Disclosure
Mallick: Google released experimental proactive audio for Gemini native audio
“One of the features that we've pushed out A little more experimental, but would love for people to test it is what we're calling proactive audio, and it's available only in the native audio, in the audio to audio architecture right now. And what this feature d…”
Assertion Supported
Howard: Google Gemini is about to release KV caching support
“Gemini is about to finally come out with KV caching, and this is something that Austin actually and Gemma.cpp had had on his roadmap for years well not years, months, long time is, is that.”
Disclosure
DeepMind Singapore Explicitly Adds AGI to Job Postings
“I think that, like, one reason why we work on these models is that we want to get to AGI, and, like, this was a right thing that we added AGI to the job posting, yeah. There is no, like, formal name of the team yet, but it's basically the Gemini theme Singapor…”