The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 27 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Lukas Petersson Jun 4, 2026 ▶ 46:27 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Prediction Not checkable as stated
Tay: Most specialized tools will be subsumed directly into model parameters
“Then the most I can see in the future is there'll be a model then that, that is, there's something that really cannot be subsumed by a model. Then you just use a tool or something, right? But my prediction is that I think most things can be subsumed by the mod…”
Yi Tay Jan 23, 2026 ▶ 19:31 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Disclosure
DeepMind Abandoned AlphaProof to Run Gemini End-to-End for IMO Math
“We wanted to try to, like, use, actually use Gemini as an end-to-end model. Basically, no, no second system with alpha proof. No second system. In, text out.”
Yi Tay Jan 23, 2026 ▶ 13:38 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Assertion Not checkable as stated
Kilpatrick: Gemini's SOTA video performance resulted from reasoning, not video engineering
“With reasoning is a great example of this where like multimodal with video understanding ended up like having this huge, like it's having this beautiful moment. The model is like soda out of the box because of all the reasoning capabilities that were baked in…”
Logan Kilpatrick Jun 2, 2025 ▶ 13:19 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Disclosure
Kilpatrick: Google's strategy is building Gemini as one single unified model
“Like we're here to make one model and like that model is Gemini. And like, I think you do need to just trust this point, like to make the capabilities work in some cases, like you do need to have these forks that like go off and make that capability and harde…”
Logan Kilpatrick Jun 2, 2025 ▶ 12:25 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Prediction Not checkable as stated
Howard: Reka's model is probably superior to GPT and Claude for certain tasks
“There's a whole model that's been trained in a different way. So there's probably a whole lot of tasks it's probably better at than you know, GPT and Gemini and Claude.”
Jeremy Howard Aug 17, 2024 ▶ 36:42 Answer.ai & AI Magic with Jeremy Howard
Opinion
Feinberg: Gemini is obviously worse at coding despite benchmark wins
“So Gemini does pretty well on Sweebench. Sometimes Gemini publishes models that win on some of those software benchmarks. Raise your hand if you're using Gemini to write code right now instead of, you know, the obvious other name competitors. No one. Like, why…”
Evan Feinberg Jun 30, 2026 ▶ 55:32 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Insight
Dean: Adding training data for hundreds of languages displaces other model capabilities
“We're always making these kind of you know, trade-offs in the data mix that we train the base Gemini models on. You know, we'd love to include Data from 200 more languages and as much data as we have for those languages. But that's going to displace some other…”
Jeff Dean Feb 12, 2026 ▶ 53:24 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Disclosure
Dean: A one-page internal memo sparked the Gemini unification effort
“I actually wrote a one-page memo saying we were being stupid by fragmenting our resources. So in particular at the time we had you know efforts within Google research on and in the brain team in particular on large language models. We also had efforts on multi…”
Jeff Dean Feb 12, 2026 ▶ 1:07:48 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Disclosure
Jeff Dean: Gemini was designed to ingest Waymo LIDAR and robotics telemetry
“I think one of the things about Gemini's multimodal aspects is we've always wanted it to be multimodal from the start. And so, you know, that sometimes to people means text and images and video sort of human-like and audio, audio, human-like modalities, but I …”
Jeff Dean Feb 12, 2026 ▶ 16:56 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
Assertion Not checkable as stated
Gemini's IMO Model Checkpoint Required Only One Week of Training
“The training process of this IMO model itself was, like, maybe a week or so.”
Yi Tay Jan 23, 2026 ▶ 16:16 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Opinion
Davis: Claude Deep Research Outperforms OpenAI, Perplexity, and Gemini
“And time and time again, over the last couple of weeks, I found that Claude has by far outperformed the others. And I guess the definition of good for me right now is not just length, but also the number of sources and diversity of response.”
Dylan Davis Jul 5, 2025 ▶ 2:57 ⚡️Anthropic vs Cognition on Multi-Agents: A Breakdown with Dylan Davis
Disclosure
Gemini 1 Proved GPT-4-Level Models Can Bootstrap Reinforcement Learning
“Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were basically at GPT-IV level or above, were capable en…”
Misha Laskin Mar 7, 2025 ▶ 4:34 Solve coding, solve AGI [Reflection.ai launch w/ CEO Misha Laskin]
Assertion Partly supported
Swyx: Claude Sonnet and Gemini Outperform o1-Preview in Coding
“Claude Sonnet so far is beating O-one on coding tasks without At least one preview without being a reasoning model and same for Gemini pro or Gemini two point O.”
Shawn Wang Jan 1, 2025 ▶ 37:15 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Supported
Molmo-72B Ranks Second Behind Only GPT-4o in Human Preference Elo
“Most preference ranked was GPT-Four-O, then Momo-Semety-Two-B, then Gemini, then Sonnet, then the Seven-B.”
Vibhu Sapra Oct 13, 2024 ▶ 39:17 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Insight
Yi Tay: Meta's Llama is corporate open weights, not grassroots open source
“To me, Lama Tree is like... Meta has an org that is hypothetically very similar to Gemini or something but they just decide to release the weights It's open weights It's open weights and everything”
Yi Tay Jul 5, 2024 ▶ 1:59:19 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Opinion
Huang: RULER benchmark is more comprehensive than Gemini's multi-needle test
“I would even argue is more comprehensive than the benchmark that, that Gemini released for their, like, multi-needle in the haystack.”
Mark Huang May 31, 2024 ▶ 44:18 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Assertion Not checkable as stated
Sanseviero: Gemma 3 outperforms stronger general models when fine-tuned on non-English languages
“If you compare Gemma III to other models from back then, maybe the other models were better than Gemma III like as general model, But if you train all of these models for, I don't know, a specific Southeast Asian language, I don't know, Vietnamese, let's say, …”
Omar Sanseviero May 24, 2026 ▶ 8:50 ⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
Assertion Supported
Non-reasoning Grok, GPT, and Gemini models exhibit recursive self-correction loops
“And so then I tried this across models, and I saw consistently across Grok, and GPT, and Gemini, that you were seeing this phenomena where models will, like, self-correct themselves quite a bit, and these were, like, non-thinking models. They were, like, the s…”
Pratyush Maini Feb 10, 2026 ▶ 7:08 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Disclosure
DeepMind Singapore is keeping its Gemini RL team small to maximize compute per capita
“We're hiring, like, my team will work on like RL and reasoning for Gemini and Gemini deep thing. I think we care more about like talent density now. So we're not like also like growing that big, this small team first, just because compute per capita is probabl…”
Yi Tay Jan 23, 2026 ▶ 1:24:52 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Insight
Swix: ChatGPT Canvas Inverts Google Docs and Gemini's Interface Architecture
“It's basically an inversion of what Google Docs is, wants to do with Gemini. It's like Google Docs on the main screen and then Gemini on the side. And right, whatnot, what ChatGPT has done is Do the chat thing first, and then the docs on the side. But it's kin…”
Shawn Wang Feb 1, 2025 ▶ 36:56 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Prediction Not checkable as stated
Swyx: Always-on vision AI assistants will dominate desktop software by late 2025
“And like this time next year, I would be willing to bet that I would just have this running on my machine. And you know, I think That assistance always on that you can talk to with vision that sees what you're seeing. I think that is where at least one hour so…”
Shawn Wang Jan 1, 2025 ▶ 1:32:43 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Disclosure
Tay: DeepMind did not optimize Gemini specifically for Pokémon
“There's actually nothing specifically done for Pokemon.”
Yi Tay Jan 23, 2026 ▶ 27:10 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Assertion Supported
Mallick: Gemini officially supports 24 languages but responds in Klingon
“We officially support 24 languages, but you can try talking to the model and cling on and it'll respond to you.”
Shrestha Basu Mallick Jun 2, 2025 ▶ 23:28 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Disclosure
Mallick: Google released experimental proactive audio for Gemini native audio
“One of the features that we've pushed out A little more experimental, but would love for people to test it is what we're calling proactive audio, and it's available only in the native audio, in the audio to audio architecture right now. And what this feature d…”
Shrestha Basu Mallick Jun 2, 2025 ▶ 20:19 [AIEWF Preview] Gemini in 2025 and Realtime Voice AI
Assertion Supported
Howard: Google Gemini is about to release KV caching support
“Gemini is about to finally come out with KV caching, and this is something that Austin actually and Gemma.cpp had had on his roadmap for years well not years, months, long time is, is that.”
Jeremy Howard Aug 17, 2024 ▶ 1:06:34 Answer.ai & AI Magic with Jeremy Howard
Disclosure
DeepMind Singapore Explicitly Adds AGI to Job Postings
“I think that, like, one reason why we work on these models is that we want to get to AGI, and, like, this was a right thing that we added AGI to the job posting, yeah. There is no, like, formal name of the team yet, but it's basically the Gemini theme Singapor…”
Yi Tay Jan 23, 2026 ▶ 1:10 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.