The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 41 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 41 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Open · timeframe Sep 2029
Anandkumar: Multi-physics models outperform single-physics models of equivalent parameter size
“And in fact, I was going to add that it turns out that having the model of the same size with multiple areas of physics does better than giving all of those parameters to each single physics. So if you had separate models and made them big enough as the origin…”
Anima Anandkumar Sep 4, 2026 ▶ 8:08 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Assertion Open · timeframe Jun 2027
Kolter: Gray Swan's Shade system outperforms human red teamers at breaking models
“However, one thing that we are finding, and this is actually, I think we're kind of crossing this point too. Is that in a lot of the latest experiments, we can do much better than people, than human red teamers now at breaking these models. When I say we, I me…”
Zico Kolter Jun 22, 2026 ▶ 12:14 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Open · timeframe Dec 2026
Up to 20% of US data centers risk cancellation from community backlash
“Up to 20% of all data centers this year in the US, my understanding is are at risk... Of not getting the community support they need to get brought up.”
Anjney Midha Jun 18, 2026 ▶ 5:19 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Assertion Open · timeframe Jun 2029
Petersson: Telling AI It Is in a Simulation Increases Bad Behavior
“One ablation we did run in, in, in Vending Bench was that we said like we added like, you're in a simulation, your actions doesn't affect anyone. And then it became even more crazy or like it did even more bad stuff.”
Lukas Petersson Jun 4, 2026 ▶ 56:50 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Open · timeframe Mar 2027
Eskildsen: Turbopuffer outperforms Lucene on long LLM search queries
“Turbo Puffer today has a fairly start of the state of the art full text search engine. We beat Lucene on some queries, in particular, very long queries that we've optimized for, because those are the text search queries we see today.”
Simon Eskildsen Mar 12, 2026 ▶ 51:26 Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer
Assertion Open · timeframe Feb 2027
Modern LLMs verbatim regurgitate JEE exam questions from two-word prompts
“We consistently saw how many of these, like, models today are being, like, massively, like, kind of fine-tuned on problems from... Like, oversight? Massively worked with. Like, even, like, imagine if I ask you the light bulb, what comes next in your mind? It w…”
Pratyush Maini Feb 10, 2026 ▶ 3:36 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Not yet assessed · timeframe Jan 2026
White: Synthesis routes for dangerous compounds are already available on Wikipedia
“You can go find the synthesis route for many dangerous compounds on Wikipedia. People know what are the targets in the human body that, like, are targeted by most biological weapons. It's not really that much of a mystery.”
Andrew White Jan 28, 2026 ▶ 47:05 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Assertion Open · timeframe Sep 2026
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”
Diego Bachman Sep 23, 2025 ▶ 7:46 ⚡️ Beyond Transformers with Power Retention
Assertion Open · timeframe Sep 2028
Bachman: Power Retention models match original base model performance
“They'll come out with a nice shiny new, a power retention architecture that has the same performance on whatever data set they want as the original base model did.”
Diego Bachman Sep 23, 2025 ▶ 24:00 ⚡️ Beyond Transformers with Power Retention
Assertion Open · timeframe Aug 2028
Morcos: Datology matches DCLM performance 12x faster with under 10% tokens
“We're able to now get to the same performance as DCLM about 12 x faster. So, you know, in fewer than 10% of the tokens we can match What you get from training to convergence.”
Ari Morcos Aug 29, 2025 ▶ 29:43 Better Data is All You Need — Ari Morcos, Datology
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Thomas Sohmers Aug 18, 2025 ▶ 16:28 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Open · timeframe Aug 2026
Krentsel: Exo autonomously modified its code to inspect Pokémon game RAM
“We've had XO running, playing, playing Pokemon. And while it's running, the system itself decided to try inspecting the like RAM of the game and then went and mapped the RAM to, and people have reversed in the past, people have reverse engineered this manually…”
Alex Krentsel Aug 15, 2026 ▶ 11:41 Exo: Harnesses should see their own code and logs — Alex Krentsel, UC Berekeley / Google Research
Assertion Open · timeframe Aug 2026
Krentsel: Exo autonomously re-architected its Discord adapter, cutting costs by 96%
“We asked it, Hey, I noticed, I asked, Hey, but how much did the last message cost in the discord adapter? And it was like, it was. It's like, are you serious? 16 cents. That's actually crazy. Like. Go work on driving that down. And so it went and re-architecte…”
Alex Krentsel Aug 15, 2026 ▶ 39:44 Exo: Harnesses should see their own code and logs — Alex Krentsel, UC Berekeley / Google Research
Assertion Open · timeframe Jun 2029
Hong: Axiom's unmodified Putnam system achieved 99% on Verina benchmark
“And we actually recently, with no modification to the Putnam system, we saw a 99% out of the 189 problems, we saw a 187, we missed only two code-wisp-proof.”
Carina Hong Jun 3, 2026 ▶ 29:46 Scaling Past Informal AI - Carina Hong, Axiom Math
Assertion Open · timeframe Jun 2029
Ethan He: Grok Imagine Video Extension Tracks Full Historical Context
“So the Glock Imagine video extension, it has historical context of all of the previous generated videos. It can it has a context of who is speaking and what objects have appeared and everything having that to generate the next video.”
Ethan He Jun 1, 2026 ▶ 55:32 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Assertion Open · timeframe May 2027
Daytona spins up a single agent sandbox in 60ms
“And so our time to spin up one is 60 milliseconds with network agency. So requests, spin up, reply, 60, the whole thing, 60 milliseconds.”
Ivan Burazin May 21, 2026 ▶ 17:40 AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
Assertion Open · timeframe May 2027
Daytona can spin up 50,000 concurrent sandboxes in 75 seconds
“But if you want to spin up 50,000 at once, we are now at about 75 seconds. So it takes about 75 seconds to spin up concurrently 50,000.”
Ivan Burazin May 21, 2026 ▶ 17:49 AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
Assertion Open · timeframe Jan 2026
White: ChemCrow paper was presented to U.S. President in 30-minute block
“I ended up visiting the white house. I guess my paper was like the only time a preprint or peer review paper was presented to the president on like their schedule for like a 30 minute block.”
Andrew White Jan 28, 2026 ▶ 10:21 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White
Assertion Open · timeframe Sep 2028
Bachman: StarCoder-3B converted to Power Retention matches baseline loss in two hours
“After just 10,000 steps of training, which this training one took about two hours, this orange curve, you see that it fully matches the original loss.”
Diego Bachman Sep 23, 2025 ▶ 16:11 ⚡️ Beyond Transformers with Power Retention
Assertion Open · timeframe Aug 2026
Sohmers: Positron AI hardware achieves 93% of theoretical memory bandwidth
“And so our fundamental architecture is enabling us, you know, today with hardware that we're shipping right now to be achieving, you know, 93% of the theoretical memory bandwidth of our device consistently across all use cases.”
Thomas Sohmers Aug 18, 2025 ▶ 15:26 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Open · timeframe May 2025
Huang: PoSE breaks down on needle-in-a-haystack at 500k tokens
“It does start to break down a little bit more on the longer, longer context. So, like, 500,000 to a million it appeared that it doesn't hold as well specifically for, like, needle in the haystack.”
Mark Huang May 31, 2024 ▶ 24:11 How to train a Million Context LLM — with Mark Huang of Gradient.ai
Assertion Open · timeframe Sep 2026
Slack: Sourcegraph serves nine of top ten public tech companies
“We have like nine of the 10 top Public tech companies as customers and like four of the six top banks and like Uber and Stripe and so on, all these companies using Sourcegraph for code search.”
Quinn Slack Sep 7, 2026 ▶ 28:39 Orbs: Shifting Coding to Cloud — Quinn Slack, Amp Code
Assertion Open · timeframe Jun 2026
Fredrikson: Skilled red teamers phish human participants 60% to 70% of the time
“But for a skilled, like, human red teamer, they could fish the human participants, like, with the 60 to 70% success.”
Matt Fredrikson Jun 22, 2026 ▶ 22:20 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Open · timeframe Jun 2026
Fredrikson: Top AI browser agents yielded only a handful of successful breaks
“There were a couple of models that seemed to be very, very robust, right? Like the red teamers found just a handful of successful breaks on them.”
Matt Fredrikson Jun 22, 2026 ▶ 22:29 AI Security After Codex and Claude Code — Zico Kolter & Matt Fredrikson, Gray Swan
Assertion Open · timeframe Jun 2029
Midha: MatX chips adopt NVIDIA reference architecture to plug into existing sites
“When they decided to pick the standard for their data center, they picked the NVIDIA reference architecture. So the Matex chips just plug in to any site that has an NVIDIA bring up planned. And you know.”
Anjney Midha Jun 18, 2026 ▶ 31:28 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Assertion Open · timeframe May 2027
The United Nations uses Chatbase on Facebook Messenger for regional crisis support
“Like, the UN is using us right now on their Facebook messenger. So when people reach out to them, specifically for, like, specific regions, it gets drafted to Chatbase, and Chatbase helps them.”
Yasser Elsaid May 2, 2026 ▶ 23:20 ⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
Assertion Open · timeframe Apr 2029
Sun: Moonlake can generate multiplayer environments and persistence databases via prompting
“So if you just actually just like prompt our Model to say, hey, like configure the multiplayer, then it'll do like this. You'll be able to configure multiplayer. Persistency database for you.”
Fan-yun Sun Apr 2, 2026 ▶ 27:03 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Assertion Open · timeframe Feb 2027
Qwen 3 memorizes benchmark questions significantly more than Qwen 1.5 or 2
“I don't, we don't see this phenomenon like the earlier versions of like QN 1.5 or even QN two, but start seeing it in QN three. So there's something about a more, like a significantly higher weight on benchmark or like a J or like any times of evaluation quest…”
Pratyush Maini Feb 10, 2026 ▶ 4:59 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Assertion Not yet assessed · timeframe Dec 2025
Altman, Nadella, and Pichai publicly committed to adopting MCP around April
“And then like, you had this like inflection point around April with like Sam Altman and Satya and Sundar and all posting about like MCP and that they're going to adopt MCP at Microsoft, at Google. At OpenAI and that was really like the big inflection point.”
David Soria Parra Dec 28, 2025 ▶ 1:44 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Assertion Open · timeframe Nov 2028
Chan: CZI Billion Cell Project takes months at fraction of historical cost
“Now we're doing the billion cell project and that is taking months and at a fraction of the price.”
Priscilla Chan Nov 6, 2025 ▶ 15:19 Priscilla Chan and Mark Zuckerberg: Frontier AI + Virtual Biology To Solve All Diseases
Assertion Open · timeframe Oct 2026
Merrill: AI models now reliably solve Terminal-Bench's ML training task
“Unfortunately we are getting to the point where models do reliably get this one.”
Mike Merrill Oct 18, 2025 ▶ 11:36 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Open · timeframe Sep 2028
Bachman: PowerCoder-3B reaches 35% HumanEval accuracy versus StarCoder's 30%
“In the end, this converges to, I believe, about 35% accuracy on human eval, whereas the star coder baseline was about 30%.”
Diego Bachman Sep 23, 2025 ▶ 17:33 ⚡️ Beyond Transformers with Power Retention
Assertion Open · timeframe Jun 2026
Vibhu: Gemma Activates Abstract Behavioral Traits Over Simple Token Completion
“It also shows internally that there's more than just token completion of, you know, this plus this equals this. No, it has some under understanding of characteristics, right? Like this is a pretty stubborn dog. It has a stubborn feature. Pretty high up that ac…”
Vibhu (Viboo) Jun 6, 2025 ▶ 21:47 The Utility of Interpretability — Emmanuel Amiesen
Assertion Open · timeframe Apr 2028
GPT-4.1 Nano and Mini are new pre-trains; base 4.1 is mid-train
“Nano is obviously a new pre-train. We also have a new pre-train for Mini, and then, ah, the larger version is, ah, a new mid-train.”
Michelle Pokrass Apr 15, 2025 ▶ 7:46 GPT 4.1: The New OpenAI Workhorse
Assertion Open · timeframe Aug 2029
Park: Simile AI agents can autonomously navigate live website URLs
“Some of the things that our agents can also do is you can be given a domain, like a website URL and actually go use it for a while.”
Joon Sung Park Aug 21, 2026 ▶ 54:02 Simulating Humanity: from Generative Agents to 8 Billion Digital Twins — Joon Sung Park, Simile AI
Assertion Open · timeframe Aug 2029
Midjourney's David Holtz explored text diffusion to storyboard entire movies
“David Holtz from Midjourney was investing in text diffusion. I don't think anything came out of it, but like the idea was that you can storyboard a long movie and then you can generate the scenes with video, normal video gen.”
Shawn Wang Aug 3, 2026 ▶ 1:27:18 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Assertion Open · timeframe Jul 2026
Bubna: Ramp trained custom tokenizers to swap into LLaMA
“Ramp actually early in the day was training their own tokenizer and, like, Swapping out the tokenizer in Lama and whatnot.”
Akshat Bubna Jul 8, 2026 ▶ 51:11 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Assertion Open · timeframe Jun 2026
Krause: DARPA and GE Aerospace synthesized 500 alloys in 12 months
“The largest alloys program was the mock program. It was run by DARPA NGE Aerospace. They did 500 alloys in about 12 months. They did a bunch of kind of AI and simulations on the front end of that, and then they synthesized 500 new alloys in that whole year.”
Joseph Krause Jun 17, 2026 ▶ 42:49 🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
Assertion Not yet assessed · timeframe Dec 2024
Swix: Bolt, Devin, and AI Agent Startups Rely on Netlify Deployments
“Both Bolt and Cognition DevIn and a bunch of other sort of agent type startups, they all use Nullify to deploy because of this one feature.”
Shawn Wang Dec 2, 2024 ▶ 18:31 0 to over $8M ARR in 2 months as a Claude Wrapper (Bolt.new, Qodo)
Assertion Not yet assessed · timeframe Oct 2024
BFCL v3 generates evaluation tasks using graph edge construction
“They basically created their own API for the sake of this testing, and then did this like mapping to create a graph edge construction and like generate tasks through that.”
Sam Julien Oct 5, 2024 ▶ 10:46 [Paper Club] Berkeley Function Calling Paper Club! — Sam Julien, Writer
Assertion Open · timeframe Aug 2027
Carlini encodes 1.44 megabytes of data onto a single sheet of paper
“Yeah, okay. So it's about, in particular, it's about 1.44 megabytes.”
Nicholas Carlini Aug 28, 2024 ▶ 36:17 Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.