The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 1,786 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 41 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Bakouch: Novel optimizer speedups are exaggerated due to undertuned AdamW baselines
“And what they find is that the speed up is greatly, greatly exaggerated. And mostly because often people like undertone the Adam W baseline.”
Elie Bakouch Oct 20, 2025 ▶ 16:06 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Contradicted
Swix: Every frontier lab now distills dense models into MoEs
“I think like, I think this is the pattern for every frontier lab now.”
Shawn Wang Oct 20, 2025 ▶ 36:32 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Not checkable as stated
Bakouch: DeepSeek and Qwen do not release all their ablation data
“We want to train our MOE because it's fun and everyone is doing that. And also I think there is a lot of different direction. And it's always good in terms of science. To, because basically the coin tree or even deep seek, they don't release, release all the a…”
Elie Bakouch Oct 20, 2025 ▶ 59:13 ⚡ Open Model Pretraining Masterclass — Elie Bakouch, HuggingFace SmolLM 3, FineWeb, FinePDF
Assertion Supported
Corbitt: OpenPipe Beat Frontier Models Using a Qwen 32B Judge
“One of the results we published was we used Quen 2.5 14 B as the model we're training, and as the judge we used Quen 2.5 32 B, which is, like, Not, I mean, it's fine, but it's like not a, it's much worse than any frontier model. Right. And even with that combi…”
Kyle Corbitt Oct 16, 2025 ▶ 53:18 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Assertion Not checkable as stated
Lenz: Local smartphone AI requires hybrid models due to KV cache limits
“So if you wanted to do something local on your phone to search your images, as an example, you can't do that without a hybrid architecture or without doing drastically changes because the model plus KVCache won't fit.”
Barak Lenz Oct 11, 2025 ▶ 13:09 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Assertion Not checkable as stated
Howard: Answer.AI runs fully in-house stack without AWS or Google Cloud
“This group of, which has averaged about 10 to 12 people, currently nine, I think, have built a pretty Transformational and complex piece of software, which we can do a quick demo of later if you're interested. Using a complete web application development platf…”
Jeremy Howard Oct 2, 2025 ▶ 4:28 The antidote to AI fatigue — Answer.ai Solveit
Assertion Supported
Feldman: Cerebras provides 2,625x more memory bandwidth than traditional GPUs
“And we have 2625 times more memory bandwidth than the GPU does.”
Andrew Feldman Oct 1, 2025 ▶ 6:22 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Assertion Supported
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Andrew Feldman Oct 1, 2025 ▶ 11:03 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Assertion Not checkable as stated
Feldman: AI startups are replacing closed-source models with fine-tuned open-source
“I think for sort of AI companies like Cognition, like all your competitors, like AlphaSense, like dozens of others, they are trying to replace closed source models with very, very fast open source models, and they're trying to drive the open source Accuracy dr…”
Andrew Feldman Oct 1, 2025 ▶ 14:20 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Assertion Not checkable as stated
Slack: No AI developer tool has retained developers past 12 months
“There's not been any tool that has stuck with devs for more than, Six or 12 months or something.”
Quinn Slack Sep 25, 2025 ▶ 3:38 Amp: The Emperor Has No Clothes
Assertion Not checkable as stated
Fanelli: Cursor Default Model Switch Cost Anthropic $200M in ARR
“When cursors switch from Sonnet to GPT-V as like the default model that was like, you know, Two hundred million our revenue for Anthropic that kind of went away and like moved on to GPT-V.”
Alessio Fanelli Sep 25, 2025 ▶ 29:57 Amp: The Emperor Has No Clothes
Assertion Not checkable as stated
Slack: Competitors discounted products up to 100% to win deals against Amp
“So we've had one head to head loss with Amp where we lost against the usual players. And the reason why is one of them discounted their other product a hundred percent for two years. The other one discounted at 85% for two years, which is just crazy.”
Quinn Slack Sep 25, 2025 ▶ 38:21 Amp: The Emperor Has No Clothes
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Shreya Rajpal Sep 25, 2025 ▶ 23:21 ⚡️Snowglobe: Simulations for your AI
Assertion Open · timeframe Sep 2026
Bachman: Power Retention Delivers 100x Inference Speedup at 64k Context
“And at 64 K tokens, We get something like a 10 X speed up at training, but at inference time, because you're not only saving flops at inference time, but also paging in and out of memory of the KV cache, you actually get a hundred X speed ups from power retent…”
Diego Bachman Sep 23, 2025 ▶ 7:46 ⚡️ Beyond Transformers with Power Retention
Assertion Open · timeframe Sep 2028
Bachman: Power Retention models match original base model performance
“They'll come out with a nice shiny new, a power retention architecture that has the same performance on whatever data set they want as the original base model did.”
Diego Bachman Sep 23, 2025 ▶ 24:00 ⚡️ Beyond Transformers with Power Retention
Assertion Not checkable as stated
Taskaya: NSFW content makes up less than 1% of Fal traffic
“Moderation is optional to a level where, like, illegal content is moderated, and we also track, like, the non-illegal content NSFW moderation, and, like, we haven't seen, like, we haven't seen more than one percent.”
Batuhan Taskaya Sep 8, 2025 ▶ 43:03 A Technical History of Generative Media
Assertion Supported
Morcos: Soft inductive biases become harmful past 1M data points in vision
“Turns out in the small data regime, and when I say small data here, I mean, say less than 500,000 data points. And this was in the context of image self-supervised learning. So in that small data regime, this is super helpful. And where this paper's actually b…”
Ari Morcos Aug 29, 2025 ▶ 6:04 Better Data is All You Need — Ari Morcos, Datology
Assertion Supported
Morcos: Kaplan and Chinchilla scaling laws incorrectly assume all data is equal
“And even if you go and you look at the scaling laws work from Kaplan and Chinchilla and all these other things, they all assume IID data which is insane. We know that all data are not created equal, that garbage in garbage out is like the oldest adage in compu…”
Ari Morcos Aug 29, 2025 ▶ 8:25 Better Data is All You Need — Ari Morcos, Datology
Assertion Contradicted
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Ari Morcos Aug 29, 2025 ▶ 18:06 Better Data is All You Need — Ari Morcos, Datology
Assertion Supported
Morcos: Proper data curation can bend neural scaling laws
“And what that paper showed was that if you use your data correctly, you can actually bend the scaling laws themselves.”
Ari Morcos Aug 29, 2025 ▶ 27:35 Better Data is All You Need — Ari Morcos, Datology
Assertion Open · timeframe Aug 2028
Morcos: Datology matches DCLM performance 12x faster with under 10% tokens
“We're able to now get to the same performance as DCLM about 12 x faster. So, you know, in fewer than 10% of the tokens we can match What you get from training to convergence.”
Ari Morcos Aug 29, 2025 ▶ 29:43 Better Data is All You Need — Ari Morcos, Datology
Assertion Not checkable as stated
Morcos: Qwen is much easier to align than Llama due to pre-training
“It's much easier to RL Quen than it is to do Lama. Likely that has to do with the fact that Quen put a lot of synthetic reasoning traces into their training data.”
Ari Morcos Aug 29, 2025 ▶ 54:07 Better Data is All You Need — Ari Morcos, Datology
Assertion Not checkable as stated
Morcos: Yann LeCun was never defining Meta's AI strategy
“I don't think he was ever you know, or at least not since the beginning in a role where he was defining AI strategy for Meta. I don't think that's the role he wanted at any point. You know, I think he really wanted to be doing that research, and I think, so I …”
Ari Morcos Aug 29, 2025 ▶ 1:16:13 Better Data is All You Need — Ari Morcos, Datology
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Thomas Sohmers Aug 18, 2025 ▶ 16:28 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Supported
Sohmers: Positron AI requires zero compilers to run Hugging Face models
“So rather than having like, we don't have a compiler whatsoever. There's no compiler. There's no translator, no tooling that's involved in actually taking those and getting that to, you know, for your common, you know, Huggy Face Transform models to be able to…”
Thomas Sohmers Aug 18, 2025 ▶ 21:32 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Contradicted
Sohmers: Google Veo and Imagen 3 are pure autoregressive transformers, not diffusion
“A lot of things have actually been moving away from diffusion to being pure autoregressive transformers for image and video generation. So like the latest, yeah, there's a VO three and since image and three on, on Google side have been pure autoregressive movi…”
Thomas Sohmers Aug 18, 2025 ▶ 43:48 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Not checkable as stated
Brockman: AI models are reaching parameter counts comparable to human synapses
“It's a hundred T synapses, which kind of corresponds to the weights of the neural net. And so there's some sort of equivalence there. And so we're starting to get to the right numbers. Let me just say that.”
Greg Brockman Aug 15, 2025 ▶ 16:11 Greg Brockman on OpenAI's Road to AGI
Assertion Not checkable as stated
Brockman: Physicists say GPT-5 re-derived research insights taking months of work
“We've seen physicists starting to kick the tires on GPT-V and say that, like, hey, this thing was able to get, this model was able to re-derive an insight that took me many months worth of research to produce.”
Greg Brockman Aug 15, 2025 ▶ 21:50 Greg Brockman on OpenAI's Road to AGI
Assertion Partly supported
The Information: OpenAI hit $12B ARR as burn rose to $8B
“We had a story yesterday about open AI and how, like, I think they've reached about twelve billion ARR and yeah, but their burn went from like They projected, like, one billion to, like, eight billion or something.”
Stephanie Palazzolo Aug 6, 2025 ▶ 40:26 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Assertion Supported
Palazzolo: Claude Code leads stayed at Cursor only two weeks
“We know that they went there, they were there for, I think, about two weeks, and they came back.”
Stephanie Palazzolo Aug 6, 2025 ▶ 41:45 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Assertion Not checkable as stated
Palazzolo: VCs are growing frustrated with unconventional AI acqui-hire deals
“Talking to investors, I think they're starting to, I think they're kind of over these types of deals. Like, they're like, hey, I mean, like, this is, like, fine, but, like, we don't invest in companies so they can get, like, a weird acquihire situation a coupl…”
Stephanie Palazzolo Aug 6, 2025 ▶ 46:05 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Assertion Supported
Ermon: Diffusion LLMs Pareto-dominate autoregressive models on inference efficiency
“On the inference side, what we're seeing is that diffusion models are much more efficient. We're actually able to Pareto dominate autoregressive models. If you think about the typical trade-off between throughput versus latency, which you kind of like cannot, …”
Stefano Ermon Aug 4, 2025 ▶ 14:30 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Assertion Partly supported
Inception generalist model matches Claude Haiku quality at 5-10x speed
“We had our generalist model evaluated by artificial analysis and the intelligence score from AA artificial analysis around 40. So it's comparable to GPT, 4.1 nano, cloud haiku, kind of like Close source speed optimized models. It's roughly comparable in terms …”
Stefano Ermon Aug 4, 2025 ▶ 16:55 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Not checkable as stated
Swix: OpenAI Deep Research was built by three people as an o3 wrapper
“As far as I know, it's three people did it. It was Isa and like the two other collaborators that she had. I don't know if they did that much on top of all three, like every indication I've had from over the eye is that deep research is more or less a thin wrap…”
Shawn Wang Jul 31, 2025 ▶ 25:36 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Nathan Lambert Jul 31, 2025 ▶ 1:16:23 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Supported
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Brendan Fortuna Jul 29, 2025 ▶ 24:01 ⚡️Using RFT to Build Clinical Superintelligence
Assertion Not checkable as stated
Mohan: Codeium quality matches Copilot and drives user churn
“The product is actually one of those products where even use Copilot and use us, it's hard to tell the difference actually. And a lot of our users have actually churned off of Copilot.”
Varun Mohan Jul 28, 2025 ▶ 17:08 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Assertion Not checkable as stated
Hou: Codeium inference costs 1/100th of competitors by avoiding third-party APIs
“It's that idea that our computation is one 100th of the cost of the competitors. We are not using APIs, and as a result, our customers and our users actually get 100 X the amount of compute that they would on another product.”
Kevin Hou Jul 28, 2025 ▶ 1:05:22 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Assertion Not checkable as stated
Wu: Autonomous coding agent capability currently doubles every 70 days
“What you see in general is that that doubling time is about every seven months, which already is pretty crazy, actually, but in code, it's actually even faster. It's every 70 days, which is two or three months, and so, you know, if you look at various software…”
Scott Wu Jul 28, 2025 ▶ 3:04:27 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Assertion Not checkable as stated
Hou: Windsurf's SWE-1 achieves near-frontier model results at lower cost
“And we've been able to achieve near frontier model results at the fraction of the cost, and with a significantly smaller team.”
Kevin Hou Jul 28, 2025 ▶ 3:31:20 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Assertion Not checkable as stated
Hou: Users choose SWE-1 at higher frequencies than Claude 3.7 and 3.5
“People are choosing SWE-ONE because it recognizes how they do work, not necessarily how to generate code. And it's contributing, actually, an even higher frequency than models like 3.7 and 3.5.”
Kevin Hou Jul 28, 2025 ▶ 3:31:45 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Assertion Not checkable as stated
Scott Wu: No customer or training data was shared in Windsurf transaction
“There was no, no, no information that was given out, for example, in terms of customer data, training data, any things like that. And so, so, you know, all of that is, is, is strictly proprietary and then remains, you know, our exclusive access.”
Scott Wu Jul 28, 2025 ▶ 3:51:39 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Assertion Supported
OpenAI's IMO performance was not officially verified by the IMO
“It turns out, like, OpenAI actually didn't involve officially with IMO. They just, like, use the problems, but, and then just, like, use their model to test the results, and ask, like, three previous IMO analysts to review them.”
Dr. Jasper Zhang Jul 24, 2025 ▶ 3:28 ⚡️Math Olympiad gold medalist explains OpenAI and Google DeepMind IMO Gold Performances
Assertion Not checkable as stated
McCloy: Gray-Hat SEO Tactics Currently Work in AI Search
“In terms of like some of the more gray area SEO tactics like that, you know, Penguin was addressing with Google back in the day. I mean, I think the reality is a lot of those things do work today. Like I can't tell you they don't work, but I would say as like,…”
Robert McCloy Jul 23, 2025 ▶ 25:07 AI is Eating Search
Assertion Supported
McCloy: ChatGPT Does Not Index or Retrieve llms.txt by Default
“I knew that there's debate about this, but I'd say the evidence is like ChatTriPT is not indexing and it's not retrieving content from LMS.txt by default.”
Robert McCloy Jul 23, 2025 ▶ 38:52 AI is Eating Search
Assertion Not checkable as stated
Kamradt: xAI Delaying Coder Model Release To Beat Specific Rival
“I heard rumors that, that Grok doesn't want to release the coding model until it's better than one specific other lab out there. So they're going to wait and see when it's actually better for the, to, they can have that marketing point.”
Greg Kamradt Jul 18, 2025 ▶ 37:18 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Pratik Bhavsar Jul 14, 2025 ▶ 8:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.