The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 2,445 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 100 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Drew Houston Oct 18, 2024 ▶ 52:59 Building the Silicon Brain - Drew Houston of Dropbox
Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Prediction Not checkable as stated
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Nathan Lambert Oct 13, 2024 ▶ 12:15 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Prediction Not checkable as stated
Goyal: Software engineers will drive AI engineering, but ML tools are unusable for them
“The real gap is that software engineers who have a particular way of thinking, a particular set of biases, a particular type of workflow that they run, are going to be the ones who are doing AI engineering, and that the tools that were built for ML are fantast…”
Ankur Goyal Oct 11, 2024 ▶ 34:05 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Simple tool-calling prompts cover 80% to 90% of AI use cases
“For probably 80 or 90% of the use cases that we see with people doing this, like very, very simple, I create a prompt, it calls some tools. I can like very ergonomically write the tools, plug into popular services, et cetera, and then just call them kind of li…”
Ankur Goyal Oct 11, 2024 ▶ 50:55 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Fewer Braintrust customers run fine-tuned models in production than six months ago
“I will say in my own experience with customers as of the recording date today, which is September or something, yeah, very few of our customers are currently fine-tuning models. And I think a very, very small fraction of them are running fine-tuned models in p…”
Ankur Goyal Oct 11, 2024 ▶ 1:21:53 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: OpenAI dominates production while Anthropic Sonnet leads side projects
“We still see an overwhelming majority of customers using OpenAI, but almost everyone is using Anthropic for the, and Sonnet specifically for their side projects, whether it's You know, via cursor or prototypes or whatever.”
Ankur Goyal Oct 11, 2024 ▶ 1:26:58 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Public clouds fail to match direct OpenAI endpoint experience and capacity
“It hasn't been a smooth journey for people to get the capacity on public clouds that they're able to get through, you know, OpenAI directly. I mean, I think a lot of this is changing, catching up, et cetera. But it hasn't been perfectly smooth. And I think the…”
Ankur Goyal Oct 11, 2024 ▶ 1:28:43 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Nearly all Braintrust clients shifted to simple code with LLM calls
“Almost everyone that we work with has gone into this model that, that I, that actually exactly what you said, which is sprinkle intelligence everywhere and make it easy to write dumb code”
Ankur Goyal Oct 11, 2024 ▶ 1:42:27 Production AI Engineering starts with Evals
Prediction Didn’t hold up
Godement: Developers will rely on continuous, automated fine-tuning within years
“The vision we have is, fast forward a couple of years, I think, like, most developers will essentially, like, have an automated, continuous, fine-tuned model. The more, like, you use the model, the more data you pass to the mobile provider, like, the model is …”
Olivier Godement Oct 4, 2024 ▶ 26:46 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Supported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Alistair Pullen Oct 4, 2024 ▶ 1:13:04 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Not checkable as stated
Altman: OpenAI reached Level 2 AGI with o1
“I think we clearly got to level two, or we clearly got to level two with O-one.”
Sam Altman Oct 4, 2024 ▶ 1:24:50 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Not checkable as stated
Altman: o1 is OpenAI's most aligned model ever by a lot
“And O-one is obviously our most capable model ever, but it's also our most aligned model ever by a lot.”
Sam Altman Oct 4, 2024 ▶ 1:33:35 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Not checkable as stated
Weil: OpenAI customer support team is 20% expected size thanks to AI
“There are things that get closer to that, I mean, there, like, customer service, we have bots internally that do what's fun about answering external questions and fielding internal people's questions on Slack and so on, and our customer success, our customer s…”
Kevin Weil Oct 4, 2024 ▶ 1:57:42 Building AGI in Real Time (OpenAI Dev Day 2024)
Prediction Not checkable as stated
Altman: Infinite AI context windows will happen within a decade
“That obviously takes some research breakthroughs, but I assume that infinite context will happen at some point. At some point, it's like, less than a decade.”
Sam Altman Oct 4, 2024 ▶ 2:05:18 Building AGI in Real Time (OpenAI Dev Day 2024)
Prediction Held up
Altman: 10-million-token fast context windows are coming within months
“Even getting to the, like, Ten million tokens of very fast and accurate context, which I expect to measure in, like, months, something like that.”
Sam Altman Oct 4, 2024 ▶ 2:05:30 Building AGI in Real Time (OpenAI Dev Day 2024)
Prediction Not checkable as stated
Altman: AI will dynamically render custom real-time interfaces for any request
“At some point in not that many years in the future, you'll walk up to a piece of glass, you will say whatever you want they will have, like, there will be incredible reasoning models, agents connected to everything, there will be a video model streaming back t…”
Sam Altman Oct 4, 2024 ▶ 2:08:17 Building AGI in Real Time (OpenAI Dev Day 2024)
Assertion Not checkable as stated
Shunyu Yao says Ilya Sutskever claimed GPT-1 had solved language
“Back in OpenAI, they did this GPT-ONE together, and Ilya just said, Karthik, you should stay, because we just solved the language.”
Shunyu Yao Sep 27, 2024 ▶ 2:12 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Prediction Not checkable as stated
Karpathy predicts LLMs will act as compilers generating bare-metal CUDA code
“If LLINs are about to become much better at coding over time, then I think you can expect that the LLIN could actually do this for any custom application over time. And so the LLINs could act as a kind of compiler What you're interested in, they're gonna do al…”
Andrej Karpathy Sep 21, 2024 ▶ 21:54 llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
Assertion Supported
Schulhoff: LLMs Rely More on Prompt Structure Than Exemplar Labels
“There are a number of papers which have found that the label of the exemplar doesn't really matter, and the model reads the exemplars and cares more about structure than label.”
Sander Schulhoff Sep 20, 2024 ▶ 26:41 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Assertion Not checkable as stated
Schulhoff: DSPy Beat 20 Hours of Manual Prompt Engineering in 10 Minutes
“And then I spent 20 hours prompt engineering for a task, and Dyspy beat me in 10 minutes, and that's when I changed my mind.”
Sander Schulhoff Sep 20, 2024 ▶ 45:18 The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Assertion Supported
Jamil: Writing in the Margins solves the lost-in-the-middle problem
“It improves the ability of any language model to extract relevant information, so solving the lost in the middle problem”
Umar Jamil Sep 19, 2024 ▶ 22:54 [Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Assertion Not checkable as stated
Howard: 80% of top unique creators have unconventional or non-mainstream backgrounds
“Like, 80% of the time, I find out the person has a really unusual background. So, like, often they'll have, like, either they, like, came from poverty and, like, didn't get an opportunity to go to good school, or they, like, you know, had dyslexia and, you kno…”
Jeremy Howard Aug 17, 2024 ▶ 15:43 Answer.ai & AI Magic with Jeremy Howard
Assertion Not checkable as stated
Howard: Decoder models must be far larger to match DeBERTa
“Now, the interesting thing is, you see, unlike Kaggle competitions, that decoder models still Are at least competitive with things like DiBerta VIII. But they have to be way bigger to be competitive with things like DiBerta VIII. And the only reason they are c…”
Jeremy Howard Aug 17, 2024 ▶ 35:33 Answer.ai & AI Magic with Jeremy Howard
Prediction Not checkable as stated
Howard: Reka's model is probably superior to GPT and Claude for certain tasks
“There's a whole model that's been trained in a different way. So there's probably a whole lot of tasks it's probably better at than you know, GPT and Gemini and Claude.”
Jeremy Howard Aug 17, 2024 ▶ 36:42 Answer.ai & AI Magic with Jeremy Howard
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas Scialom Jul 23, 2024 ▶ 37:43 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Prediction Not checkable as stated
Scialom: Agentic systems will yield order-of-magnitude scaling gains over pre-training
“I expect some incremental and significant progress on pre-training and post-training, but I'm really hopeful that we can gain some order of magnitude of scaling by interconnecting well models into agents as a more complex system that can do planning, that can …”
Thomas Scialom Jul 23, 2024 ▶ 47:14 Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Assertion Supported
Tay: Zero-shot benchmark scores at 1B model scale are random chance
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance …”
Yi Tay Jul 5, 2024 ▶ 1:46:43 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Assertion Supported
Albrecht: Benchmark performance differences vanish once ambiguous questions are cleaned
“The main takeaway from any of the, like, actual performance is like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it's important to look at like, you know, like when you're talking about performance on NL…”
Josh Albrecht Jun 25, 2024 ▶ 1:01:43 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Supported
Albrecht: AgentBench paper's appendix examples are actually incorrect solutions
“Like we were looking at the agent bench paper, I think just last week for our paper club. And one of the things that we noticed is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not…”
Josh Albrecht Jun 25, 2024 ▶ 1:06:05 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Assertion Not checkable as stated
Frankle: No Databricks enterprise customer asks for abstract reasoning AI
“I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems.”
Jonathan Frankle Jun 25, 2024 ▶ 1:13:34 State of the Art: Training 70B LLMs on 10,000 H100 clusters
Prediction Not checkable as stated
Brady: Prompt Engineering Won't Be a Durable Differentiating Skill
“I don't think that prompt engineering is going to be a kind of durable differential skill that people will hold. I do think that the way that you set up the ML problem to kind of ask the right questions, if you see what I mean, rather than the specific phrasin…”
James Brady Jun 21, 2024 ▶ 31:31 How To Hire AI Engineers (ft. James Brady and Adam Wiggins of Elicit)
Assertion Not checkable as stated
Conover: AI model developers are absolutely overfitting to public evaluation benchmarks
“And I think the work around over, you know, overfitting on the test, I think is like that. 100% is happening.”
Mike Conover Jun 11, 2024 ▶ 58:21 How AI is Eating Finance - with Mike Conover of Brightwave
Prediction Not checkable as stated
Conover: Economic incentives to pre-train commodity foundation models from scratch are diminishing
“The incentives, the economic incentives for companies to train their own foundation models, I think, are diminishing. So the, like, window in which you are the dominant pre-train, and let's say that you spend five to forty million dollars, you know, for like a…”
Mike Conover Jun 11, 2024 ▶ 58:37 How AI is Eating Finance - with Mike Conover of Brightwave
Assertion Partly supported
Bach: LLMs Demonstrate Theory of Mind from Conversational Context
“When you ask the LLM to make inferences about your mental state based on the conversation that you have, it's able to demonstrate that it has a theory of mind.”
Joscha Bach Apr 27, 2024 ▶ 1:03:48 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Prediction Held up
Bach: Smaller, more powerful models will ensure unconstrained AI remains accessible
“Yes, but there will also be better jailbroken models or models that have never been jailed before, because we find out how to make smaller models that are more powerful.”
Joscha Bach Apr 27, 2024 ▶ 1:28:21 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Contradicted
Bach: AI has casually passed the Turing test in recent years
“At some point in the last few years, we casually skipped the Turing test, right? We broke through it.”
Joscha Bach Apr 27, 2024 ▶ 1:45:41 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Prediction Not checkable as stated
Liu: XGBoost rankers will beat LLMs for large-scale tool selection
“Yeah, my money is on the rankers because you can do those so easily, right? You could just say, well, given the embeddings of my search query and the embeddings of the description, I can just train XGBoost and just make sure that I have very high, like, MRR, w…”
Jason Liu Apr 24, 2024 ▶ 19:54 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Assertion Supported
Murphy: Five-minute voice calls cost 6.5 cents on Deepgram versus ElevenLabs.
“And then on the text-to-speech side, and doing something like this with an 11 labs would be about maybe a dollar 20. And just to give you an idea of comparison. So you can do a five minute call here for about six and a half cents.”
Damien Murphy Apr 6, 2024 ▶ 16:52 Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Prediction Not checkable as stated
Luan: AI will converge into a universal byte model across all modalities
“Multimodal models are becoming more of a thing, we're behavioral cloning the visual world, but really what we're just going to have is this like universal byte model, right? Where like tokens of data that have high signal come in, and then all of those pattern…”
David Luan Mar 27, 2024 ▶ 6:15 Why Google failed to make GPT-3 -- with David Luan of Adept
Prediction Not checkable as stated
Luan: Future AI value will shift from base models to agents
“In a world where foundation models are looking more and more commodity. And if, and I think a huge amount of gain is going to happen from how do you use foundation models as like the, like well learned behavioral cloner to go solve agents.”
David Luan Mar 27, 2024 ▶ 23:50 Why Google failed to make GPT-3 -- with David Luan of Adept
Prediction Held up
Multimodal models will completely supplant text-only large language models
“I actually think like it's really clear today. Multimodal models are the default foundation model, right? It's just going to supplant LLMs. Like why did you just train a giant multimodal model?”
David Luan Mar 27, 2024 ▶ 28:05 Why Google failed to make GPT-3 -- with David Luan of Adept
Prediction Not checkable as stated
Pure-play foundation model companies will be commoditized by Llama and big tech
“I think pure play foundation model companies are just gonna be pinched by how good The next couple of llamas are going to be, and the next, like, what next good open source thing, and then seeing the really big players put ridiculous amounts of compute behind …”
David Luan Mar 27, 2024 ▶ 42:31 Why Google failed to make GPT-3 -- with David Luan of Adept
Prediction Not checkable as stated
Luan: Training foundation models for robotics via behavioral cloning will work
“One, I'm so excited for someone to train a foundation model of robots. Like, it's just, I think it's just gonna work. Like, I will die on this hill. I mean, like, again, this whole time, like, we've been on this podcast, just continually saying, you know, like…”
David Luan Mar 27, 2024 ▶ 44:47 Why Google failed to make GPT-3 -- with David Luan of Adept
Prediction Not checkable as stated
Chintala: LLM inference market will become a low-margin laundromat business
“My view of the LLM inference market in general is that it's like the laundromat model. Like you, the margins are going to drive down towards the bare minimum, like It's gonna be all kinds of arbitrage between how much you can get the hardware for, and then how…”
Soumith Chintala Mar 6, 2024 ▶ 28:12 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Assertion Not checkable as stated
Chintala: Aggregate open source AI model usage rivals GPT
“Maybe open source models are being as used as GPT is at this point in, like, all kinds of, in a very fragmented way. Like, in aggregate, all the open source models together are probably being used as much as GPT is. Maybe, you know, close to that.”
Soumith Chintala Mar 6, 2024 ▶ 1:12:31 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Prediction Not checkable as stated
Chintala: Centralized feedback could trigger open source runaway over OpenAI
“If that central sinkhole is there, who's gonna go coordinate all of this integration across all of these, like, open source frontends? But I think if we do that, if that actually happens, I think that probably has a real chance of the open source models having…”
Soumith Chintala Mar 6, 2024 ▶ 1:14:45 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Prediction Not checkable as stated
Chintala: Open source cannot beat Google's feedback distribution advantage
“Probably doesn't have a chance against Google, because, you know, Google has Android, and Chrome, and Gmail, and Google Docs, and everything, you know. So people just use that a lot.”
Soumith Chintala Mar 6, 2024 ▶ 1:15:10 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.