The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Chollet: AGI Codebase Will Be Under 10,000 Lines on 1980s Compute
“I do believe that, you know, when you create a GI retrospectively, it will turn out that it's a code base that's less than 10,000 lines of code. And that if you had known about it back in the 19 eighties, you could have done a GI back then using the computer r…”
Chollet: Symbolic Models Will Eventually Replicate and Outperform Deep Learning
“And so everything you're doing with machine learning today, with parametric curves, we should be able to do it. With symbolic models in the future in a way that will be much, much closer to optimality. Much closer to optimality in the sense that you're going t…”
Chollet: AI in 50 Years Will Not Use Today's LLM Stack
“I personally don't think that machine learning or AI in 50 years is still going to be built on this stack.”
Chollet: All inventive AI systems rely on discrete search
“All known AI systems today that are capable of some kind of invention, some kind of creativity, they rely on discrete search.”
Chollet: 50,000x LLM scaling yielded flat progress on ARC benchmark
“Because between, like, GPT-II and GPT-IV. There's been this 50,000 X scale up of base models that has resulted in, in, in basically a flat curve. On Arc.”
Chollet: LLM Progress in Non-Verifiable Domains Will Slow or Stall
“Progress of reasoning models and base LLMs on this type of domain is, is, you know, it's going to be very slow because the stack we're using, like the LLM stack is very, very reliant on its trained data. It's basically just operationalizing the trained data. A…”
Chollet: Future AI Won't Just Be Layered Harnesses on Base LLMs
“Future AI in a few decades it's not going to be this harness on top of a reasoning model on top of a base LLM.”
Chollet: AI built from first principles will be more efficient than human brains
“It's an implementation of fundamental principles, the fundamental principles of intelligence, which, you know, I think we can identify these principles and re-implement intelligence from scratch, from first principles, in a way it will be much more efficient t…”
Chollet: Artificial General Intelligence Will Likely Arrive Around ARC-6 or ARC-7
“My timeline to AGI, you know, if you just try to extrapolate from the current rate of progress and the amount of investment that's going into, not just the LLM stack, but also like side ideas, side bets that might work out, like, you know, India, for instance,…”
Chollet: Scaling LLMs 50,000x Only Lifted ARC Accuracy to 10%
“And from at the time, back in 2019 to now, with a model like GPT 4.5, for instance, there's been a roughly 50,000 X scale up of basal alarms. And we went from zero percent accuracy on that benchmark to roughly 10%, which is not a lot.”
Chollet: Base LLMs score 0% and static reasoning scores 1-2% on ARC-2
“Well, if you take Bazel Alums, model Slack, GPT-IV-IV-V, LAMA-IV, it's simple, they get zero percent. There is simply no way to do these tasks simply via memorization. Next, if you look at static reasoning systems, so systems that use a single chain of tasks t…”
Chollet: Gradient descent requires 3 to 4 orders of magnitude more data than humans
“Gradient descent requires vast amounts of data to distill simple abstractions. Many orders of magnitude more data than what humans need. Roughly three to four orders of magnitude more.”
Chollet: SOTA test-time adaptation takes thousands in compute to solve ARC-1
“Even the latest set of the art CTA techniques they still need thousands of dollars of compute to solve arc one at human level. And that doesn't even scale to arc two.”
Chollet: AI will evolve into meta-learners synthesizing software on the fly
“AI is going to move towards systems that are more like programmers that approach a new task by writing software for it. And when faced with a new task, your programmer like MetaLearner will synthesize on the fly a program or model that is adapted to the task.”
Chollet: GPT-4 lacks fluid intelligence, but OpenAI's o3 model has it
“GPT-IV does not have fluid intelligence, for instance, but O-III does.”
Chollet: Commercial AI models will increasingly adopt test-time search architectures
“Increasingly, you're gonna see commercial models that use test-time search, where instead of just trying to generate one single COT to adapt to the task, they're actually gonna run through this, you know, search.”
Chollet: Latest base LLMs score zero percent on ARC-AGI-2
“Today the latest base alarms, they're doing something like 10% on ARK-I. But on Arc two, they are doing zero percent.”
Chollet: NDEA will solve problems previously unsolved by humans in verifiable domains
“The kind of technology we're building on, it's differential advantage that it's going to be capable of solving problems that have never been solved by humans before. That's, you know, that's a very different deal than LLMs, for instance, but effectively only i…”
Chollet: Mathematics AI Revolution Is Coming in the Next Few Years
“I think mathematics is also, it's also primed to see a revolution in the next few years for the same reasons, again, because The domain just gives you verifiable rewards.”
Chollet: Base LLMs Score Under 10% on ARC-AGI-1
“So basal alarms were scoring extremely low on V-one, like sub-ten percent, basically. And, I mean, it was true of the original, like, GPT-III actually scoring zero, but that's even true of the latest basal alarms today, you know, as of March.”
Chollet: Fine-Tuned OpenAI o3 Reached Human-Level Performance on ARC
“So in particular, in December last year, OpenAI previewed its, ah, all three model, and they used a version of it that was, ah, fine-tuned specifically on Arc, and that showed human-level performance on that benchmark versus time.”
Chollet: Ten random people with majority voting score 100% on ARC-2
“And all tasks in Arc-II were sold by at least two other people that saw it. And each task was seen on average by about seven people. And so what that tells you is that a group of 10 random people with majority voting would score 100% on Arc-II.”
Chollet: OpenAI o3 cost $10k–$20k per ARC puzzle on maximum compute
“For instance OpenAI O.S. On the highest compute settings that we tried it on for Arc, it was consuming somewhere between, like, 10,000 dollars to 20,000 dollars per task, like, for one little puzzle, which you could normally solve with a base of an API for a f…”
Chollet: OpenAI used about 75% of ARC training tasks to adapt o3
“So they told us that they were using a significant fraction, I think they said something like 75%, of the training tasks to, you know, to adapt the model in some way.”
Chollet: MindAI dropped out of ARC Prize over open-source requirements
“They ended up dropping out because they did not want to open source their solution. And of course, that meant that they were not eligible for the prize.”
Chollet: Every High-Performing ARC AI Method Uses Test-Time Adaptation
“And today, every single AI approach that performs well on Arc is using one of these techniques.”
Chollet: Average human test score on ARC-AGI-2 is about 60%
“Based on our own testing, an average person in our test sample would score about 60%.”