Everything Jeremy Howard said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Howard: Closed US AI ecosystems will cause China to move faster
“Whereas, oddly, the US companies that used to lead the way have all drawn up the drawbridges, and that's gonna cause China to keep moving faster, because when you're in that more, both collaborative and competitive environment, you just Go way ahead, as we've …”
Howard: OpenAI compute spending grows exponentially while model utility scales logarithmically
“They kept on kind of exponentially increasing the amount they were spending on their models, whilst the Return, you know, the kind of utility of those models was only increasing logarithmically, and you kind of very quickly hit this point where it's like, oh, …”
Howard: Test-time compute scaling will hit diminishing returns within two years
“It's a thing where you get most of the juice out of it in the first year or two, so we're still in that. Phase at the moment, and we'll start to hit the curve off point pretty soon. Just like we did for training.”
Howard: No more evidence for near-term ASI today than 15 years ago
“I don't think we have any more evidence that ASI might be close now than we did 15 years ago, 15 years before that.”
Howard: AI growth will continue tailing off as low-hanging fruit disappears
“The amount of people and money being put into AI, it's going to keep going up, but again, the low hanging fruits being done now. So yeah, I think we'll continue to see the tailing off of growth.”
Howard: Traditional AI labs favor paper citations over bold innovation
“It's very different to a research lab, which is like, hey, let's try and get lots of citations on a paper in a field that's so highly recognizable to its peers that they recognize it as being something they want to appear at their next conference, you know, wh…”
Howard: Autonomous AI coding agent Devin produces low-quality, useless software
“Devon's an agent you know, a bunch of tools, tool calls, and tied together with an OpenAI model, if I understand correctly. Yeah, and it actually ended up definitely supporting our thesis, which is, there are so many places we wished we could have got involved…”
Jeremy Howard sees no reason to believe Artificial Superintelligence will occur
“I have no particular reason to believe that will happen, and for every point until we get there, if we ever get there, by definition, there's going to be humans interacting with AI, and so that's what I care about, is how do we do that In an optimal way.”
Howard: Alec Radford Built OpenAI's GPT After Reading ULMFiT
“I organized a chat for both of us with Kate Metz in the New York Times, and Kate Metz answered, sorry, and Alec answered this question for Kate, and Kate just like, so how did, you know, GPT come about? And he said, well, I was pretty sure that pre-training on…”
Howard: Meta 'blew it' on Code Llama due to catastrophic forgetting
“So Code Llama was a, I think it was like a five hundred billion token fine-tuning of Llama II using code. And also prose about code that Meta did. And honestly, they kind of blew it. Because Code Llama is good at coding, but it's bad at everything else.”
Howard: TensorFlow 2 was a failure that Google avoided internally
“Then in the end, you know, Google didn't follow through, which is fair enough, like, asking everybody to, you know, learn a new programming language is going to be tough, but, like, it was very obvious, very, very obvious at that time that TensorFlow II was go…”
Howard: JAX was a grassroots Google reaction against TensorFlow 2
“But I mean, in the meantime, I will say, you know, Google now does have a backup plan. You know, they have JAX, which was never a strategy. It was just a bunch of people who also recognized TensorFlow two as shit, and they just decided to build something else.”
Howard: RAG is an inefficient hack compared to fine-tuning
“RAG is like such a inefficient hack, really, isn't it? It's like, You know, segment up my data in some somewhat arbitrary way, embed it, ask questions about that, you know, hope that my embedding, you know, model embeds questions in the same embedding space as…”
Howard: Answer.AI runs fully in-house stack without AWS or Google Cloud
“This group of, which has averaged about 10 to 12 people, currently nine, I think, have built a pretty Transformational and complex piece of software, which we can do a quick demo of later if you're interested. Using a complete web application development platf…”
Howard: Correcting LLM errors in chat history degrades subsequent model answers
“The autoregressive nature of language models means that if they make a mistake, and you correct it, and then say, no, that was a mistake, please do it this way instead. The more often you do that, the worse the dialogue answers get. Because it's in the trainin…”
Howard: Longer dialogues improve AI outputs when humans edit intermediate results
“The nice thing is that with the dialogue engineering we discussed, the longer your dialogue is, the better the AI gets, which is the opposite to what we're used to, right? Because you can edit the outputs that aren't great.”
Howard: There was no technological breakthrough 'DeepSeek moment'
“For me, there was no technology DeepSeek moment.”
Howard: OpenAI is shutting down GPT-4.5
“I think they're shutting down that product or they've shut down that product, if I understand correctly.”
Jeremy Howard: Academic approaches to predictive modeling are far less successful than practical experience
“Much to my surprise, the academic approach is, was way less successful.”
Howard: Answer.ai aims to launch 5,000 successful products with 14 people
“If we're going to have 12 to 14 people create five to 10,000 extremely commercially successful products, we're going to have to be extremely efficient, you know, at every level.”
Howard: Training AI models from random weights is almost never justified
“If you're training for random weights, you better have a really good reason, you know, because it seems so unlikely to me that nobody has ever trained on data that has any similarity whatsoever to the general class of data you're working with, and that's the o…”
Howard: Pre-training data mixes should be continuous per-batch functions, not discrete phases
“So the point at which they're doing proper continued pre-training is the point at which that becomes a continuum rather than a phase. So the only difference with what I was describing last time is to say, like, oh, they should, you know, There's a function or …”
Howard: Non-profit boards cannot control commercial entities with equity-compensated staff
“This didn't make sense to have like a so-called non-profit where then there are people working at a commercial company that's owned by or controlled nominally by the non-profit where the people in the company are being given the equivalent of stock options. Li…”
Howard: Corporations are sociopathic by design due to fiduciary duty
“Companies are sociopathic, like, by design. And so the alignment problem, as it relates to companies, has not been solved. Like, companies become huge, they devour their founders, they devour their communities, and they do things where even the CEOs, you know,…”