Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q kind of core question that everyone's asking right now is, does more compute equal an increased level of performance, or have we reached a point where it is misaligned, and more compute will not create that significant spike in performance? Kevin Scott at Microsoft says, absolutely, we have a lot more room to run. Why are you skeptical? And Have we gotten to a stage of diminishing returns on compute?
A So if we look at what's happened historically, the way in which, uh, compute has improved model performance is with companies building bigger models, right? So I, in my, in my view, at least the biggest thing that changed between GPT, 3.5 and GPT four was the size of the model. And it was, you know, also trained with, uh, more data, presumably, although they haven't made the details of that public and more compute and so forth. So I think that's running out. I think we're not going to be, we're not going to have too many more cycles, possibly zero more cycles of a model that's, you know, almost an order of magnitude bigger in terms of the number of parameters than what came before, and thereby more powerful. And I think a reason for that is data becoming a bottleneck. These models are already trained on essentially all of the data that companies can get their hands on. While data is becoming a bottleneck, I think more compute still helps, but maybe not as much as it used to. And the reason for that is that, uh, perhaps ironically, more compute allows one to build smaller models with the same capability level. And that's actually the trend we've been seeing over the last year or so. As you know, you know, the models today have gotten somewhat smaller and cheaper than when GPT-IV initially came out, but with the same capability level. So I think that's probably going to continue.…
AI assessment note: “more compute still helps, but maybe not as much as it used to.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Do you think people dramatically overestimate the fear of job replacement? We always see job replacement with any new technology, and then it tends to create a lot more jobs than previously were. Do you think that is the case here, or do you think job replacement fears are justified?
A I think for now they are very much overblown. My favorite example of the thing you said of technology creating jobs is bank tellers. When ATMs became a thing, you know, It would have been reasonable to assume that bank tellers were just going to go away, but in fact, the number of tellers increased, and the reason for that is that it became much cheaper for banks to open regional branches, and once they did open those regional branches, they did need humans for some of the things that you couldn't do with an ATM. You know, the more abstract way of saying that is, as economists would put it, jobs are bundles of tasks, And AI automates tasks, not jobs. So if there are, you know, 20 different tasks that comprise a job, uh, the odds that AI is gonna be able to automate all 20 of them are pretty low. And so there are some occupations, certainly, that have already been affected a lot by AI, like translation or stock photography. But, you know, for, for most jobs out there, I don't think we're anywhere close to that.
AI assessment note: “I think for now they are very much overblown.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Why does it become a barrier and why does it not?
A There's this interesting concept called Jevons paradox. And this was first in the context of, uh, uh, coal in England in the 18th century. I think when coal mining got cheaper, there was more demand for coal. And so the amount invested into coal mining actually increased. And I predict that we're going to see the same thing with models when models get cheaper. They're put into a lot more things, and so, uh, the total amount that companies are spending on inference is actually going to increase. On a task like, uh, in an application like a chatbot, let's say, you know, it's text in, text out, no big deal. I think costs are going to come down. Even if someone is chatting with, uh, a chatbot all day, it's probably not going to get too expensive. On the other hand, if you want to scan all of someone's emails, for instance, right? If a model gets cheaper, You know, you're just going to have it running always on in the background, and then from emails you're going to get to all their documents, right, and some of those attachments might be many megabytes long. Uh, and so there, even with Moore's Law, I think cost is going to be significant in the medium term. And then you get to applications like writing code, where what we're seeing is that it's actually very beneficial to let the model do the same task tens of times, thousands of times, sometimes literally millions of times, and pi…
AI assessment note: “in an application like a chatbot... costs are going to come down. On the other hand”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q So we have smaller models, and they're effective, as we said, they're because of cost, and they're popular because of cost. What does that do to the requirements in terms of compute?
A So there is trading compute, which is when the developer is building the model, and then there is inference compute, when the model is being deployed and the user is using it to do something. And it might seem like Like, uh, really the training cost is the one we should worry about since, you know, it's trained on all of the text on the internet or whatever, but it turns out that over the lifetime of a model, when you have billions of people using it, the inference cost actually adds up, and for many of the popular models, that's the cost that dominates. Let's talk about each of those two costs. With respect to training cost, if you want to build a smaller model at the same level of capability or without compromising capability too much, you have to actually train it for longer. So that increases training costs, but that's maybe ok, because you have a smaller model, you can push it to the consumer device, or, you know, even if it's running on the cloud, your server costs are lower, so your training cost increases, your inference cost decreases, but because it's the inference cost that dominates, the total cost is probably going to come down.
AI assessment note: “your training cost increases, your inference cost decreases”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Do you think people dramatically overestimate the fear of job replacement? We always see job replacement with any new technology, and then it tends to create a lot more jobs than previously were. Do you think that is the case here, or do you think job replacement fears are justified?
A I think for now they are very much overblown. My favorite example of the thing you said of technology creating jobs is bank tellers. When ATMs became a thing, you know, It would have been reasonable to assume that bank tellers were just going to go away, but in fact, the number of tellers increased, and the reason for that is that it became much cheaper for banks to open regional branches, and once they did open those regional branches, they did need humans for some of the things that you couldn't do with an ATM. You know, the more abstract way of saying that is, as economists would put it, jobs are bundles of tasks, And AI automates tasks, not jobs. So if there are, you know, 20 different tasks that comprise a job, uh, the odds that AI is gonna be able to automate all 20 of them are pretty low. And so there are some occupations, certainly, that have already been affected a lot by AI, like translation or stock photography. But, you know, for, for most jobs out there, I don't think we're anywhere close to that.
AI assessment note: “I think for now they are very much overblown.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So we have smaller models, and they're effective, as we said, they're because of cost, and they're popular because of cost. What does that do to the requirements in terms of compute?
A So there is trading compute, which is when the developer is building the model, and then there is inference compute, when the model is being deployed and the user is using it to do something. And it might seem like Like, uh, really the training cost is the one we should worry about since, you know, it's trained on all of the text on the internet or whatever, but it turns out that over the lifetime of a model, when you have billions of people using it, the inference cost actually adds up, and for many of the popular models, that's the cost that dominates. Let's talk about each of those two costs. With respect to training cost, if you want to build a smaller model at the same level of capability or without compromising capability too much, you have to actually train it for longer. So that increases training costs, but that's maybe ok, because you have a smaller model, you can push it to the consumer device, or, you know, even if it's running on the cloud, your server costs are lower, so your training cost increases, your inference cost decreases, but because it's the inference cost that dominates, the total cost is probably going to come down.
AI assessment note: “your training cost increases, your inference cost decreases”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Speaking of that commoditization, the thing that I'm interested by there is kind of the benchmarking or the determination that they are suddenly commoditized or kind of equal performance. You said before LLM evaluation is a minefield. Help me understand why is LLM evaluation a minefield?
A A big part of it is this issue of vibes, right? So you evaluate LMs on these benchmarks, but then, you know, it seems to perform really well on the benchmarks, but then the vibes are off. In other words, you start using it, and somehow it doesn't feel adequate. It makes a lot of mistakes in ways that are not captured in the benchmark, and the reason for that is simply that when there is so much pressure to do well on these benchmarks, developers are intentionally or unintentionally optimizing these models In ways that look good on the benchmarks, but, uh, don't look good in real world evaluation. So when GPT-IV came out and OpenAI claimed that it passed the bar exam and the medical licensing exam, uh, people were very excited slash, uh, scared about what this means for doctors and lawyers, and the answer turned out to be approximately nothing, right? Because it's not like a, a lawyer's job is to answer bar exam questions all day. Uh, these, Benchmarks that models are being tested on don't really capture what we would use them for in the real world. So that's one reason why LLM evaluation is a minefield. And there's also just a, uh, very, uh, simple factor of contamination. Maybe the model has already trained on the answers that it's being evaluated on in the benchmark. And so if you ask it new questions, it's gonna struggle, uh, and there are various other pitfalls. So I think,…
AI assessment note: “A big part of it is this issue of vibes, right?”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I feel like I worry more about this than you on the content misinformation side, and so I'm intrigued on the concerns that you have. What would you say is a more pressing concern for you?
A So when we were talking about deepfakes, I'm much less worried about misinformation deepfakes and more worried about, uh, deepfake nudes that I was talking about, right? So those are things that can destroy a person's life. It's been shocking to me how little attention this got from the press and from policymakers until it happened to Taylor Swift a few months ago, uh, and then it got a lot of attention. So there were deepfake nudes of Taylor Swift posted on Twitter slash X, uh, and after that, You know, policymakers started paying attention, but it has been happening for many years now, even before the latest wave of generative AI tools. So that's the type of misuse, you know, that, that is very clear. And then there are other kinds of misuses that are not necessarily dangerous in the same way, but impose a lot of costs on society. So when students are using AI to do their homework, for instance, now, you know, high school teachers and college teachers everywhere have to Uh, revamp how they're teaching in order to account for the fact that students are doing this and there's no way really to catch AI-generated, uh, text or, or homework answers. And so these are costs upon society. I'm not saying that the availability of AI makes education worse. I don't think that's necessarily the case. But, uh, you know, it forces a lot of costs upon the education system, and ideally, AI com…
AI assessment note: “I'm much less worried about misinformation deepfakes and more worried about, uh, deepfake nudes”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q kind of core question that everyone's asking right now is, does more compute equal an increased level of performance, or have we reached a point where it is misaligned, and more compute will not create that significant spike in performance? Kevin Scott at Microsoft says, absolutely, we have a lot more room to run. Why are you skeptical? And Have we gotten to a stage of diminishing returns on compute?
A So if we look at what's happened historically, the way in which, uh, compute has improved model performance is with companies building bigger models, right? So I, in my, in my view, at least the biggest thing that changed between GPT, 3.5 and GPT four was the size of the model. And it was, you know, also trained with, uh, more data, presumably, although they haven't made the details of that public and more compute and so forth. So I think that's running out. I think we're not going to be, we're not going to have too many more cycles, possibly zero more cycles of a model that's, you know, almost an order of magnitude bigger in terms of the number of parameters than what came before, and thereby more powerful. And I think a reason for that is data becoming a bottleneck. These models are already trained on essentially all of the data that companies can get their hands on. While data is becoming a bottleneck, I think more compute still helps, but maybe not as much as it used to. And the reason for that is that, uh, perhaps ironically, more compute allows one to build smaller models with the same capability level. And that's actually the trend we've been seeing over the last year or so. As you know, you know, the models today have gotten somewhat smaller and cheaper than when GPT-IV initially came out, but with the same capability level. So I think that's probably going to continue.…
AI assessment note: “I think that's running out... a reason for that is data becoming a bottleneck.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Speaking of that commoditization, the thing that I'm interested by there is kind of the benchmarking or the determination that they are suddenly commoditized or kind of equal performance. You said before LLM evaluation is a minefield. Help me understand why is LLM evaluation a minefield?
A A big part of it is this issue of vibes, right? So you evaluate LMs on these benchmarks, but then, you know, it seems to perform really well on the benchmarks, but then the vibes are off. In other words, you start using it, and somehow it doesn't feel adequate. It makes a lot of mistakes in ways that are not captured in the benchmark, and the reason for that is simply that when there is so much pressure to do well on these benchmarks, developers are intentionally or unintentionally optimizing these models In ways that look good on the benchmarks, but, uh, don't look good in real world evaluation. So when GPT-IV came out and OpenAI claimed that it passed the bar exam and the medical licensing exam, uh, people were very excited slash, uh, scared about what this means for doctors and lawyers, and the answer turned out to be approximately nothing, right? Because it's not like a, a lawyer's job is to answer bar exam questions all day. Uh, these, Benchmarks that models are being tested on don't really capture what we would use them for in the real world. So that's one reason why LLM evaluation is a minefield. And there's also just a, uh, very, uh, simple factor of contamination. Maybe the model has already trained on the answers that it's being evaluated on in the benchmark. And so if you ask it new questions, it's gonna struggle, uh, and there are various other pitfalls. So I think,…
AI assessment note: “developers are intentionally or unintentionally optimizing these models In ways that look good on the benchmarks”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Help me just understand again. I'm sorry, I, I, the show's very successful, Arvin, because I think I asked the questions that everyone asked, but they're too afraid to actually admit they don't know the answers to. Um, Why are we seeing this trend towards smaller models? And why do we think that is the most likely outcome in the model landscape to have a world of many smaller models?
A My view is that in a lot of cases, the adoption of these models is not bottlenecked by capability. If these models were actually deployed today to do all the tasks that they're capable of, it would truly be a striking economic transformation. The bottlenecks are things other than capability, and one of the big ones is cost. And cost, of course, is roughly proportional to the size of the model, and that's putting a lot of downward pressure on model size. Once you get a model small enough that you can run it on device, Excuse me. That, of course, opens up a lot of new possibilities, uh, both in terms of, uh, privacy. You know, people are much more comfortable with on-device models, especially if it's something that's going to be listening to their phone conversations or looking at their desktop screenshots, which are exactly the kinds of AI assistants that companies are, uh, building and pushing. Uh, and just, you know, from the perspective of cost, you don't have to dedicate servers to run that model. So I think those are a lot of the reasons why, Uh, companies are furiously working on making models smaller without a big hit in capability.
AI assessment note: “The bottlenecks are things other than capability, and one of the big ones is cost.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What did you mean when you said to me that AI companies should pivot from creating gods to building products?
A In the past, uh, you know, they didn't have this balance. They, um, were so enamored by this prospect of creating AGI that they didn't think there was a need to build products at all. And, you know, the craziest example for me is when OpenAI put out ChatGPT, there was no mobile app for six months, uh, and the Android app took even longer than that. And there was this, you know, there was this assumption that ChatGPT was just going to be this kind of, uh, uh, really demo to show off the capabilities of the models. OpenAI was, you know, in the business of building these models and Uh, third-party developers would take the API and put it into products, but really AGI was coming so quickly that, that, you know, even the notion of productization seemed obsolete. This was, you know, I, I'm not trying to put words in anyone's mouth, but this was kind of a coherent, but in my view, incorrect philosophy that I think a lot of AI developers had. Um, and I think that has, uh, changed quite a bit now, and I think that's a good thing. Uh, so if they had to pick one, I think they should pick building products, but it certainly doesn't make sense for a company to be just an AGI company and not try to build products, not try to build something that people want, and just assuming that AI is going to be so general that it's just going to, you know, do everything that people want and, and, and tha…
AI assessment note: “They, um, were so enamored by this prospect of creating AGI”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q If you, uh, Suggesting, as you said at the beginning about kind of your work on policy, you have US regulators and European regulators. What would you put forward as the most proactive and effective policy for US and European regulation around AI and models?
A So in a sense, AI regulation is a misnomer. Let me give you an example from just this morning. The FTC, uh, has been worried about, uh, the Federal Trade Commission in the US, uh, Uh, you know, which is, um, an antitrust and consumer protection authority has been worried about, uh, people writing fake reviews for their products, and this has, of course, been a problem for many years. It's become a lot easier to do that with AI. So now, someone who thinks about this in terms of AI regulation might say, oh, you know, regulators have to ensure that AI companies don't allow their products to be used for generating fake reviews. And I think this is a losing proposition. Like, how would an AI model know whether something is a fake review or a real, real review, right? It just depends on who's, uh, writing their review. But instead, you know, that's not the approach that the FTC took. They recognized correctly that it's a problem whether AI is generating the fake review or people are. So what they actually banned is fake reviews, right? And so what is often thought of as AI regulation is better understood as regulating certain harmful activities, whether or not AI is Used as a tool for doing those harmful activities. So I think, you know, 80% of what gets called AI regulation is better seen this way.
AI assessment note: “AI regulation is better understood as regulating certain harmful activities”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I feel like I worry more about this than you on the content misinformation side, and so I'm intrigued on the concerns that you have. What would you say is a more pressing concern for you?
A So when we were talking about deepfakes, I'm much less worried about misinformation deepfakes and more worried about, uh, deepfake nudes that I was talking about, right? So those are things that can destroy a person's life. It's been shocking to me how little attention this got from the press and from policymakers until it happened to Taylor Swift a few months ago, uh, and then it got a lot of attention. So there were deepfake nudes of Taylor Swift posted on Twitter slash X, uh, and after that, You know, policymakers started paying attention, but it has been happening for many years now, even before the latest wave of generative AI tools. So that's the type of misuse, you know, that, that is very clear. And then there are other kinds of misuses that are not necessarily dangerous in the same way, but impose a lot of costs on society. So when students are using AI to do their homework, for instance, now, you know, high school teachers and college teachers everywhere have to Uh, revamp how they're teaching in order to account for the fact that students are doing this and there's no way really to catch AI-generated, uh, text or, or homework answers. And so these are costs upon society. I'm not saying that the availability of AI makes education worse. I don't think that's necessarily the case. But, uh, you know, it forces a lot of costs upon the education system, and ideally, AI com…
AI assessment note: “more worried about, uh, deepfake nudes that I was talking about”
Answered raw tape
D 5 · C 5 · P 4 · Cm 5 4.75
Q you know, Zuck has committed fifty billion dollars over the next three years. When you look at how much OpenAI has raised over the last three years and they carry on that run rate, it's something crazy like that. It'd still be thirty-eight billion dollars short of a Zuck spend over a three-year period. Can you create AGEI-like products or God-like products unless you are Google, Amazon, Apple, or Facebook?
A You know, we've been in this kind of historically Um, interesting period where a lot of progress has come from building bigger and bigger models that need not continue in the future. It might, or what might happen is that the models themselves get commoditized, and a lot of the interesting development happens in a layer above the models. We're starting to see a lot of that happen now with AI agents, and if that's the case, great ideas could come from anywhere, right? It could come from a two-person startup. It could come from an academic lab. Uh, and my hope is that we will transition to that kind of, ah, mode of progress in AI development, ah, relatively soon.
AI assessment note: “great ideas could come from anywhere, right? It could come from a two-person startup.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What about synthetic data? What about the creation of new data that doesn't exist yet?
A So there's two ways to look at this, right? So one is the way in which synthetic data is being used today, which is not to increase the volume of training data, but it's actually to Overcome limitations in the quality of the training data that we do have. Uh, so for instance, if in a particular language there's too little data, you can try to augment that, or you can try to, um, uh, let's say have a model, uh, you know, solve a bunch of mathematical equations, throw that into the training data, uh, and so for the next training run, that's going to be Part of the pre-training, and so the model will get better at doing that. And the other way to look at synthetic data is, okay, you take one trillion tokens, you train a model on it, and then you output 10 trillion tokens, so you get to the next bigger model, and then you use that to output a hundred trillion tokens. You know, I'll, I'll bet that's just not going to happen. That's just a snake eating its own tail, and what we've learned in the last two years is that the quality of data matters a lot more than the quantity of data. So, If you're using synthetic data, uh, to, uh, try to augment the, the quantity, I think it's just coming at the expense of quality. You're, you're not learning new things from the data. You're only learning things that are already there.
AI assessment note: “So there's two ways to look at this, right?”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What do you think society's biggest misconception of AI is today?
A I think our intuitions are too powerfully shaped by sci-fi portrayals of AI, and I think that's really a big problem. Uh, you know, this idea that AI can become self-aware. When we look at the way that AI is architected today, that kind of fear has no basis in reality. Uh, maybe one day in the future, you know, people are going to build, uh, AI systems where that becomes, uh, at least somewhat possible. And we should, you know, we should have, uh, visibility, transparency, monitoring, regulation around these systems to make sure that developers don't. But that would be a choice. That's a choice that society can make, that governments and companies can make. It's not that despite our best efforts, AI is going to become conscious and have agency and do things that are harmful to humanity. That whole line of fear, I think, is, uh, completely unfounded.
AI assessment note: “this idea that AI can become self-aware”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q you know, Zuck has committed fifty billion dollars over the next three years. When you look at how much OpenAI has raised over the last three years and they carry on that run rate, it's something crazy like that. It'd still be thirty-eight billion dollars short of a Zuck spend over a three-year period. Can you create AGEI-like products or God-like products unless you are Google, Amazon, Apple, or Facebook?
A You know, we've been in this kind of historically Um, interesting period where a lot of progress has come from building bigger and bigger models that need not continue in the future. It might, or what might happen is that the models themselves get commoditized, and a lot of the interesting development happens in a layer above the models. We're starting to see a lot of that happen now with AI agents, and if that's the case, great ideas could come from anywhere, right? It could come from a two-person startup. It could come from an academic lab. Uh, and my hope is that we will transition to that kind of, ah, mode of progress in AI development, ah, relatively soon.
AI assessment note: “if that's the case, great ideas could come from anywhere”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Help me just understand again. I'm sorry, I, I, the show's very successful, Arvin, because I think I asked the questions that everyone asked, but they're too afraid to actually admit they don't know the answers to. Um, Why are we seeing this trend towards smaller models? And why do we think that is the most likely outcome in the model landscape to have a world of many smaller models?
A My view is that in a lot of cases, the adoption of these models is not bottlenecked by capability. If these models were actually deployed today to do all the tasks that they're capable of, it would truly be a striking economic transformation. The bottlenecks are things other than capability, and one of the big ones is cost. And cost, of course, is roughly proportional to the size of the model, and that's putting a lot of downward pressure on model size. Once you get a model small enough that you can run it on device, Excuse me. That, of course, opens up a lot of new possibilities, uh, both in terms of, uh, privacy. You know, people are much more comfortable with on-device models, especially if it's something that's going to be listening to their phone conversations or looking at their desktop screenshots, which are exactly the kinds of AI assistants that companies are, uh, building and pushing. Uh, and just, you know, from the perspective of cost, you don't have to dedicate servers to run that model. So I think those are a lot of the reasons why, Uh, companies are furiously working on making models smaller without a big hit in capability.
AI assessment note: “The bottlenecks are things other than capability, and one of the big ones is cost.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Why does it become a barrier and why does it not?
A There's this interesting concept called Jevons paradox. And this was first in the context of, uh, uh, coal in England in the 18th century. I think when coal mining got cheaper, there was more demand for coal. And so the amount invested into coal mining actually increased. And I predict that we're going to see the same thing with models when models get cheaper. They're put into a lot more things, and so, uh, the total amount that companies are spending on inference is actually going to increase. On a task like, uh, in an application like a chatbot, let's say, you know, it's text in, text out, no big deal. I think costs are going to come down. Even if someone is chatting with, uh, a chatbot all day, it's probably not going to get too expensive. On the other hand, if you want to scan all of someone's emails, for instance, right? If a model gets cheaper, You know, you're just going to have it running always on in the background, and then from emails you're going to get to all their documents, right, and some of those attachments might be many megabytes long. Uh, and so there, even with Moore's Law, I think cost is going to be significant in the medium term. And then you get to applications like writing code, where what we're seeing is that it's actually very beneficial to let the model do the same task tens of times, thousands of times, sometimes literally millions of times, and pi…
AI assessment note: “it doesn't matter how much cost goes down. You're gonna just proportionally increase”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What did you mean when you said to me that AI companies should pivot from creating gods to building products?
A In the past, uh, you know, they didn't have this balance. They, um, were so enamored by this prospect of creating AGI that they didn't think there was a need to build products at all. And, you know, the craziest example for me is when OpenAI put out ChatGPT, there was no mobile app for six months, uh, and the Android app took even longer than that. And there was this, you know, there was this assumption that ChatGPT was just going to be this kind of, uh, uh, really demo to show off the capabilities of the models. OpenAI was, you know, in the business of building these models and Uh, third-party developers would take the API and put it into products, but really AGI was coming so quickly that, that, you know, even the notion of productization seemed obsolete. This was, you know, I, I'm not trying to put words in anyone's mouth, but this was kind of a coherent, but in my view, incorrect philosophy that I think a lot of AI developers had. Um, and I think that has, uh, changed quite a bit now, and I think that's a good thing. Uh, so if they had to pick one, I think they should pick building products, but it certainly doesn't make sense for a company to be just an AGI company and not try to build products, not try to build something that people want, and just assuming that AI is going to be so general that it's just going to, you know, do everything that people want and, and, and tha…
AI assessment note: “they were so enamored by this prospect of creating AGI that they didn't think”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Arvin, what did you believe about Kind of the developments that we've seen in AI over the last two years that you now no longer believe?
A So I think, like a lot of people, I was fooled by how quickly after GPT-III.V, GPT-IV came out. It was just, you know, three months or so, but it had been in training for 18 months. That was only revealed later. So it gave a lot of people, including me, an inflated idea of how quickly AI has, was progressing. And what we've seen in the nearly year and a half since GPT-IV came out is that we haven't really had models That have, uh, surpassed it in a meaningful way. Uh, and this is not based on benchmarks. Again, I think benchmarks are not that useful. It's more based on vibes. When you get people using these things, what do they say? I don't think models have, you know, really qualitatively improved on, on GPT-IV. And I don't think things are moving as quickly as I did 12 months ago.
AI assessment note: “It gave a lot of people, including me, an inflated idea of how quickly AI”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Can I say another one that does worry me is actually defense. You know, we had Alex Wang from Scale On I mentioned earlier, he said that AI has the potential to be a bigger weapon than nuclear weapons. How do you think about that? And if that is the case, should we really have open models?
A I think, you know, it's, it's a good question to ask. I think it's a bit of a, a category error there. I mean, a nuclear weapon is an actual weapon. AI is not a weapon. AI is something that, you know, might enable, uh, adversaries to do certain things more effectively to, um, you know, for example, find, uh, vulnerabilities, cybersecurity vulnerabilities in critical infrastructure, right? So that's one way in which, uh, AI could be used on the, Quote, unquote, battlefield. That being the case, I think it would be a big mistake to view it analogously to a weapon and to argue that it should be closed up for a couple of reasons. First of all, it's not going to work at all. Uh, so I think we have, uh, you know, close to state of the art AI models that can already run on people's personal devices, and I think that trend is only going to accelerate. We talked earlier about Moore's law, and it still continues to apply Uh, to these models, and even if one country decides that models should be closed, the odds of getting every country to enact that kind of, ah, ah, rule are, you know, just vanishingly small. So if our approach to safety with AI is going to be premised on ensuring that quote unquote bad guys don't get access to it, we've already lost, because it's only a matter of time before it becomes impossible to do that. And instead, I think we should radically embrace the opposite,…
AI assessment note: “it would be a big mistake to view it analogously to a weapon and to argue that it should be closed”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q GP in your pocket with AI. Are you high? Like, GPs feel your elbow. They look at x-rays. They, um, look inside your ear and see very specific things. They look up your nose. You're not gonna shove your smartphone up your nostril. You know, it can't feel your arm. Can you help me understand why I'm wrong and why AI will revolutionize medical with a GP in everyone's pocket?
A Sure. Uh, so I don't think you're wrong. Uh, I think the reason there is a lot of talk about this is, uh, it goes back to something we've observed over and over, which is that when there are problems with an institution like the medical system, Right? Like the wait times are too long, or it's too costly, or, uh, in a lot of countries, you know, people don't even have access, you know, in developing countries, there might be entire villages with, uh, no physician, uh, then this kind of technological band-aid becomes very appealing. So I think that's what's going on here. I think the responsible way to use AI in medicine is for it to be integrated into the medical system. And actually the, the medical system has been a very enthusiastic adopter of technology, including AI. So you can consider, you know, CAT scans, for instance, to be a form of AI to be able to reconstruct, um, you know, what's going on inside a person based on, uh, certain imaging, and now with generative AI as well, there's a lot of interest from the medical system in figuring out, you know, can this be useful for diagnosis or for more mundane things like summarizing medical notes and so forth. So I think that work is really important. I think that should continue. Uh, it still does leave us with the harder question of, you know, here in America, you know, if it takes me three weeks to get a GP appointment, it's…
AI assessment note: “so I don't think you're wrong”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q have a tutor in your pocket. Yeah, I get you, but we do also have your videos that we can watch at home. Like a tutor has personal relationships. It's one-to-one where I want to impress you, Arvind, and I have that personal desire to fulfill, you know, abilities, potentials that doesn't have. How do you think AI impacts the future of education, one-on-one tuition, and that up-leveling of students?
A I think there's, Uh, you know, different populations of students. I think, you know, there's a small subset of learners who are very self-motivated, will learn very well, even if there's no, uh, you know, physical tutor, uh, whether it's at the, uh, uh, the primary school level, or it's at the college level, or at the expert level. I think those, there are those kinds of learners at, uh, at all different levels. And then there's the vast majority of learners for whom the social aspect of learning is really the most critical thing. And if you take that away, um, they're just not going to be able to learn very well. And I think this is often forgotten, especially because in the AI developer community, there are a lot of these, uh, self-taught learners. I'm among them, right? I just paid zero attention throughout school and college and everything that I know literally is stuff that I taught myself. So I grew up in India. The education system wasn't very great there. Uh, our geography teacher thought that India was in the southern hemisphere. True story. Right, right. So again, I, I literally mean it when I say everything that I know I taught myself. Um, and so, you know, you have a lot of AI developers who are thinking of themselves as the typical learner, and they're not. And I think for someone like me, AI is on a daily basis, uh, an incredible, uh, tool for, for learning. I use…
AI assessment note: “the vast majority of learners for whom the social aspect of learning is really”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What do you think society's biggest misconception of AI is today?
A I think our intuitions are too powerfully shaped by sci-fi portrayals of AI, and I think that's really a big problem. Uh, you know, this idea that AI can become self-aware. When we look at the way that AI is architected today, that kind of fear has no basis in reality. Uh, maybe one day in the future, you know, people are going to build, uh, AI systems where that becomes, uh, at least somewhat possible. And we should, you know, we should have, uh, visibility, transparency, monitoring, regulation around these systems to make sure that developers don't. But that would be a choice. That's a choice that society can make, that governments and companies can make. It's not that despite our best efforts, AI is going to become conscious and have agency and do things that are harmful to humanity. That whole line of fear, I think, is, uh, completely unfounded.
AI assessment note: “this idea that AI can become self-aware. ... that kind of fear has no basis”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What about synthetic data? What about the creation of new data that doesn't exist yet?
A So there's two ways to look at this, right? So one is the way in which synthetic data is being used today, which is not to increase the volume of training data, but it's actually to Overcome limitations in the quality of the training data that we do have. Uh, so for instance, if in a particular language there's too little data, you can try to augment that, or you can try to, um, uh, let's say have a model, uh, you know, solve a bunch of mathematical equations, throw that into the training data, uh, and so for the next training run, that's going to be Part of the pre-training, and so the model will get better at doing that. And the other way to look at synthetic data is, okay, you take one trillion tokens, you train a model on it, and then you output 10 trillion tokens, so you get to the next bigger model, and then you use that to output a hundred trillion tokens. You know, I'll, I'll bet that's just not going to happen. That's just a snake eating its own tail, and what we've learned in the last two years is that the quality of data matters a lot more than the quantity of data. So, If you're using synthetic data, uh, to, uh, try to augment the, the quantity, I think it's just coming at the expense of quality. You're, you're not learning new things from the data. You're only learning things that are already there.
AI assessment note: “So there's two ways to look at this, right?”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q GP in your pocket with AI. Are you high? Like, GPs feel your elbow. They look at x-rays. They, um, look inside your ear and see very specific things. They look up your nose. You're not gonna shove your smartphone up your nostril. You know, it can't feel your arm. Can you help me understand why I'm wrong and why AI will revolutionize medical with a GP in everyone's pocket?
A Sure. Uh, so I don't think you're wrong. Uh, I think the reason there is a lot of talk about this is, uh, it goes back to something we've observed over and over, which is that when there are problems with an institution like the medical system, Right? Like the wait times are too long, or it's too costly, or, uh, in a lot of countries, you know, people don't even have access, you know, in developing countries, there might be entire villages with, uh, no physician, uh, then this kind of technological band-aid becomes very appealing. So I think that's what's going on here. I think the responsible way to use AI in medicine is for it to be integrated into the medical system. And actually the, the medical system has been a very enthusiastic adopter of technology, including AI. So you can consider, you know, CAT scans, for instance, to be a form of AI to be able to reconstruct, um, you know, what's going on inside a person based on, uh, certain imaging, and now with generative AI as well, there's a lot of interest from the medical system in figuring out, you know, can this be useful for diagnosis or for more mundane things like summarizing medical notes and so forth. So I think that work is really important. I think that should continue. Uh, it still does leave us with the harder question of, you know, here in America, you know, if it takes me three weeks to get a GP appointment, it's…
AI assessment note: “I don't think you're wrong... then this kind of technological band-aid becomes very appealing.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Can I say another one that does worry me is actually defense. You know, we had Alex Wang from Scale On I mentioned earlier, he said that AI has the potential to be a bigger weapon than nuclear weapons. How do you think about that? And if that is the case, should we really have open models?
A I think, you know, it's, it's a good question to ask. I think it's a bit of a, a category error there. I mean, a nuclear weapon is an actual weapon. AI is not a weapon. AI is something that, you know, might enable, uh, adversaries to do certain things more effectively to, um, you know, for example, find, uh, vulnerabilities, cybersecurity vulnerabilities in critical infrastructure, right? So that's one way in which, uh, AI could be used on the, Quote, unquote, battlefield. That being the case, I think it would be a big mistake to view it analogously to a weapon and to argue that it should be closed up for a couple of reasons. First of all, it's not going to work at all. Uh, so I think we have, uh, you know, close to state of the art AI models that can already run on people's personal devices, and I think that trend is only going to accelerate. We talked earlier about Moore's law, and it still continues to apply Uh, to these models, and even if one country decides that models should be closed, the odds of getting every country to enact that kind of, ah, ah, rule are, you know, just vanishingly small. So if our approach to safety with AI is going to be premised on ensuring that quote unquote bad guys don't get access to it, we've already lost, because it's only a matter of time before it becomes impossible to do that. And instead, I think we should radically embrace the opposite,…
AI assessment note: “I think it would be a big mistake to view it analogously to a weapon”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Arvin, what did you believe about Kind of the developments that we've seen in AI over the last two years that you now no longer believe?
A So I think, like a lot of people, I was fooled by how quickly after GPT-III.V, GPT-IV came out. It was just, you know, three months or so, but it had been in training for 18 months. That was only revealed later. So it gave a lot of people, including me, an inflated idea of how quickly AI has, was progressing. And what we've seen in the nearly year and a half since GPT-IV came out is that we haven't really had models That have, uh, surpassed it in a meaningful way. Uh, and this is not based on benchmarks. Again, I think benchmarks are not that useful. It's more based on vibes. When you get people using these things, what do they say? I don't think models have, you know, really qualitatively improved on, on GPT-IV. And I don't think things are moving as quickly as I did 12 months ago.
AI assessment note: “it gave a lot of people, including me, an inflated idea of how quickly AI”