Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q making this move. Uh, where it's trying to build for, I think for the first time since GPT for its own model that competes with it. How should we read into what Microsoft is doing there? Do you think it is a loss of faith in open AI? It's a hedge against open AI, sort of an unnecessary move, even though it has such an important partner. What's your take?
A Yeah. One of my friends who works at DeepMind told me that Microsoft is basically reversing what Google has managed to do over the last few years. And in fact, making the same mistake that Google initially made, which was to have its training distributed, uh, or split up between two different corporations or institutions. For Google, it was a brain, Google brain and DeepMind. And so Microsoft has a company which is in the lead, right? OpenAI. And I guess instead of doubling down on it, they're trying to hedge their bets in this way. I think if you think that open AI is like another product where you have multiple vendors, so you can be sure that if one of them, uh, you know, has, decides to go a different route, you have some leverage over them. That, that might make sense for another kind of product. The thing with AI is if you buy scaling in this picture, that is, you make the models bigger, they get much smarter. Then I don't think it makes sense to hedge your bets in this way. I think you should just double down, give, give one of them a hundred billion dollars and just say like, go make me, go make me super intelligence. You know what I mean? Like, cause then you're just splitting up your efforts and yeah, like it would be much better to have one GPD four than two companies that have, uh, you own two companies that have a GPD 3.5.
AI assessment note: “making the same mistake that Google initially made, which was to have its training distributed”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q it's going to take to make this stuff work and Are we ever going to get to the place where a lot of these people want to get to talking about like adding more compute and data and energy, and eventually you get to the point where you can train better large language models and see what the scaling law really looks like at its limit. What do you think?
A Yeah, I think compute will be less of a bottleneck than energy. At sort of the seven trillion, I, yeah, I imagine, um, well, the backing up, the reason I think compute will be a less bottleneck than energy is because right now you have one company, NVIDIA, which is making the sort of, uh, GPUs, and other than Google, nobody has a clear competitor, and so the, the thing that was bottlenecking NVIDIA so far is that some of their The components that the need for these GPUs, uh, COOS and HBM, uh, they just weren't able to get enough allocation or get, uh, TSMC to build facilities for these, because TSMC was like, I don't know if we buy all this AI stuff, but, uh, because then they had to make this huge investment into building it out, but now it seems like the, uh, fabs are building it out, and also all these companies have accelerator programs where they're gonna try to ship their own chips, so I think Compute will become more and more available. And that's what Zuckerberg said on the podcast that now the compute constraints are decreasing. Then the question that Zuckerberg pointed to was, well, will there be energy? And the, the key constraint with energy is not necessarily, is there enough energy in the world, but more so for training, is there enough energy in one place? Because to do a training run, it has to usually, at least from what it seems like publicly, the training met…
AI assessment note: “I think compute will be less of a bottleneck than energy.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q rate to get proof of concepts out the door is pretty small. One out of every five actually gets shipped into production, and often it's a scaled down version of that. So what you're saying is interesting. You're saying it's not their fault. It's that these models are not reliable enough to do what they need to do because they don't learn on the job. Am I getting that right?
A Yeah, and you're talking about reliability. It's just, um, they just can't do it. So if you think about what makes humans valuable, It's not their raw intelligence, right? Any person who goes onto their job the first day, even their first couple of months, maybe they're just not going to be that useful because they don't have a lot of context. What makes human employees useful is their ability to build up this context, to interrogate their failures, to build up these small improvements and efficiencies as they practice a task. And these models just can't do that, right? You're stuck with the abilities that you get out of the box and they are quite smart. So you will get five out of 10 on a lot of different tasks that they'll Often they'll, on any random task, they'll probably might be better than an average human. It's just that they won't get any better. Um, I, for my own podcast, I have a bunch of little scripts that I've tried to write with LLMs where I'll get them to rewrite parts of ION scripts to make them more, uh, turn auto-generated transcripts and do like human written like transcripts or to help me identify clips that I can tweet out. So these are things which are just like short horizon language in language out tasks, right? This is the kind of thing that the LLM should be Just amazing guy because it's a debt center in their, uh, of what should be in their repertoir…
AI assessment note: “Yeah, and you're talking about reliability. It's just, um, they just can't do it.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q dollars. There is a pressure to deliver to investors. There are reports that safety is becoming less of a priority as market pressure makes them go and ship without the typical reviews. So is this kind of a risk for the world here that these companies are developing this stuff? Many started with the focus on safety and now it seems like safety is taking a backseat to financial returns.
A Yeah, I think it's definitely a concern. Like, We might be facing a tragedy of the common situation where obviously all of us want our society and civilization to survive. Um, but maybe the immediate incentive for any lab CEO is to Look, if there is an intelligence explosion, it has a really tough dynamic because if you're a month ahead, you will kick off this loop much faster than anybody else. And what that means is that you will, uh, you will be a month ahead to super intelligence, but nobody else will have it right. Like you will get, you'll get the 1000 X multiplier on research much faster than anybody else. And so it could be a sort of winner take all kind of dynamic there. Um, and Therefore they might be incentivized. Like I think to keep this system, keep this process in check might require slowing down, um, using these alignment techniques to like that, which might be sort of a tax on the speed of the system. And so, yeah, I do worry about the, the pressures here.
AI assessment note: “Yeah, I think it's definitely a concern. Like, We might be facing a tragedy”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q it wishes it understood what the audio sounded like, because it knows that that's an important data point. But anyway, let's just put that aside for a moment. The memory thing is interesting. Do you think there are easy ways to To then like have, have a persistent conversation with one of these bots, or is that going to be like another tough problem that we won't solve for awhile?
A I, it could plausibly be very tough because I don't think it's a member. It's a, it's a matter of just keeping like storing memory. I think it's like, what kind of thing are you and are you a chat bot or are you, is your persona like I am an entity that, you know, it's, it's not just about like, I'm storing these things somewhere. It's like, you have to train it to act as an agent and compared to just pre-training tokens on the internet where Yeah. It knows how to complete statements. Does it know how to act as an agent? There's not necessarily a good way to structure that. So people have been talking about long horizon RL, which is the training method. You need to get something like this, where you go tell it to do something and then you reward it at the end for having achieved that outcome. But the difficulty with those kinds of approaches and the difficulty with RL in general is sparse reward and, uh, non-stationary distributions, which is like, uh, you know, like You failed to book me my right, the right appointments based on like reading out my inbox and like talking with me about it. Why did you fail? There's like so many different reasons you could have failed. That's hard to attribute to any one of them. You know what I mean? It's like hard to learn from that. It's kind of an interesting question, honestly, like why humans are so good at, uh, learning from these sparse …
AI assessment note: “it could plausibly be very tough because I don't think it's”
Partly raw tape
D 3 · C 4 · P 2 · Cm 2 2.90
Q Yeah. Then the second question we had was, uh, someone says, give us the Dwarkesh interview prep playbook. That's his innovation, and if he's able to explain it in the way that can be replicated or at least approximated by others, we'll have many more interesting interviews. Okay, I'm honestly self-interested in this as well, so how do you do it?
A Yeah. Well, I know it sounds like first to say this or, but I honestly just like I prep a lot, and I think there's also a flywheel by doing interviews. I learned a lot of things, and because of that, I can get better interviews, learn more things. I think the main flywheel, honestly, is that I make the podcast better, smarter people listen, some of those smart people I become friends with, and they teach me a bunch of things. And now I can do an even better interview. Now I have a bunch, I can get connected to a bunch of other smart people. They teach me more things. So I think that's like, that's going to be a big part of the flywheel that people may not know about.
AI assessment note: “I honestly just like I prep a lot, and I think there's also a flywheel”