Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q sort of twofold. Um, on the one hand, you're doing a large scale, uh, custom model Um, specifically in part focus on code, and then you're also building sort of at the product suite that can really help, um, address, uh, coding and working, uh, on the, the software development side. How did you decide to start Magic, and why, why focus on that, um, versus other aspects of AI?
A It sort of came from a place of working backwards from AGL. If you, uh, your end goal is to have a system that can do everything, uh, you can reduce that to building a system that can build that system. And so that minimal system is a system that writes code and comes up with ideas and can validate those, uh, by writing code and running experiments, which is still, like, in the same order of complexity as the full thing, but at least we don't have to train Zora. Uh, and, you know, we don't have to think about ten billion other use cases that everyone building general domain products has to think about. We only have to think about code. So it's a lot simpler on all aspects except compute and slightly simpler on the aspect and slightly cheaper on the aspect of compute. I, uh, I think it's not a lot cheaper. I probably overestimated how much cheaper it would get, um, on the compute site, uh, at the beginning. Uh, but the other things are simpler, I think.
AI assessment note: “It sort of came from a place of working backwards from AGL.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q The, uh, the other area of, um, expansive exploration right now amongst researchers for better AI code generation tends to be test time search. Um, so, you know, more compute at inference time, uh, you know, speaking of a, speaking of a mutual friend in Noam, like, how do you think about this?
A Well, so you can think of model performance as some function of training compute times some function of inference time compute. Now those are specific functions that are just scaling law things that you can like model, but the, the general Way to think about it is some function and some function, and then you would want to estimate how much inference you're going to do, and how much, what your total budget is, and then you would want to create the optimal trade-off, um, in your allocation of money. You also want to consider the distribution of outputs. There'll be users who will want to spend less money, and users will want to spend more money, and this is likely going to follow some sort of very sort of spiky distribution, where there'll be like four users spending a million dollars, and, you know, four billion users spending 10 cents, and A curve in between. It seems, um, strictly beneficial to be able to provide that choice. So instead of training the, putting all the compute into training and having that one million dollar inference performance be purely from the training compute, which is just hilariously inefficient, you can, you can allow the user to choose their, I think. So, um, or rather you can just deploy multiple things. The reason I can talk about this now is because everyone, like, is getting, like, everyone gets this. Um, but, but basically you clearly want to b…
AI assessment note: “model performance as some function of training compute times some function of inference time compute”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q So is the eval bar like, I'm not going to do code review?
A For the, so that's like the fully automation thing, right? And then you can launch something where you're like, I'm going to do code review, but it's not frustrating. Uh, if you have to do code review, and it's like really taxing, and you have to fix half of the problems, like half of the PRs or whatever. So I think the bar for this product is just high. It's not that like, we are like so ambitious, and you know, I think that too, but I, I just genuinely think that there is a gigantic market that gets unlocked in a step function moment where users decide that they're no longer, they're no longer, uh, going to use, uh, VS code to write code and send it to their colleagues. They're going to use magic or whoever ends up doing this well, uh, first, uh, to, to write their code for them and then briefly look at it and correct every now and then what, what has been done and then eventually not, right? But that is a step function moment, I think. It's not, like, you're not gonna use this for, like, five percent of your tasks. You're gonna use this for 90% of your tasks or zero.
AI assessment note: “you're like, I'm going to do code review, but it's not frustrating”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q And do you think the existence of that, I guess one could argue it can increase GDP, but it may also, um, decrease, uh, a lot of, sort of, human-driven activities. Um, how do you think about the eventual implications or impact of AGI?
A This is going to take longer. I'm going to try to compress it, but it's actually quite complicated. One of the biggest problems with this question is that everyone tries to simplify it by picking one side of the argument. For example, And this is just one example. Um, centralizing power is terrible, so therefore you should open source everything. If I stop now, this is reasonable, right? Please don't make a quote out of this because I don't believe this. Um, and then you could say, well, I should not give these, I should not, I should learn how to do PR. I should not say these sentences. Anyway, you can keep this in. So there's one way to say this. And then the other way to say this is, well, this is like nukes. We all fear this existential risk thing. Like maybe we care less about the, you know, we, we, or we care, it's not that we care less, it's just that we think these intermediate problems are completely solvable. Uh, but, but you can't open source how to build a nuke. This is just terrible. So, so the problem is that both of these things are true. And the problem is that there are 10 questions like this, and both of the answers are true in all of them. Um, and so what you get is people arguing on X, Claiming one side and ignoring the other. Um, and so, so I think there are like, 10,000 possible futures, and which one we end up with depends entirely on which answer we choo…
AI assessment note: “which one we end up with depends entirely on which answer we choose”
Answered raw tape
D 4 · C 3 · P 2 · Cm 2 2.90
Q So I guess there's the model side of it, and then there's sort of the productization of that model or the ways people access or interact with that. Is there anything you can share at this point regarding those types of things?
A As we have built the UX, each iteration of the, like, next UX internally, um, we thought about launching it. We were at this interesting stage of, it feels like an uncanny valley, um, where Like, well, you know, clearly you can see signs of life. First of all, I mean, completion is a trivial one, right? Like, like, we decided not to launch completions, because it's just obviously going to get killed by the next thing. And then we're like, okay, it's going to take us a few months to get a prototype of the next thing. We got a prototype of the next thing. And then when I get looking at this, it was like, you can see signs of life, but, you know, like, you guys tried it when you decided to invest. So it's, you know, it's like, would you be using this to, like, write your idea? No, not yet. But anyway, we can train the next model. And then, like, that model can do it. But then that model, like, can also do all these other things that, like, would go into final shape. So you just enter this, like, stupid recursive loop, uh, until the point we got, we got to this, we got to the point where we're like, okay, like, what's the final UX? Let's just, like, let's just make sure this never happens again. And so, so, um, we, we're trying to, like, meet the bar of the, that UX now, um, which I hope we will, um, sooner rather than later. The closer you get to it, the sort of dumber it feels to…
AI assessment note: “we're trying to, like, meet the bar of the, that UX now”
Answered raw tape
D 3 · C 3 · P 3 · Cm 2 2.85
Q I guess if the goal of the company is eventually to build AGI, um, how does that impact your choices from a design perspective relative to some of these trade-offs? Or is it more, you were going to iterate till we have a system that's very good at writing code and then it bootstraps its own next version? Or how do you think about your roadmap relative to AGI itself?
A Yeah, that's a great question. So I remember one of our first conversations, we were talking about sort of the step function relevance of safety risks, let's say, right? We're like, there's a lot of stuff people are panicking about that really doesn't matter in the short term, uh, for like the grand scheme of society. And like, some people will get pissed at me saying this, but it's just what I actually believe. Uh, I, I just don't think this is like that they, the current complaints are like similar to what we have seen in like all other technologies and like totally resolvable. And, but, but then there comes like the evolutionary one and I was like, okay, shit. Like, humans succeeded in the world because we're smarter than chimps, and we're smarter than bunnies, and, uh, you know, like, bunnies do not rule the world, uh, and, um, that would be cute, probably, but also, um, it's, it probably wouldn't be as nice for us, uh, and, like, here we are turning ourselves into bunnies and apes and creating this thing that, you know, we're all thinking is going to be way smarter than we are. This is insane, and, um, so the reason I think like the sort of recursive approach that, that you're mentioning here, and then obviously, you know, this is what, you know, this is, this is what we're founded to do, um, is exactly the, I think the only way to, to, to sort of reasonably approach this …
AI assessment note: “the only way to sort of reasonably approach this is to iteratively ask your model”