The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mustafa Suleyman no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I'm not saying that's what's gonna happen with you. Um, and maybe it makes sense just to, you know, buy off the shelf or, uh, use open source. And in fact, that seemed like that was the strategy that you had for a long time. It seems logical. Um, and so I'm curious, like why you would disagree with that? Why is it so important to build your own models?

A I mean, we, we're going through a foundational platform shift, um, you know, in software, um, from the operating system to apps, from browsers, search engines, mobile, social. This is the next major platform, and it's going to be bigger than all of the other platforms put together. So the idea that a three trillion dollar company with three hundred billion dollars of revenue and 80% of the S&P 500 on our Azure stack and M three five stack Um, you know, could, could depend on a third party. This it's, you know, just in perpetuity. It doesn't make sense. So we, we, you know, this is a company that's been around for 50 years, uh, and navigated many of the past platform shifts incredibly well. And that's the, that's the journey that we're on. We have to be AI self-sufficient. There's an important mission that, uh, Satya set last year. And I think that we're, we're now on a path to be able to do that.

AI assessment note: “The idea that a three trillion dollar company... could depend on a third party”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Is it possible that something can be super intelligent, but not generally intelligent? Like, is it possible that maybe super intelligence happens without AGI because AGI is all about generality and what you're talking about is not?

A It's not possible, I don't think. I think they need to be general. Um, they need to transfer knowledge from one domain to another. They need to, um, you know, um, have generalist reasoning capabilities. But when you apply it, and you put it into production, and you let it have more autonomy to make decisions, or you let it generate, uh, arbitrary code to solve a particular problem, Or you let it write its own evals so that it can modify its own code and generate new prompts to generate new training data, to write new evals, to then iterate on its own performance. These capabilities, autonomy, goal setting, um, writing code, uh, modifying itself, you know, if you add to that then also a perfectly generalized model or, or, or sort of general purpose model, That's a very, very, very powerful system, which today I don't think anybody really knows how we would contain or align something like that. And so it's not to say that we should not do any one of those dimensions. It's just to outline a roadmap of capabilities which we're all working on, which add more risk, especially when they compound with one another and you combine them all together. And so, you know, my claim is that we should just approach this with caution, remembering That we don't want to bundle together all these capabilities so that there's a higher risk of a, you know, recursively self-improving exponential takeof…

AI assessment note: “It's not possible, I don't think. I think they need to be general.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So is it really just making it more personable than the others? Like that's the way to differentiate?

A Yeah, I, I think so. I think like at the end of the day, um, we are like at the very beginning of a new era where there are going to be as many copilots or AI companions as there are people. There are going to be agents in the workplace that are doing work on our behalf. And so everyone is going to be trying to build these things. And what is going to differentiate is real attention to detail, like true attention to The personality design. I've been saying for many years now, we are actually personality engineers. We're no longer just engineering pixels. We're engineering tokens that create feelings, that create lasting, meaningful relationships. And that's why we've been obsessed with the memory, the personal adaptation, the style, and really just declaring that it is an AI companion, you know, not a tool, right? A tool is something that does exactly You know, what you intend, what you direct it to. Whereas, you know, an AI companion is going to have a much richer, more kind of emergent, dynamic, interactive style. It will change, you know, every time you interact with it, it will give a slightly different response. So I think it's going to feel quite different to past waves of technology.

AI assessment note: “Yeah, I, I think so... what is going to differentiate is... The personality design.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q we're definitely seeing these new modalities come out. Voice, of course, we've obviously this, we're in the middle of a firestorm with images, um, but I am, okay, I guess, let me ask the previous question a little bit differently. What, do you think that there are diminishing returns on pre-training right now? Basically scaling up the biggest possible model and then building from there? You're shaking your head no.

A Maybe specifically on pre-training, it's been a little, um, slower than it was in the previous four orders of magnitude. But the same computation, um, the same flops, or the, the units of calculation that go into turning data and compute into some insight into the model, that is just a different application of the compute. We're, we're using compute at a different stage. We're either using it at post training, or we're using it at inference time, where we generate lots of synthetic data to sample from. So net net, we're still spending as much on computation. Um, it's just that we're using it in a different part of the process, but, but for, for as far as everyone else's, you know, should be concerned aside from the technical details, we're definitely still seeing massive improvements in capabilities. And I think that's, that's for sure going to continue.

AI assessment note: “Maybe specifically on pre-training, it's been a little, um, slower”

Answered raw tape D 4 · C 5 · P 4 · Cm 3 4.15

Q And so how does it decide which test to order and how to optimize cost? Is that another feature within the model or?

A Well, I mean, you can think of an unnecessary test as a kind of error, a human error. So really what the model is trying to do is to get to the best diagnosis with the minimum number of tests. And, you know, obviously, you know, the, the model has much broader, you know, sort of range in terms of Awareness of which test results tend to correlate with which particular, ah, diagnostic outcomes, and so given that it's seen so many more cases than any given human, then it's obviously showing that it can do a better job of judging in this instance, given this case history that it already knows about a patient, what is the sort of minimum number of necessary tests to get the next piece of information to be able to continue the, um, the diagnosis and get it more accurate.

AI assessment note: “what the model is trying to do is to get to the best diagnosis”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q the possibility Um, that we might have some serious change here come to our work. And, uh, you had said, um, that AI is gonna create a serious number of losers in white collar work. Maybe it already is. I've sort of changed my tune in thinking that, you know, we're all fine in the white collar work world, and now thinking, well, it's anyone's guess. So what's coming, Mustafa?

A I do think that that is the big Story that we should be talking about. That's the transition that's going to happen over the next 15 years, um, is that it is going to be a cheap and basically abundant resource to have these reasoning models that can take action in your workplace, that can orchestrate your apps and get things done for you on your desktop. Like, that really is quite a profound shift in, in how we work today, and I do think that, like, The, the, your day-to-day workflow just isn't going to look like this in 10 or 15 years time. It's going to be much more about you managing your AI agent, you asking it to go do things, checking in on its quality, getting feedback, and getting into this, like, symbiotic relationship where you iterate with it and you create with it and solve with it. Uh, that's going to be massively more efficient, and I do think it's going to make everybody a lot more creative and productive. I mean, after all, It is intelligence that has produced everything that is of value, um, in our human civilization. Like everything around us is a product of smart human beings getting together, organizing, creating, inventing, and producing everything that you see in your, you know, line of sight at this very moment. And we're now about to make that very same technique, those set of capabilities, really cheap, um, if not like zero marginal cost. And so, you kn…

AI assessment note: “your day-to-day workflow just isn't going to look like this in 10 or 15 years”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.