Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q foundry for AI. Here, we talk about what's next for Scale as models, uh, approach and go beyond human abilities. With great power comes great responsibility. Um, if, uh, you know, if these AI systems are what we think they are in terms of societal impact, like, trust in those systems is a crucial question. Like, how do you guys think about this as part of your work at Scale?
A A lot of what we think about is how do we utilize, how does the data foundry, um, enhance the entire AI lifecycle, right? And that lifecycle goes from, you know, A, ensuring that there's data abundance, as well as data quality going into the systems, but also being able to measure the AI systems, which builds confidence in, in AI, and also enables for further development and further adoption of the technology. And this is, this is the fundamental loop that I think every AI company goes through, you know, They, they get a bunch of data, or they generate a bunch of data, they train their models, they evaluate those systems, and they sort of, you know, uh, go again in the loop. And so evaluation and measurement of the AI systems is a critical component of the life cycle, but also a critical component, I think, of, of society being able to build trust in these systems. You know, how are governments gonna know that these AI systems are, are safe and secure and fit for, uh, you know, broader adoption within their countries? How do, how are enterprises going to know that when they deploy an AI agent or an AI system that it's actually going to be good for the consumers, and that it's not going to create greater risk for them? How do, um, how are labs going to be able to consistently measure what are the intelligences of my, of the AI systems that we build, and how are we going to, you …
AI assessment note: “evaluation and measurement of the AI systems is a critical component... of society being able to build trust”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q With great power comes great responsibility. Um, if, uh, you know, if these AI systems are what we think they are in terms of societal impact, like, trust in those systems is a crucial question. Like, how do you guys think about this as part of your work at scale?
A A lot of what we think about is how do we utilize, how does the data foundry, um, enhance the entire AI lifecycle, right? And that lifecycle goes from, you know, A, ensuring that there's data abundance, as well as data quality going into the systems, but also being able to measure the AI systems, which builds confidence in, in AI, and also enables for further development and further adoption of the technology. And this is, this is the fundamental loop that I think every AI company goes through, you know, They, they get a bunch of data, or they generate a bunch of data, they train their models, they evaluate those systems, and they sort of, you know, uh, go again in the loop. And so evaluation and measurement of the AI systems is a critical component of the life cycle, but also a critical component, I think, of, of society being able to build trust in these systems. You know, how are governments gonna know that these AI systems are, are safe and secure and fit for, uh, you know, broader adoption within their countries? How do, how are enterprises going to know that when they deploy an AI agent or an AI system, that it's actually going to be good for the consumers, and that it's not going to create greater risk for them? How do, um, how are labs going to be able to consistently measure what are the intelligences of my, of the AI systems that we build, and how are we going to, you…
AI assessment note: “evaluation and measurement of the AI systems is a critical component... of society being able to build trust”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Can you give our listeners a little bit of intuition for, like, what makes evals hard?
A One of the hard things that, you know, because we're building systems that We're trying to approximate and, and build human intelligence. Um, grading one of these AI systems is, is not something that's very easy to do automatically. And it's, it's sort of like, um, you know, you have to kind of build IQ tests for these models, which in and of itself is a very fraught philosophical question. It's like, how do you measure the intelligence of a system? And this, there's very practical problems as well. So most of the benchmarks that we as a community look at For the academic benchmarks. Yeah. The academic benchmarks that are what the industry used to measure the performance of these algorithms are fraught with issues. Many of the models are overfit on these benchmarks. They're sort of in the training data sets of these models. Um, and so.
AI assessment note: “grading one of these AI systems is, is not something that's very easy to do”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I was just curious. I mean, one thing related to that is you mentioned that, for example, JP Morgan has a 150 petabytes of data relative, you know, it's a 150 times what, um, some early GPT models trained on. How do you work with enterprises around those loops or what are the types of customer needs that you're seeing right now or application areas?
A One of the things that every, you know, all the model developers understand well, but the enterprises understand super well is that, um, you know, not all data is created equal and high quality data or frontier data is, is, can be, you know, 10,000 times more valuable than just any run of the mill data, uh, within an enterprise. And so a lot of the challenge in, you know, or a lot of the problems that we solve with enterprises are how do you go from this, Giant mountain of data that sort of is like all over, is truly all over the place, um, and distributed everywhere within the enterprise to what are the, how do you compress that down and filter it down to the high quality data that you can actually use to, you know, fine tune or train, um, or continue to enhance these models to actually drive differentiated performance.
AI assessment note: “a lot of the problems that we solve with enterprises are how do you go from this”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q Does the new OpenAI or Google, ah, release, ah, change your point of view on anything fundamentally? Multi-modality, um, you know, the applicability of voice agents, et cetera?
A You know, I think you tweeted about this, but one, one very interesting element is, uh, the, the direction that we're going in terms of consumer focus. And it's, it's fascinating. I mean, I think multimodality, well, taking a step back, first off, I think it points to where there's, where there's still huge data needs. So multimodality as an entire space is one where, um, for the same reasons that we've like exhaust a lot of the internet data, there's a lot of scarcity for good multimodal data that can empower these personal agents and these personal Use cases. So I think there's, um, you know, as we want to keep improving these systems and improving these personal agent use cases, there's, you know, we think about this a lot. What are the data needs that are actually, um, that are going to be required to actually fuel that? Um, I think the other thing that's fascinating is, is, um, the convergence actually. So both labs have been working, um, independently on, on various technologies and You know, Astra, which is Google's, uh, major sort of, uh, hubcap release, as well as four. Oh, you know, they're both shockingly similar, um, and, uh, sort of, uh, you know, demonstrations of the technology. And so there's, I think that was, that was very fascinating. The labs were sort of converging on the same end use cases or the same visionary use cases for the technology.
AI assessment note: “I think multimodality, well, taking a step back, first off, I think it points to”
Redirected raw tape
D 2 · C 4 · P 3 · Cm 3 3.00
Q is there's other types of parties that are starting to engage with this technology. So you're obviously now working with a lot of the technology giants, with government, um, with automotive companies. It seems like there's emergence now of enterprise customers and a platform for that. There's emergence of sovereign AI. How are you engaging with these other massive use cases that are coming now on the generative AI side?
A It's quite an exciting time because I think for the first time in a, in maybe the entire history of AI, AI truly feels like a general purpose technology, which can be applied in, you know, a very large number of business use cases. I contrast this to, you know, the autonomous vehicle era where it really felt like we were building a very specific use case that happened to be very, very valuable. Now it's general purpose can be, it can, uh, be encompassed across the, the broad span. And as we think about what are the infrastructure requirements to support this broad industry, And what is the, what is the broad arc of the technology? Um, it's really one where we think, how do we empower data abundance, right? Um, there's a, there's this question that comes up a lot, you know, are we going to run out of tokens? Uh, and, uh, and what happens when we do? And I think that that's a choice. I think we as an industry can either choose data abundance or data scarcity. Um, and we view our role and our job in the ecosystem to be, to build data abundance. Um, The key to the scaling of these large language models and the, you know, these, these, uh, uh, language models in general is the ability to scale data. And I think that one of the fundamental bottlenecks to, you know, what's, what's in the way of us getting from GPT-IV to GPT-X is, you know, data abundance. Are we going to have the data…
AI assessment note: “we view our role and our job in the ecosystem to be, to build data abundance.”