The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Misha Laskin no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Maybe six months ago, I think you, you, you said like, I think it's possible we have my definition of ASI in a couple of years. Um, do you still believe that's true?

A I, I still do believe that's true. Um, I think that where I think we'll be in a couple of years from now is that there'll be kind of definitive, um, super intelligence in, Some meaningful categories of work. And so, for example, when I say coding, I don't mean all of coding there, but there will be a super intelligence within some kind of slivers, some meaningful slivers of coding that are driving, um, I'd say immense progress in the companies that can benefit from that. And the reason why I would say that the problem of ASI would have been solved by then is because you've kind of, um, at that point, it's just a matter of operationalizing. Like what, you know, you know, it just so happened that these particular categories, like you might have a super intelligent front-end developer because there's so much data distribution for that on the internet, and it's easier to make synthetic data for that. But at that point, you have the recipe, and it's just a matter of, um, making kind of economic decisions of, is it worth sinking in X amount of dollars to get the data in this category, um, to get kind of something, um, close to super intelligence there. Um, an example of that is, um, What happened with reinforcement learning before language models? Um, effectively, the blueprint for building super intelligent systems was developed. It happened with, um, the Atari games, AlphaGo, um, y…

AI assessment note: “I, I still do believe that's true.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What is an example as you guys open up the waitlist that you want users to try where it should just be like obvious that the answers are, are better than other coding agents?

A I think the kinds of, um, queries that it tends to be better at are I guess what we would call Semantic queries. So let's say like an example of a query where this is not the best system to use. It's like file level. If you're looking at a file, and there's like a specific thing in that file, and you're just trying to get a quick answer to it, you don't really need the hammer of like a deep research like experience. Um, you don't need to wait, you know, like tens of seconds or a minute or two, uh, to, to get that answer, because that should just be delivered snappily. But if you, um, Don't exactly know where you're looking for, and you, you know, you don't know the function name, or you don't, you know, something, and this is kind of the hard problems that engineers are usually in. Like, there's a flaky test. I mean, you know that this test is flaky, but that's where your knowledge stops, right? And that's when you usually go to Slack and ask some engineers, like, this test is flaky. What's going on? Does anyone know? Um, you know, we've had, uh, the way we've used it is when you're training these models, there's a lot of infrastructure work that goes into it. And, um, it fails in interesting ways all the time. Uh, and asking things like, you know, my jobs are running slowly, five times more slowly than usually. Why is that? That's kind of a vague query that would be very hard …

AI assessment note: “asking things like, you know, my jobs are running slowly... Why is that?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You have to have a concept of authority, right? People are going to say things that are wrong.

A The way it's worked with customers we've started working with is, uh, they typically have, they, they want to start off with kind of a group of trusted kind of senior, you know, staff level plus engineers who are kind of the gatekeepers, which is a very, I think, common notion. Um, you have permissions, right? And ownerships, uh, ownership structure and code bases. And they basically are the ones who kind of populate the memory first. Um, and then sort of expand the scope, but I think it works. It's actually a much more complex feature to build, uh, because it touches on, um, yeah, org wide permissions. Um, there's some parts of the code where a certain engineer should be able to edit the memory, but other engineers shouldn't. Um, and so it, it actually starts looking like the new way of, um, versioning code effectively, right? It's kind of a GitHub plus plus, uh, because you're not versioning the code, you're kind of versioning the meta knowledge around it. That helps language models understand it better. Uh, but definitely that is something that we built, but I think it's a thing to iterate a lot until you kind of get the right design here because you're effectively building and yeah, and you, and you get from scratch.

AI assessment note: “staff level plus engineers who are kind of the gatekeepers”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay, last question, Misha, where would you characterize us as, like, being on the path toward deployment of these capabilities in, in different fields?

A I think we're a lot earlier than most people think, uh, that this is going to be one of those areas where the technological building blocks, um, outpace their deployment. And so, yeah, within the next couple of years, uh, the blueprint roughly for, you know, how to build ASIs will have been set more or less. Like, uh, maybe there's still some, um, efficiency Breakthroughs that need to happen, um, but more or less there'll be a blueprint for how do you build a super intelligence in a particular category? Actually going in and deploying it and, and building it for, you know, specific categories of work. There are going to be a lot of product and kind of research innovation specific to those categories, um, that will probably make this a multi-decade thing. Um, so I don't think that it's a couple of years from now and, uh, GDP starts growing 10%, um, you know, year over year globally. I think We're actually going to get there, uh, but it's going to be a, kind of, multi-decade, uh, endeavor. I tend to, kind of, see a lot of patterns, uh, now in, kind of, real-world deployment with, uh, reinforcement learning, um, research as it worked, again, before large language models. Um, and before large language models, it used to be, kind of, you pick an environment, like, you pick Go, you pick, um, StarCraft, you pick something else, and You go and try to solve it with, you know, some combi…

AI assessment note: “I think we're a lot earlier than most people think”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Yeah, I definitely, uh, I definitely want to talk about what you think is solved and unsolved here. Um, the entire field has clearly gotten more focused on, um, deep reinforcement learning over the last 18 months. You have this, uh, huge product launch this week with Asimov. Um, can you just sort of describe what it is?

A So Asimov is, uh, the best Code research agent in the world. It's a comprehension agent, meaning that it's really designed to kind of feel almost like a deep research for large code bases. The way a developer is supposed to feel interacting with it is effectively like they have a principal level engineer who deeply understands their organization at their fingertips. Uh, so it's very different from the existing set of tools that I focus primarily on code generation. Like every single coding tool has some code generation and some comprehension aspect. But as we spent a lot of time kind of with our customers, um, trying to understand why Coding tools. And this is enterprise specific. So I think, I think the world is different with startups, but within enterprises, when you, you know, they're adopting coding tools and you see the impact that this is having, um, on their actual productivity. And I think it's much lower than people, uh, expect. Um, so it's, uh, in fact, it's, it's sometimes negative, sometimes negligible.

AI assessment note: “So Asimov is, uh, the best Code research agent in the world. It's a comprehension agent”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q you had reason through it with context of the system, you never would have made such a mistake or like a, you know, works and works in my environment type problem. Um, and, and so I think that very much mirrors my, you know, intuitive understanding of engineering here. That's great as problem formation. What, what makes Asimov different in terms of ability to understand better versus just generate code?

A There are a few things. So I think this is kind of where, you know, why it is so important to co-design research and product, because as a researcher, you'd go in and say the answer is entirely in the agent design or the model or something like this. And as a product person, you would say, well, it's in these product, you know, differentiators, like being able to draw not just from your code base, but knowledge that lives, you know, in other sources of information or being able to learn from I have the engineering team to offload, uh, their tribal knowledge. So, right, an engineer can go in and teach Asimov, like, hey, uh, we deploy our, you know, when we say environment jobs on our team, we mean this specific thing, which we mean kind of Google that job. So now when another engineer asks a question about environment jobs in the future, the system just knows what they're talking about. A lot of knowledge is stored in engineers' heads, and I think you need, um, both of these things. You need to Understand your customer really closely and develop differentiated product almost independently, right, of the models that are towering it. Um, but then you also need to innovate on the research, uh, in terms of agent design and model training to actually drive the capabilities that you want to see out of the system. And this becomes an evaluation problem, which is basically at the heart …

AI assessment note: “being able to draw not just from your code base, but knowledge that lives”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.