The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Maxim Bar Kogan no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I think I see, I need some guardian, guardian spirits around it. Um, given your deployments today and talking to large enterprises, what is the state of deployment, right? Uh, like how much do you see that's within these more scoped, like studio-like platforms versus, uh, you know, uh, free, free riding coding agents, you know, how, how much are you actually seeing in large enterprises and in different sectors?

A Yeah. So I think right now in our typical enterprise, we're going to see if we break it down to three categories. So we break it down to various SaaS platforms that are typically more low code and where people build agents in this drag and drag way. And they're not really autonomous agents, right? They're kind of the same kind of, I would think of them more as the automations. And then there are, um, first party agents, people are building in their cloud, Potentially because it's an application they want inside the company or even a product they're planning to release to the customers that is agentic. And then the third category is very autonomous coding agents and assistants. Of these categories, I would say roughly at this point, over 50% is the autonomous, uh, coding agents and assistants in the average enterprise. Then probably 45%, uh, is, uh, is those, uh, Uh, low code automations. And the last two percent are really the first party ones that they're building themselves because obviously it's much harder to build effective agents. So, and it's much easier to adopt agents off the shelf or, or build them with low code. So, and that's what we're seeing. And we do see that the autonomous users are also the fastest growing category. So it used to be that only developers and we would see cloud code growing like fire in our customer base. And now we're seeing a cloud cowork grow…

AI assessment note: “over 50% is the autonomous, uh, coding agents and assistants in the average enterprise”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I asked you the same question. I said, is anyone going to do this before you run out of money?

A And, and I think there was a good chance that, uh, I would have run out of money before, because I think you were right. Like, I think there was an element of chance here, but then I think the market did happen. So we had suddenly reasoning models that could do long horizon tasks. We had a cloud code, which became like the really first, uh, widely used autonomous agent. And then we had co-work and open claw. And, and I think We're starting to see now that these types of agents that are very autonomous, even though they're like, uh, everyone was afraid to build them. So everyone started building these low code platforms that were much more limited, much more based on connectors. And those platforms ended up being quite limited so that we didn't get the productivity gains from those limited platforms. But when we started getting the crazy benefits from these very unleashed agents that could do everything, that had much less controls baked into them, And even very large enterprises decided they're going to adopt it, you know, like Anthropix revenue is coming from enterprises that are paying for cloud code to do a lot of the work that developers used to do. That was a bit of a kind of how we started, and we definitely were in luck that very autonomous agents appeared before it was too late.

AI assessment note: “I think there was a good chance that, uh, I would have run out of money”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Our level, like human intelligence or our level of your models? Okay. Human intelligence.

A Yeah, I think like, yeah, exactly. I think as, as humans, it might still be very difficult to understand what weights and activations mean, and maybe mechanistic interpretability, it seems like, oh, maybe that's too hard or shouldn't be possible. But as we're starting to have models that are much smarter than us, at least in some important ways, we think that, uh, we'll be able to start tracking mechanistic capability much more effectively. Um, And I, and I think it's gonna be extremely rewarding, by the way, long term for understanding intelligence in general, like not just overseeing, but just understanding what intelligence is, how it works, what's the difference between the smarter model and the less smart model.

AI assessment note: “Yeah, I think like, yeah, exactly. I think as, as humans”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you have a point of view on the phased rollout or controlled rollout with Glasswing and Daybreak from, from Ant and OpenAI in this area?

A I don't have a strong opinion, but I think it's, uh, On the one hand, like, if we knew that there's not going to be anyone who's going to release a Mythos-level model soon, I think that would be great, because it gives enough time for everyone to prepare, to build the know-how, to build the playbooks, to share that around in the community, and to make sure that we're not starting to see airlines go down and power plants go down, and really, like, disastrous effects that could happen. The problem is that if anyone gets to a mythos level model earlier, then in retrospect, it would look like a huge mistake because we could have at least given companies the choice to start moving very quickly and give more companies access to mythos. Now they're all vulnerable because, you know, there's a Chinese model that's mythos level and there's nothing they can do about it. So I think hopefully we'll manage to do the gradual rollout correctly. I would really encourage that we expand The amount of companies that get access to this and make it much easier for people to get. I would advise everyone to assume that these models are coming anyway. The only thing you can do right now is to invest in these foundational controls that will stop the downstream effects of these vulnerabilities are going to be found in their systems.

AI assessment note: “I don't have a strong opinion, but I think it's, uh, On the one hand”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q labs could ever do? Or do you think it's a structural thing that I, I ask because the number one question amongst the startup ecosystem in the Bay area today is, you know, if you assume capability improves or, you know, when the labs just gets hungrier from their already currently ambitious stance, Uh, why wouldn't they do this too? And, and so I, I ask you the same question.

A Today, if you're, if you're a private person or if you're a security buyer, there are some places where you don't want to trust the same person that you're buying it from. So, you know, maybe, you know, if you're buying a car, you're not going to have the same guy that you're buying it from certified that the car is good, right? And maybe you're going to have someone else do it. And if you're a security, You're not going to trust the vendor of a product to tell you that this product is not going to mess your environment. You're going to want to have an independent party whose whole business depends on telling you that this thing is correct and being right. This, this thing is legitimate and being right. So that's like, there's the buyer psychology in this space that I think really goes in our favor. And then I think there's the core problems, like why are models even making mistakes? Why are agents even making mistakes? Right? So I would broadly Categorize it into two things. One is You know, there's the jagged intelligence of these models and there's like sometimes kind of very silly mistakes that they make. And I think that problem will go away. I think we're heading for much smarter models that make less silly mistakes and, and our role is not going to be to prevent silly mistakes. That will be taken care of by the, the model vendors because they're very incentivized to do i…

AI assessment note: “You're going to want to have an independent party whose whole business depends on telling”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Part of the solution for Onyx has been training its own models. Like what can you say about that?

A If you, if you try today, let's say you were trying to build a solution to oversee and kind of control how other agents are operating. Maybe the first thing a lot of our listeners might think is say, well, I'll just ask Cloud Code to do it. And, and in a sense, they would be right because Cloud Code is great. And maybe we can ask it to Spawn a version of itself for every agent that we have and kind of keep monitoring everything that agent is trying to do. And if you think that there's a problem, um, intervene. So that approach, it has, obviously it's pretty naive and there are some ways in which it totally fails, we could talk about, but it has some merit to it, right? So it does seem intuitive that it's a good idea to have Uh, capable agents reviewing what other agents are doing. Same as we have capable humans reviewing what other humans are doing. Right. But then the problems that you're going to run into is how do I make this work from, uh, uh, cost latency and reliability perspective? Because if I need to run an agent for every agent you're running as your security vendor, uh, you're going to be paying for me More than you're paying for your AI, right? So it's not, it's pretty much a deal breaker. And also it's going to be so slow. So you're not going to be happy with whatever latency you're going to get. And so the challenge then becomes, how do I know what are the times w…

AI assessment note: “how do I make this work from, uh, uh, cost latency and reliability perspective?”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.