Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Uh, it was actually, it's a, it's a, you can ask any Googler, it's like, just like a thing that, that, that, I mean, look like, yeah, limited resources, you gotta have some kind of marketplace, right? You know, you could, sometimes it's explicit, sometimes it's in, you know, just political favors.
A You could. And so then, like, basically everyone's assigned a credit, right? So if you have a credit, you get to buy, you get to buy n chips according to supply and demand. So if you want to go do a giant job, you have to convince, like, 19 or 20 of your colleagues not to do work. And if that's how it works, it's like, it's, it's, it's really hard to get that bottom-up critical mass to go scale these things. And like, um, and the team at Google were fighting valiantly, but like, we were able to beat them simply because we We took big swings and we focused. And I think again, that's like part of the narrative of like this phase one of AI, right? Of like this modern AI era to phase two. And I think in the same way, I think phase three companies can out execute phase two companies because of the same, like, like asymmetry of success.
AI assessment note: “basically everyone's assigned a credit, right? So if you have a credit, you get to buy”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q is LIDAR versus camera. Um, and I feel like Most agent companies that I'm tracking are all moving towards camera approach, which is like the multimodal approach that we're doing. You know, multimodal vision, very, very heavy vision. All the Fuyu stuff that you're doing, you're focusing on that, uh, including charts and tables and, um, do you find, like, inspiration there from, like, uh, the, the, the self-driving world?
A That's a good question. I think sometimes the most useful inspiration I've found from self-driving is The, um, is the levels analogy. Uh, and I think that's great. I think that's awesome. Um, but I think that, um, our number one goal is for agents not to look like self-driving in that we want to minimize the chances that agents are sort of a thing that you just, uh, have to bang your head at for a long time to get to like two discontinuous milestones, which is basically what's happened in self-driving. Um, we want to be living in a world where you have the data flywheel immediately. And that takes you all the way up to the top. But similarly, I mean, like compared to self-driving, like two things that people really undervalue is like really easy to get the, like driving a car down highway one-on-one in a sunny day demo, right? Like that actually doesn't prove anything anymore. And I think the second thing is that, um, as, as a non-self-driving expert, I think one of the things that, um, that we believe really strongly is that Um, everyone under values the importance of really good sensors and actuators. And actually a lot of what's helped us get a lot of reliability is like a really strong focus on like, actually, why does the model not do this thing? And the non-trivial amount of time, the time the model doesn't actually do the thing is because if you're a wizard of Oz-ing it …
AI assessment note: “most useful inspiration I've found from self-driving is The, um, is the levels analogy”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q readiness scale, because you have reliability, you also have cost, you also have speed. Speed is a huge emphasis for that. Um, All of that seems to, tends towards wanting to reduce, the tendency or the temptation is to reduce generality to improve reliability and to improve cost, improve speed. Um, do you perceive a trade-off? Do you have any insights that, that solve those trade-offs for, for you guys?
A There's definitely a trade-off, um, if you're at the Prado frontier. Um, I think a lot of folks aren't actually at the Prado frontier. Um, and I think the way you get there is basically like, How do you frame the fundamental agent problem in a way that just continues to benefit from data? And, um, and I think that, um, I think like one of the main ways of like being able to solve that particular trade-off is like you, uh, you basically just want to formulate the problem such that every particular use case just looks like you collecting more data to go make that use case possible. I think that's how you really solve it. Then you get into the other problems like, ah, okay, are you overfitting on these end use cases? Right. But like, you're not Doing a thing where you're like being super prescriptive for the end steps and, uh, that the model, uh, that the model can only do, for example.
AI assessment note: “There's definitely a trade-off, um, if you're at the Prado frontier.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q So you don't consider yourself a pure play foundation model company?
A No, because if we were a pure play foundation model company, we would be training, um, General foundation models that do summarization and all, and, uh, and, uh, dedicated towards the agent. Yeah. And our business is an agent business. We're not here to sell you tokens. Right. And I think like, um, selling tokens, um, uh, unless there's like, yeah, I love it. It was like, if you have a, if you have a particular area of specialty, right. Then, um, you won't get caught in the fact that like, everyone's just scaling to ridiculous, ridiculous levels of compute. But if you don't have a specialty, I find that I, I, I think it's gonna be a little tougher.
AI assessment note: “No, because if we were a pure play foundation model company”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q to do and the software just does it for you. When you think about these use cases, do the users still go in and like Look at the agent kind of like doing the things and can intervene or like, are these like fully removed from them? Like the truck thing is like, that's the truck just show up or like, are there people in the middle, like checking in?
A Yeah. So actually what's been really interesting is you could question whether they're funneled, but I think there's two current flaws in the framing for services as software, or I think what you just said, I think that one of them is like in our experience, as we've been rolling out Adept, um, the people who actually do the jobs are the most excited about it. Because they don't go from, I do this job to, I don't do this job. They go from, I do this job for everything, including the shitty wrote stuff to I'm a supervisor. And I literally like, it's pretty magical when you, when you watch the thing being used, because like now it parallelizes a bunch of the things that you were, you had to do sequentially by hand as a human. And you can just click into any one of them, be like, Hey, I want to watch the trajectory that the, the, the agent went through to go solve this. And, um, the nice thing about agent execution as opposed to like LLM generations is that Um, a good chunk of the time when the agent fails to execute, it doesn't give you the wrong result. It just fails to execute and the whole trajectory is just broken and dead and the agent knows it, right? So then those are the ones that the human then goes and solves. And so then they become a troubleshooter. They work on the more challenging stuff. They get way, way more stuff done and they're really excited about it. I think …
AI assessment note: “They go from, I do this job for everything... to I'm a supervisor.”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q You mentioned Dota. Any memories thinking from, like, the switch from RL to Transformers at the time and kind of how the, the industry was, um, evolving more in the LLM side and leaving behind some of the more agent simulation, uh, work?
A You know, I actually think that people, like, zooming way out, I think agents are just absolutely the correct long-term direction, right? You just go to find what AGI is, right? You're like, hey, like, Well, first off, actually, I don't love AGI definitions that involve human replacement because I don't think that's actually how it's going to happen. I think even this definition of, like, AGI is something that outperforms humans at economically valuable tasks is, like, is kind of a, you know, like, uh, implicit, like, view of the world about, like, how, what, what are, what's going to be the role of people. I think, um, I think what I'm more interested in is, like, a definition of AGI that's oriented around, like, a model that can do anything a human can do on a computer. Um, and I think, like, if you go think about that, which is, like, super tractable, Then like agent is like just a natural consequence of that definition. And so like, like what, what did all the work we did on our own stuff like that get us was it got us a really clear formulation. Like you have a goal and you want to maximize the goal and you want to maximize reward, right? Like natural LLM formulation doesn't come with that out of the box, right? So like, I think that we, um, as a field got a lot right by thinking about, Hey, how do we solve problems of that caliber? And then the thing we forgot is like, Li…
AI assessment note: “what did all the work we did on our own stuff like that get us”