The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Howie Liu no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Um, and so what we can do. Like, we would put this in the eval, right?

A Yeah. You could do both. So one is, like, you could immediately go and turn this, like, or update the skill based on this feedback. You could also have it immediately just, like, turn around, like, a new draft of these tweets, right, to sound more colloquial. And then finally, to your point, I could go and create a rubric that actually says, like, okay, like, here's the five dimensions I care about, and then auto-evaluate every future output, right? Um, so, You kind of have a number of different options, like, depending on how far you want to go right now. Like, if you just want to get your job done right now, you don't want to bother with rubric, you don't have to, right? But eventually, like, you get to the point where you want to set up a scalable system for this to just constantly work and get better and better, and that's the point at which you would do a rubric, which is not that hard, actually. Like, it, you know, you can either go in through the UI and build one, or you can actually, in this chat, like, say, help me build a rubric to score great Greg-style content, which I'll queue up for after It updates the skill, um, and, and then it will go and help me create that rubric, save it, pin it to this agent or to this skill, and then automatically run every future, um, uh, time I create content.

AI assessment note: “to your point, I could go and create a rubric that actually says”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Howie, some people are like, they're, they're, you know, um, they're, what they love doing is, like, obsessing over tuning every single detail and stuff like that, and those people, you know, That, an open claw might be for them, right? But if you're, if you want more.

A Like, yeah. That, but I believe that, like, you don't have to sacrifice the tunability, right? Or the, like, the power. And so, you know, one of our strong design philosophies here is that, like, hyperagent still does give you a lot of control. Like, you can go and tweak, you know, kind of, like, agent configuration if you want to. If you want to, like, choose the exact model and system prompt and tools and, like, give it a lot of refinement, you can. And, Like, you can go quite far in terms of curating memories. Uh, we actually just shipped yesterday a kind of, like, a defrag tool for your memory, so that as you accumulate more and more memories across all these different agents, you have this, like, really elegant way of, like, defragging them, right? Where, like, we can auto-suggest here related memories clustered by both, like, you know, keyword as well as, like, embeddings similarity, so that we're actually understanding the content of the memories, and you can consolidate them, but They're like, you know, we want to really serve both people who are like power users who want control over how the agent is set up so they can get maximum bleeding edge performance, but then also, you know, like you shouldn't have to do all that to get value out of the product. So it really is about the range. I think it's more just that like if you are truly, you know, happy just like doing it…

AI assessment note: “I believe that, like, you don't have to sacrifice the tunability”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q be honest. Yeah, exactly. And I've, I just need your reaction to some, to just some things I've been thinking about. Yeah. So this chart over here, it's by Sequoia. In what domains are AI agents deployed? You can see software engineering is at almost 50%, back office at nine percent, marketing and copywriting four percent, sales and CRM 4.3%, and down. When you see this, like, what's your reaction?

A I mean, I think Two things. One is, I think it absolutely reflects the under penetration of AI in industries that clearly could already be disrupted or benefit with even today's AI capabilities, right? If you took like frontier agents today and deployed them into every one of these categories, you should get to a hundred percent. And then two, I think even the higher numbers like software engineering is actually kind of an overestimate Meaning, you know, like as I think frontier developers and companies applying frontier agentic development practices are finding like, you know, the new model of software development is not even just like every engineer using AI autocomplete, like tab autocomplete, which like we all figured out like three years ago, right? Uh, with even GitHub copilot, but it's now like, you don't even need the IDE, right? Like the, the way I develop on hyper agent is I have like 30 different cloud code instances running in parallel, and each one is coupled up to like a browser, Fully autonomous. It can go and like get other agents to comment on any PRs it creates. And so like this modality shift of like, you know, no AI to like kind of what I would call gen one AI, which is like basically like AI augmentation for still like very human driven development workflows. Andre Karpathy talked about like, you know, in October, November is when he completely inverted fro…

AI assessment note: “Two things. One is, I think it absolutely reflects the under penetration of AI”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Um, two more quick graphs, and then I want to get into, uh, Hyperagent. Um, percent of enterprise apps with embedded AI agents, um, you know, this is the fastest adoption curve in enterprise history, right? So like, when you see this, you know, how do you react?

A I am not surprised. And I think even this reflects the pace at which like incumbents can even like integrate AI into their products. Right. And I think even that is like stimied by like just incumbency and like, you know, kind of how, how seriously did enterprises, um, you know, uh, enterprise apps or enterprise app makers or internal app teams like take this. I think the real show of how profound This growth curve is, is like, if you take the aggregate revenue created from, from zero of all the leading AI companies, right? Or companies like doing AI things, like take opening Ionanthropic alone, right? Let's just say they have a combined revenue probably of like eighty million, uh, plus, right? Or eighty billion, sorry, plus right now up from like basically zero a few years ago. Like what in, in the history of software, like has there ever been An industry where, like, any company, let alone, like, or even an aggregate, like, you know, across all the companies, you got a category that went from zero to, like, you know, eighty billion plus, right? And that's not even including, like, all of the other AI providers, inference, uh, inference providers, and, like, you know, tooling, et cetera, like, out there. Like, the, the revenue of, like, I think the AI category is an even sharper curve, and I think that really reflects, like, just how profound this lightning in a bottle is.

AI assessment note: “I am not surprised. And I think even this reflects the pace”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q who is the judge around the output? The judge is you, a human being, right? It's not opus 4.6. It's a human being. So if you're trying to actually create what we were talking about before, which is like an agent first business, you know, managing a ton of agents, you're realistically, you're not going to have the bandwidth to be looking at every single output at all stages, right?

A Yeah. It's kind of like management one-on-one, right? But like applied to agents now where it's like, as you scale up, if you're the CEO of a business, like you just literally don't have time to go and like look at every single thing that every single person in the company has done. And so you need to create like better automated checks and balances to oversee what the agents are doing. Right. And like inspect quality of work, right? Like this would be like, if you actually had like a giant army of human content creators, like, You would want some way of like, you know, um, in a scalable way, like to detect, like if they're posting good or bad content or not. Right. And then know like, okay, we got to tweak like the guidelines for each of these people. Okay, so now we have the Greg Eisenberg contrarian draft skill, um, and I'm gonna go ahead and save this skill, and I'm gonna try seeing, like, okay, let's do a dry run. It's gonna scan today's AI and, uh, news and trends, and then create some contrarian drafts, right? And the whole idea here is, like, look, like, it's probably gonna do an okay job on, like, the, the first effort here. Like, it did some research about you. It kind of, like, you know, it has a lot of, like, um, context about how you work, right? If I wanted to see more about this, uh, skill, I could actually open it up. Here's, uh, when it should be used for. Here…

AI assessment note: “you need to create like better automated checks and balances to oversee what the agents”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q I'm assuming, well, do you have linear built in here?

A Uh, we do have a connection to linear, but actually, maybe, maybe Twilio could be a good example, right? Like where, I don't think you can OAuth into Twilio, so it has to be an API skill, and we may have a pre-built connector, but I'm going to have it, like, build a custom skill, uh, regardless, so can you still help me build a custom skill to integrate with Twilio via API, right? And Um, so now what it's gonna hap- what's gonna happen is like, it can go and like, research the Twilio API docs, create a skill for itself to use the API, and then actually ask me to enter my credentials in a safe way, uh, and then be able to like, use the Twilio API fully, right? Um, So I think, like, the powerful thing now is a frontier agent should be able to, like, literally do anything, right? Like, but it's just a matter of, like, you have to give it access to the right context, and you have to, like, you know, tell it, like, hey, like, yeah, you should build a skill for this, so then it can do it every single future time effortlessly. Um, let's say what we want to do, SMS, voice, for now, maybe phone numbers, we'll do an API key, auth, and, um, any specific workflows, think like, maybe, actually I want to build a voice and SMS service that can call restaurants for reservations or something, right?

AI assessment note: “we do have a connection to linear, but actually, maybe, maybe Twilio could be”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.