The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

640exchanges match
640on raw tape
31redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q get to open frontiers, but one more question about just the NSF, which I think you, uh, you mentioned once for, for about Databricks, but also I think you're trying to target the LOD grants as sort of NSF level prestige and, and, and, and impact. I think one thing I, maybe it's like a spicy question is, well, what's broken about the NSF process are you trying to fix?

A Uh, I don't think NSF is broken. I love NSF. It has been probably the best investment the American popular public has ever made. Thousands X return on their investments. I mentioned Google earlier, Databricks, all these researchers came from NSF funding, and it's just been A paradigm generation. So DARPA was important for, at some point, there's been these, these paradigm shifts in how open research has been funded. DARPA was a key one. NSF's been a key one. And now NSF is not, is not big enough. It was one billion dollars a year for computer science and that they're trying to cut that into half of that, but we need 10 to a hundred billion dollars to do frontier AI research. So it's not broken. They are trying to break it. You need more is insufficient.

AI assessment note: “I don't think NSF is broken. I love NSF.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Was there a moment for you where you were, uh, I'm sure you were debating it yourself. You had other opportunities. What was the deciding factor for you?

A It became clear that the only way to scale what we were building was to build a company out of it. That the world really needed something like Arena. Arena being really a place to sort of measure, understand, and, and advance the frontier AI capabilities in, on real world users, on real world usage. Based on organic feedback and that in order to achieve the scale and, you know, distribution necessary and the quality of course of the platform necessary to do this effectively, we would need to start a company out of it. You know, we considered other options. Are we going to keep doing this as an academic project? Are we going to do it as a nonprofit? Blah, blah, blah. But ultimately under those constructs, we didn't feel like we'd have the resources necessary to accomplish our mission.

AI assessment note: “It became clear that the only way to scale what we were building was to build a company”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Well, I think a lot of people are trying to place MCP versus skills. Obviously they're not overlapping, but how do you view it?

A Yeah, I agree. I, I think that's the interesting part is like, they're not overlapping. I think they, they solve different things. I think skills are super great. And you know, they're, I think that the first that really like, they've been built from the principle is progressive discovery. But I think the mechanism of progressive discovery, that's just universal to any type of thing you can do with the model. But what skills do, they like, they give you the domain knowledge for like a specific Set of tasks, like how you are, how you behave, how should the model behave as a data scientist, or how should the model this, um, behave as an accountant or whatever. But MCP gives you the connectiveness of the actual actions that you can take with the outside world. And so I think they're somewhat, um, orthogonal in like, in terms of like the skills really gives you this domain knowledge, just like kind of vertical. And then like MCP gives you this horizontal of like, okay, you know, Give me that one action. And of course, skills can take actions. They can take actions because you can have code and scripts in there. And that's great, but it has two interesting aspects that I think people got. The first one is you need an execution environment. So you need to, you need to use a way to execute your machine. Yeah. Yes. And that, that's, that's perfectly fine for, you know, if you like run …

AI assessment note: “skills really gives you this domain knowledge... MCP gives you the connectiveness of the actual actions”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And on, on the opposite, what's my incentive to bring my project to? I have a, uh, project with adoption. It's well maintained. It's healthy. What's the benefit that I get from, um, donating to the foundation?

A I mean, I can start that, but I'd love to hear from these guys as well. I think what you, all technology is an implicit futures contract, right? And so, you know, if there's technology that has traction and, uh, that traction sort of wants to be built upon, having that technology at a neutral place, like the Agentec AI Foundation, where the whole industry is making decisions about How to invest. And when I mean, when I say investment, I don't mean like becoming a member of the foundation, because you don't need to become a member to participate on the technical side. It's decisions about, hey, I'm gonna, you know, assign 10 of my company's engineers to co-develop this with your organization, the contributing organization. And, you know, the, that's a way that we can all essentially co-develop together, and that will provide Better support, more development velocity, higher code quality because more people are participating in it. And that's a massive incentive if, you know, you want your technology to actually be used and adopted in industry and get more feedback and kind of a, a positive feedback loop of great project that gets great products in the market. That market feedback then, you know, allows companies to make money off of them. They then pay engineers to improve the project. Better products, more profits, better project. And, you know, that's the incentive, which is p…

AI assessment note: “Better support, more development velocity, higher code quality because more people are participating in it.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. I noticed, you know, you, you also did the thing to me where you just start the conversation, right? You don't have a, well, here's the intro. Here's your birth story. Here's your origin story. Uh, which I try to do sequentially a little bit, but that's one of your tricks, right?

A Well, I would say I go through An extreme level of detail to make sure that the guest feels very comfortable when they sit down. So one example of that is, you know, how you're greeted at the door, water, all those things. The second is recording just starts. There is no like, okay, are you ready? Because the minute that somebody says like, okay, are you ready? Go. You claim up. You're like, okay, I'm going to be the guest that I want to be. It's like, you know, when you're sleeping at night before you go into a podcast, you're like, okay, how am I going to sound? What am I going to say? That's going to make me feel smart. You know what I mean? Make me sound smart. And so you start to build this like idealized version of yourself that you want to project to the world, which is like not real. And so start talking as soon as you sit down the temperature of the room. Like I like the temperature to be cold. I don't want people to feel like they're sweating or hot. It feels like kind of cool in here, right? Uh, the way that the lights are, you'll notice the lights are all up, not down. Like I, I think it bounces off. Yeah. I think it's important to not make it feel spotlighty. Yeah. If that makes sense.

AI assessment note: “The second is recording just starts. There is no like, okay, are you ready?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You have signing. Why do you need 3000 engineers?

A It sounds crazy, but like you want to nail Europe. You need a different product. You need a different team to your local data centers because of the compliance. You cannot just run your data centers from the U.S. So you need a local team there. Oh, and by the way, the way to do digital signature in Europe, totally different. So like the stack itself is different. So like the way to make a digital signature is different. Not the same standards in the same ways. So you need dedicated team to maintain that thing. The same way some people want to have DocuSign on-prem. So you need a team building appliance to basically plug and play and like, okay, you have your DocuSign appliance.

AI assessment note: “You need a different product. You need a different team”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q How did you design the tools to get to the agent? Just maybe give people an overview of, like, the framework, what it looks like. Like, how are you structuring these interactions? Is there just one superhuman agent that does everything, or like, do you have separate ones?

A We have separated tools, uh, clearly. So, uh, even an agent, like, I would call about, I would say tools. So there's a bunch of tools. Tools to detect your availability. Tools to understand who are the people you interact with. A tool to write an email. The tool to, like, so every single action is very tool specific, so it's not a magic B tool that can do pretty much everything. It's a set of small tools that are used, uh, within the Argentiq framework. So, like, there's a first step that is like, hey, what is the best tool to do this? Kind of like building a plan, like, for each step, what is the tool, and then making the calls.

AI assessment note: “every single action is very tool specific, so it's not a magic B tool”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I think like, I think that the classic thing is, well, what is a unit of output of a software engineer? Is it a PR? Is it a story point? It's extremely unclear and it's very basically unsolved. Like, I mean, don't tell me you've solved it. You know, like what's, Maybe you have, I don't know, but I'm, I'm default skeptical on the, well, what gets measured gets gained.

A Yeah. We do use story points, but you're right that it's, it's easy to game it, right? Like, like if we were to hire somebody who just, like, if you think about a technical system, right? A smart hacker will find ways to exploit it. And the easy way to exploit the story point system is to deflate the concept of the story point and decide that, okay, Any line of code, like lines of code are going to be directly proportional and equal to story points. Well, then of course you've hacked the system, right? But your clients will churn and you'll probably get let go of, and it just won't work long-term. And so what we found is that hiring two, what we found is that this problem gets solved in the hiring process and it gets solved by hiring people who fall into two buckets. One is people who are selfish, but they're long-term selfish. Everybody's selfish, but we need to look for people who are long-term selfish, people who understand that these incentives are longer than just today's story points. They're forever, right? And we need to think about how do we maintain the client relationship. And that means that we're going to give them very robust story points so that we can maintain the relationship and continue to make money. But the other is that we hire people who just like writing code and like working with really smart people and, and they're not sharp elbowed and they just want …

AI assessment note: “We do use story points, but you're right that it's, it's easy to game it”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I mean, well, so yeah, but you're gonna, it's very anecdotal, right? Like, don't you need more comprehensive evals? Because otherwise it's like you are just behaving or believing things based on the luck of the draw.

A I think at this state, did a samurai have a measurably better sword than the person to their left or right? No, right? At a certain point, I think a, a, a, a warrior's weapon becomes something of a feel. And I think that at this point, a lot of these, like, The coding agents are so good. Like, yes, you can have evals that, that provably show that one is better than the other. But for a lot of these things, it really is feel it's like, Hey, this agent actually, like it just, I can work better with it on a warm blooded level, or it writes code more like I like to, or whatever. And at least that's what we've noticed.

AI assessment note: “yes, you can have evals... But for a lot of these things, it really is feel”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q so since then, I wanted to start with Glean, obviously, because of, you know, we, uh, we're going to cover a lot of startups in this episode. So Glean has, Glean was like a billion dollars, I think, as based on my research, and now it's at seven billion dollars. So your, your, your options are good. What's your take on, like, how Glean's going and the market in general?

A I would say that Now being on venture side, I have a bit of a different take than I would have had at Glean. But broadly, one of the things that I love about Glean is it's such a boring, unsexy company that became sexy later. So from 2019, I remember going to parties in the Bay Area, and I would say enterprise search, and it's a shutting down the conversation right there. You know, like nobody would ever ask a counter question if you said enterprise search. They're like, oh, that sounds boring as hell. Leave me alone. Like, um, and, and fast forward to 2022, Enterprise search gets more, um, got more conversations. It was like, interesting. Tell me how you're doing this, this search. I think what was nice about that observation is in those three years, we did a lot of work and not, didn't take shortcuts on a lot of things that ended up generating a lot of value for us now. And I can go into what, what all of those things are, but if you look at glean from a high level business, it is top down enterprise sales. It's very hard to rip and replace. We have, we expand contracts very easily because the TAM is so large. It's every knowledge worker could use a version of enterprise search and then the AI on top. I still call it search, but information retrieval in the enterprise. And we've, we've, we solved a lot of critical problems. I can go into that too, in order to get there. Then …

AI assessment note: “if you look at glean from a high level business, it is top down enterprise sales.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Just a question on that. Was there any, because you know, oh, you have a new search tools, like go search. And it's like, what am I searching? You know, like what was that blank canvas onboarding for people?

A Um, several different things worked well for us. Uh, I can think of two at the moment, but I'm sure there were many, many more. Um, I'll say one of them was say for, for a handful of companies, like many companies, actually, we would say, we want to take over your new tab page. And then the critical part was tell us what we need to do to earn the right to do that. No one wants to give away their new tab page. So, so, so we, Went the last mile. There were companies were like, well, we have a new tab page. We're pretty happy with it. So we'd ask, do you have a search bar on it? And they'd be like, well, yes. I'm like, okay, what is, what is that using? And they'd be like, well, it's using our internal thing. I'm like, do you like it? Clearly not. That's why you're referring to us. So let's just rip and replace that. But doing that extra mile was pretty important. So that's one new tab. The second one that we liked was a Chrome extension and then doing the, I forget what we call this, but When you were on your native product and you were issuing a search query, we ran a lot of evals and we thought we were better at every product at their own search. So if you were searching on Google Drive, we will do a Glean replace of the search bar and the page pretty natively. And it would teach people, it would teach people to use Glean and be like, okay, that's pretty useful. I think these r…

AI assessment note: “one of them was say for, for a handful of companies... we want to take over your new tab page”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q is like, for me, it's like a survey episode of like, here's everything. We're also catching up with the former guests. It's always nice. Maybe we can end it on this like coding interview thing, uh, which, which literally you tweeted about today. What is the situation that, you know, I guess engineers should be aware of? And I think this like maybe ties into LLM psychosis a little bit.

A You know, like, so I tweeted, I'll just cover the tweet first. I tweeted about this, um, Guy who wrote a blog post about, he was in an interview from a, I didn't think it was a legit account. He thought it was a legit LinkedIn message where he was interviewing for the company. They sent him a coding interview. They said, clone this repo, run this code, make this edit. Kind of not untraditional. So it's pretty, pretty run of the mill type interview. It happens. And in that interview, he claims that he went to cursor and asked whether the code had anything Any vulnerabilities or anything you should be aware of. And it revealed that it had some link. They had a byte array that compiled into a link that would go and take a bunch of private information from you. So that was the TLDR. And, and I tweeted about that saying, you know, like, The, the world, interestingly enough, it was solved by vibe coding, but it could very easily, the world of vibe coders who don't really look at code, I imagine are more susceptible to being in attacks like this and in the future. And, uh, and it got me thinking about a lot of things like, what is, what do attack vectors even look like if people aren't looking at code? There's so much that can go wrong. And what are the implications on model safety and how models behave in those environments? So that's one. But I think the broader thing, and I'm curio…

AI assessment note: “I tweeted about this, um, Guy who wrote a blog post about”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. Not, not to digress too far. I wanted to, to obviously go on to Discovery Engine. So I have your slides here. Why don't you tell us a little bit about what we're looking at?

A Yeah. So this, this is a screenshot of Discovery Engine. What we've done is, is like, take this insight that is something like Deep neural networks find patterns in data that humans miss, and quite often those patterns are novel, they are new to us, we can use this to learn a new thing about the world, and we have automated it. So Discovery Engine is an end-to-end system, takes in arbitrary scientific data set, automatically trains a bunch of neural networks on it, and then We systematically, with our interpretability methods, which is the real secret, um, extract the patterns that have been learned by those models, and then we contextualize them. With existing literature, we, we rank them by novelty and prevalence in the data, stuff like that. And we make them human possible. And, uh, we made a bunch of novel scientific discoveries doing this.

AI assessment note: “Discovery Engine is an end-to-end system, takes in arbitrary scientific data set”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q lot of internal tooling and that's great, but also there's a lot of great tooling out there. How do you navigate this? Obviously you have a lot of unique internal context, but you also have You work with external vendors. Like people, I guess, want to know how to work with you, but also people in your shoes at peer companies also want to know how you do this decision.

A Yeah. I think for us, it's not an either or. It's very much an and when it comes to build versus buy. And some of that and is sequential, right? So You, you and I were talking earlier about, like, when, when GPT, 3.5, I think, like, first hit the scene where, like, oh, everybody at Stripe needs to have access to LLMs, but, like, we don't quite know how to do that in a way that's, like, enterprise grade, safe, and we feel good about. Like, we don't see a provider there right now, and so we built it. But now we use, like, open source, LibraChat, you know, so I think there's, like, there's, like, an evolution over time. And one of the things that I think can be really hard, especially for the team who has tunnel vision for the products they own, You know, you love your product. You want to make it better over time is you can get stuck in a lot of hill climbing. Like we could have taken GoLM and been like, oh, we should figure out a way to like give it access to tool shed. Oh, we should figure out a way to make it do like deep research or, you know, like we, we, we could have done that. Um, and sometimes you just need like sort of more of an outside in perspective of, hey, if I ignore the sunk cost fallacy, ignore my emotional connection to the thing that I spent nights and weekends building, First principles, like if I were to do this today, what would I do? And some of that is al…

AI assessment note: “it's not an either or. It's very much an and when it comes to build versus buy.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Can you double click on why it's hard to do the sandboxing? Because in principle, we just capture all the inputs.

A Yeah. Well, you don't need to just capture all the inputs. You need, you need a system that reacts the same way your production system does. That's, and in many different ways. And, um, so let's say you're, you're Airbnb, right? And I'm bringing this up because this is like an example of one that like, you know, companies have gone out and built sandboxes. Like if you're Airbnb and you're trying to, Um, you want to train an agent to, like, maybe you're not Airbnb, fine. You're, you're a company like us that's trying to train an agent to, like, do really well at operating Airbnb and booking on your behalf, right? Like, you have to build a copy of the Airbnb website that reacts to you as the user the exact same way that the real one does with the same failure modes, right? Because if you don't include the same failure modes and bugs they have, then, like, one of those, when one of those bugs comes up in production, your agent's gonna have no idea what to do with it. It's just gonna fall over. You also need to simulate if this is like a sort of cooperative agent, right, where it's getting human input as well and kind of like working with the human to get something done, which in practice is the way a lot of these are deployed. You also need to simulate the user. And I mean, you can do the naive thing and just say, oh, we're going to have a separate LLM that, you know, with a syste…

AI assessment note: “You need a system that reacts the same way your production system does.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Um, maybe you want to introduce Agent Kit to the audience?

A Yeah, so we launched Agent Kit today. Um, full set of solutions to build, deploy, and optimize agents. Um, I think a lot of this comes from working with API customers and realizing how hard it actually is to take, to build agents and then actually take them into production. Hard to get kind of that confidence and the iterative loop and writing prompts, optimizing them, writing evals. Um, all takes a lot of expertise, and so kind of taking those learnings and packaging them into a set of tools that makes it a lot easier and kind of intuitive to know what you need to do. Um, and so there's a few different building blocks that can be used independently, but they're kind of stronger together because you then get the whole end to end system and, um, releasing that today for people to try out and see what they build.

AI assessment note: “full set of solutions to build, deploy, and optimize agents.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What's the entry point? So are developers also supposed to come here and then do the two code, um, export, like, just segment, like, the use cases?

A Yeah, I mean, so I think, like, the two reasons that you would come to Agent Builder are one, um, kind of more as a, as a playground, right, to kind of model and iterate on, um, your systems and write your prompts and optimize them and test them out, and then you can export it and run it in your own systems using Agents SDK, using kind of You know, other models as well. Um, the second would be kind of to get all of the benefits of us deploying that for you too. So you can kind of use maybe like natural language to describe what type of agent you want to build, um, model it out, bring in subject matter experts so that you really have this canvas for iterating on it and getting feedback, you know, building data sets and kind of getting feedback from those subject matter experts as well, and then being able to deploy it all without needing to, to handle that on your, on your own. And, um, that's a lot of the philosophy around How we're building it with chat kit as well, right? You can kind of take pieces of it. You can have a more advanced integration where it's much more customized. Um, but you also get, um, a really natural path of going live, um, without, like, with really kind of easy defaults as well. Um, yeah.

AI assessment note: “the two reasons that you would come to Agent Builder are one”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q the causal ML work background that you guys have. So can, can you maybe now show the forks? So there's kind of like the way that a lot of companies does it, which is like, Hey, you, you build the context. Now you ask the LLM to figure out why this was happening. What are, like, some of the things that you're doing differently, um, that you can talk about?

A So I'll note a few things I'll let Roz also, I guess. Uh, there's Building the context itself is where a huge amount of the, the, the pain is, right? Because you're dealing with, let's say petabytes of data, like let's say a billion possible symptoms, and you have to get it down to a small number of symptoms. And that entire, so it's this like massive data processing pipeline where you're reducing, we're getting more and more relevant context over time. And there's LLMs throughout that process. And each of those LLMs, there's basically AI agents throughout the process helping to, to, um, winnow down the The relevant amount of information and the way it's windowing down that information is by running these kinds of statistical tests that we have built in a proprietary way, because most of the data you're looking at is like time series data. And these LLMs are really bad at, at processing time series data, right? And that's really where like good statistics comes in. So those are basically like statistical tests of the toolkit that the AI agent has access to. And like, almost like a data scientist is deciding, you know, uh, dynamically what is the next best Collection of statistical tests to run to filter down the search space and also use at the same time, the semantic information associated with the logs, the metrics, the traces, and so on and so forth. It's like a, it's like a…

AI assessment note: “the way it's windowing down that information is by running these kinds of statistical tests”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q it's like, what do you think is really unique about the approach you took that really resonated with people, especially when you don't really have a background that is like in this space, right? Like, I think like, you know, in a black box, you would imagine, okay, the founder of Traversal is probably somebody who ran DevOps or something at one of these big clouds or something like that.

A Yeah, I think it was two things. As Sean said, like fundamentally, this is a space of show not tell. And I think, um, we just had good references of, of customers actually using the product at scale and dealing with complex incidents. I think that the, the long term, uh, technical edge will be when you deal with complex incidents. And I'd say most of the companies that have been in the space definitely put a product out that works and people are using it, um, All the time. And also their use case has typically been like smaller, easy runbook automation for alerts versus like actually dealing with incidents where no one has any answer. And so I think the biggest thing was just showing and having, you know, being able to point to customers that are actually using it. And so I think that's probably the number one thing. And then I think the second part of it is explaining why this is such a hard AI problem, where it's not just like, oh, let's throw a chat GPT wrapper around your telemetry and, and something magical happens and why that's not possible. And so I think explaining that really carefully, and we can talk about what those key AI challenges are. So I think, I think it was a mix of those two things, which is the, you know, actually having it work. And I don't think we've seen any other company in our space having something actually work at production. And then second, a ve…

AI assessment note: “I think it was two things. As Sean said, like fundamentally, this is a space”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Does it feel like it's changing in a way the tools that you need to build? And how do you think about, yeah, you mentioned using AI for like, you know, editing the design and whatnot. Do you feel like natural language is becoming more and more the interface, even in design, that the work is going to be done? Or, yeah, what are like the pieces in your mind?

A Yeah, lots to unpack there. I'll start with just the, is natural language the interface? Yes, right now. I've said before, but I really believe it. I think we'll look back on this era as like the MS-DOS era of AI, and the prompting and natural language that everyone's doing today, I think is just sort of like the start of how we're going to create interfaces to explore it in space. So I'm just like, cannot wait for an explosion of creativity there, because I kind of think of these models as like, they're almost like a, uh, N-dimensional compass that lets you explore this, this wild unknown fog of war in laden space, and you can kind of push the models in different directions through natural language, but if you have a more constrained end there and you're able to dimensionality reduce a bit so you can push different ways, there should be other interfaces available than text. These might be more intuitive, but they also might be more fun to explore. And I think sometimes constraints unlock creativity in ways people don't expect. So I'm excited for that. But right now, yes, natural language is where we're at. And while I'm excited to push that forward, meet people where they are, I think is usually a good model for product development before you get to the point where you've really refined. Going back to your triad, I think, uh, maybe we're going to start with the spec. Like, I t…

AI assessment note: “I'll start with just the, is natural language the interface? Yes, right now.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q think some people should have some sort of, like, personal stigma, almost. Whereas, like, you know, when you're generating images on ChatGVT or MidJourney, you should have some sort of, like, aesthetics to, to draw from. How do you think about that evolution? Because you're going from a world where, like, only designers work on your product to now it becomes a core part of, like, a lot more constituents.

A Yeah, I, I think, um, in a world where design is the way you win, it's only natural that we need to get more people involved in the design process. That is not going to diminish, diminish the role of designers. In fact, I think it expands the role of designers because then you have to shepherd people through the design process and help them, uh, go from, okay, I mean, it's kind of a journey you go on, right? Like, not even being aware of design. You know, kind of, like, blindly going through the world to, oh man, aesthetics, uh, they, they matter. To, uh, okay, can we make it pop? Can we make it cool? And then people start to actually think about, oh, well, wait a second, like, I'm looking at one screen. What's the actual experience here? What is the entire flow? And then it's, okay, well, let's take mental models of how we can think about this experience And consider it in different ways. What are the potential different paths, metaphors, uh, and experiences that we can create here? And what are the abstractions that matter? And then from there, it's like, okay, wait a second. Well, this all exists in the context of like our brand, the greater culture of the moment, and, uh, business constraints and all sorts of other things you might be optimizing for. And I think, um, more people coming to the design process, that can help add in context as well. And there's no reason why so…

AI assessment note: “it expands the role of designers because then you have to shepherd people”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q have a code diff, you can give it to the model and the model reads the code and understands what's being changed. It's kind of hard for the model to understand the aesthetic change in a way. Like, how do you think about explaining to the model these things? Like, have you come up with like a good model of like taste semantics to put it in the latent space?

A I actually think Figma did a really good job with Make in that they, um, you can tell that there's much more of a bridge between the sort of underlying design model and what the LLMs are doing in kind of coordination between Sonnet, and that's, I think that's part of it, which is, um, Giving the model a sense of sort of the design building blocks is important rather than just the finished product. And I think you can reason about that as well. But the second part, and I think we need to make some strides here too, going forward is the models don't see as well as they could. Um, they see, okay, you know, you ask them analyze a complex photo and they're able to do it, but I want them to be as persnickety as a like really good visual designer. Like, no, that looks, the baseline looks a little off, you know, or this needs to be good. And I think That's going to come from sort of additional vision capabilities that, you know, we'll work on, but I think that's going to be a really key piece of, of closing that loop so that, you know, model, and I've seen people do this with cloud code and like take an MCP with Playwright, for example, in a browser and basically do the loop of, right, you generated the UI, now look at what you did, and is it right? And can you iterate on that? And I think the, is it right, still needs some, some visual help before it's, it's fully done.

AI assessment note: “Giving the model a sense of sort of the design building blocks is important”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What are those things? What are like the failure modes that you heard from customers where it's like, hey, we tried AMP and it just didn't work at doing X, Y, Z. Is there a collection of those that you guys use as almost like a North Star as you keep building or?

A But I think like one of the things is, um, the whole vibe coding stuff where people just use it and, you know, they're like, Hey, I spent 10 bucks in tokens and it didn't build me the fall app or something. Um, the failure mode of outsourcing the thinking, but not the typing, which I think it should be the opposite. You still have to know engineering. You still have to know how to program. You still have to know your application and its architecture, how it's deployed. And then basically use the agent to do the work that you would have done But you have to know what the desired outcome is and whatnot. Like that's a common one where people just, you know, hands off the wheel, agent, you go and write this for me. And then turns out a couple hours later, oh, actually this, nobody understands that it's spaghetti code.

AI assessment note: “The failure mode of outsourcing the thinking, but not the typing”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q whatnot. JetTPT obviously has this open AI. So if you give it a memory thing, Yeah, it might use this. But then you have the issue of, well, if I give it this other custom-made MCP that we built internally, and our processes don't map to anything that OpenAI and Anthropic have seen or trained for, it won't be used, and you won't get good results. And super strange, right?

A Yeah, I wrote this article for the GPT-V release about, um, models self-improving for coding. So I basically asked GPT-V What are tools that will be useful to you to be a better software engineer? It's like, well, you know, give a list of like 10 tools. And I'm like, okay, implement them, wrote all the tools. And then I asked it to do the same task I'd done before, but with those tools. And then it goes through the whole task and I'm like, which of the tools did you use? And it's like, oh, I didn't use any of them. And I'm like, why did you not? It's like, you know, to be honest, I don't really need the tools. I can just do this task, you know? And, and I think that's like a good, Metaphor just for, like, the trend of the models, which is like, hey, they're going to use less and less of this, like, custom-made tools to fix today's issue. I think the things that we can bet on, and I'm curious to hear your thoughts, is like, they're always going to have some sort of, like, test runtime. Like, I don't think there's going to be a world in which the model is not going to run tests and say, I'm sure this is going to work. The other one is there's always going to be some sort of, like, Infrastructure as code to then handle the deployment side. So I think whenever there's going to be some runtime issue, they're going to need to understand where they're running, you know, so I think lik…

AI assessment note: “they're going to use less and less of this, like, custom-made tools”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah, no, and I think, like, uh, what, you know, if I'm in your shoes, I'm like, man, these people don't even know what Anthropic is. Why would I go be the palantir for enterprises, you know?

A I honestly got off the call sometimes with these customers and just went, man, you can give me two engineers in, like, a week or two, and we could probably save you a million dollars, but it was just the, and it wasn't this company's fault, they're just not technical companies. You know, I bucketed these companies into tech companies and non-tech companies, and within non-tech, it's Are they technical or non-technical? You know, do they have good in-house engineering resources? So take like Goldman Sachs. They're actually quite technical. They've got tens of thousands of engineers, whereas a lot of other non-tech companies are really not technical and they're outsourcing to Accenture. And so one of the key things that I saw was these companies knew that AI was adding value. They saw on the news, hey, Klarna has saved tens of millions of dollars by, uh, you know, customer service automations, and Amazon saves this amount, and this company saves hundreds of millions of dollars here. So they, they saw that AI was creating real value. They were playing around with themselves saying it was creating value and they just, they had no idea how to actually get it done. And so what I ended up seeing is they were all going to Accenture. And what was happening is the board was yelling at the CEO saying, where's our AI strategy. The CEO was yelling at the SVPs and other C-suites saying, what…

AI assessment note: “I honestly got off the call sometimes with these customers and just went, man”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. Amazing. What was the hardest part of the, you know, this exploration? Anything that you struggled with, like figuring out? In terms of putting the agent together.

A So I think one of the big ones is, is the user experience and a lot of the design decisions around, you know, is there a one-to-one mapping between environment and conversation, or do you have multiple conversations in the environment and how do people think about that? And this is something that will clearly evolve, but making that accessible, not only to the really plugged in one percent that is listening to this podcast or Folks who really know their stuff, but also to maybe the more AI skeptical or the ones who think that tap tap autocomplete as of today is the pinnacle of, you know, AI based writing, writing code and making these decisions, sorry, there's a camera popping in here, making these decisions, making this really accessible, I think has been one of the challenges.

AI assessment note: “one of the big ones is, is the user experience and a lot of the design decisions”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. Is there anything that it, that it cannot do or you warn people not to, not to do, you know, like what, what is like within, well within normal capabilities and what is flakier, like in your experience?

A So I think we're seeing the very similar problems that, that a lot of other folks are seeing where, um, some, some prompting, some input is reasonably naive. So the under-specification problem, we, we also see a lot with, uh, especially people who come to this for the first time. And the second is the decomposition piece. You know, if you give it sort of a vaguely specified too large task, it's gonna go astray much like it would with, with any other agent. The advice that we generally people give to people is like, start small and work your way up, learn how to decompose, learn how to really be excessively specific. About what you want. Don't make assumptions necessarily. And that, that gets folks quite far. So there is a learning curve we find, you know, it's not, I mean, the audience will also know that of course, like this isn't some magic fairy dust that you just give some, it can't read your mind, you know, you have to specify what you want. And this is part of the challenge also in building this product of how you educate folks who sit in front of that prompt box.

AI assessment note: “if you give it sort of a vaguely specified too large task, it's gonna go astray”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q have only so much whatever. And like you, you shut it down. So there's like a cold start and all that. Uh, I opened up my laptop. It's a super powerful machine. That's like M for Mac or whatever. And it's on, it's always warm because I'm on it. So, you know, like, how does this work? How does, how does, how does the ID go away? The local ID.

A It's interesting. I think the, um, the reason, I think there are two main reasons why folks really love their IDE. One, we're all creatures of habit. Like this has been conditioned for 30 years. We've been writing software like this for decades. You know, there's a long tail in that change for sure. Like just changing the habit alone is going to take a long time. And it shows in the way, you know, folks have set up their key bindings and really made it their home. The second is, I think the way you work, like the, the way you write code influences the tool you want to use. And so as long as you're writing code in this deep, mono-focused work, where you're doing one thing at a time, an IDE is a good tool. It is built for that. And we've spent, again, decades optimizing this one particular way of working. What's changing with agents is as they gain more and more autonomy, We'll be asked to turn this autonomy into productivity, and the only way we can do that is by doing multiple things in parallel, as we've shown. The moment you do multiple things in parallel, IDEs really aren't the way. They're not built for this. It's cognitive overload. If you try and alt-tab between a bunch of cursor windows, good luck to you. You know, you're not going to have a good time. So we will need new interfaces that help make this parallel work not only easy, but highly enjoyable. You know, how do y…

AI assessment note: “The moment you do multiple things in parallel, IDEs really aren't the way.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q this thing is broken. Kind of sucks. How do you think about problems that you want to solve versus research that you do to highlight some of the problems and then hoping that other people will participate? Like does everything that you talk about is on the Chroma roadmap basically, or Are you just advising people, hey, this is bad, work around it, but don't ask us to fix it.

A Going back to what I said a moment ago, like, Chroma's broad mandate is to make the process of building applications more like engineering and less like alchemy. Um, and so, you know, it's a pretty broad tent, but we're a small team and we can only focus on so many things. We've chosen to focus very much on one thing for now. And so I don't think that, I don't have the hubris to think that we can ourselves solve This stuff conclusively for a very dynamic and large and emerging industry. I think it does take a community. It does take, like, a rising tide of people all working together. We intentionally wanted to, like, make very clear that, like, we do not have any, like, commercial motivations in this research. You know, we do not posit any solutions. We don't tell people to use Chroma. It's just, here's the, here's the problem.

AI assessment note: “we do not have any, like, commercial motivations in this research.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What was your first story? How did you, how did you get started covering?

A Yeah, I mean, I think, I mean, I'm sure that my first story was probably just, like, some silly, like, write-up of, like, a report or something, but I think the first story I remember doing about AI was during summer, 22. I, I live in New York, so every couple months I go visit SF, and specifically on that trip I went, and everyone was like, oh my god, let me show you this, like, funny app called Dolly, and, like, we're gonna make pictures of, like, cats floating in space, and it's, like, so funny and cool, and then it just started to come up in so many conversations that I was like, huh, like, this is, like, a fun Little trend. Like, maybe I should write something about it. So then I ended up doing a piece that was, like, why VCs are obsessed with, like, this new area called generative AI. And then obviously, like, three months later, it was, like, blown up and Chat TVT was, was released. Um, but yeah, I remember that was my, like, first generative AI-specific piece that I did, which was, like, honestly kind of just, like, an accident. Like, I just happened to talk with a bunch of VCs who were excited about it. Um, but yeah, and then ever since then, I've been, been covering the space, so.

AI assessment note: “I think the first story I remember doing about AI was during summer, 22.”

← previous page 6 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.