The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

640exchanges match
640on raw tape
31redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do we have to try to fit large language models, the compute that it requires, into kind of the human box? No, we don't. Why does it have to be, well, if it's not exactly how humans learn, then it's not the right way.

A No, I absolutely agree with what you're saying there, which is, uh, I call this a sour lesson, right? Like, every time you try to make a human analogy to machines, you probably fail because machines develop very differently from humans. Uh, so I, no, I, I agree with that, and I also think that it can still be an alien form of more efficient learning. Doesn't matter. We just know it is super inefficient. Like, that, that is something that we know is a unmitigated negative. So let's make it more efficient and better, uh, and that means that, you know, if, in order, instead of 2000 examples to learn one thing, what about 20 examples? What about two examples? Um, and that scales a lot more, um, and that means, You know, we can, we can get, we can actually get to a point with continual learning that we can actually have, ah, agents that adapt and, and build up a real world model. Otherwise, we're always stuck to the pre-trained, post-trained paradigm that is probably hitting some kind of limit right now.

AI assessment note: “every time you try to make a human analogy to machines, you probably fail”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q generation Opus, uh, GPT-Five. The, the amount of value that can be extracted from those models still seems, at least to me, to be, uh, critical. And, and now we have a whole new A generation of model that we're even going to get more Model LiveWay hang from. What do you, do you agree with that? Do you think you need to keep building the tools around the model?

A AI engineer exists in the white surface area between the peak capability and deploying it everywhere else, right? So the more model research peaks and spikes capabilities in one domain, but it's not evenly distributed in all products yet, that's where engineers have a job. Forever, basically. Um, so I'm very pro that. Um, I, I, I think capability over overhang will exist for a long time. I think it does keep, have waves of consolidation where you're like, actually all this stuff I built out, I don't need it anymore because the next model has got it from just like a single prompt. So I'm going to throw it out. But like, we do spring cleaning every now and then, like that's normal. And we build it up on previous gen model assumptions that then go away. And I don't think we should feel any attachment to the code. At the end of the day, we're all just trying to like serve customers better, do work cheaper, faster, easier. Yeah.

AI assessment note: “so I'm very pro that. Um, I think capability overhang will exist for a long time.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do we have to try to fit large language models, the compute that it requires, into kind of the human box? No, we don't. Why does it have to be, well, if it's not exactly how humans learn, then it's not the right way.

A No, I absolutely agree with what you're saying there, which is, uh, I call this a sour lesson, right? Like, every time you try to make a human analogy to machines, you probably fail because machines develop very differently from humans. Uh, so I, no, I, I agree with that, and I also think that it can still be an alien form of more efficient learning. Doesn't matter. We just know it is super inefficient. Like, that, that is something that we know is a unmitigated negative. So let's make it more efficient and better, uh, and that means that, you know, if, in order, instead of 2000 examples to learn one thing, what about 20 examples? What about two examples? Um, and that scales a lot more, um, and that means, You know, we can, we can get, we can actually get to a point with continual learning that we can actually have, ah, agents that adapt and, and build up a real world model. Otherwise, we're always stuck to the pre-trained, post-trained paradigm that is probably hitting some kind of limit right now.

AI assessment note: “every time you try to make a human analogy to machines, you probably fail”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q generation Opus, uh, GPT-Five. The, the amount of value that can be extracted from those models still seems, at least to me, to be, uh, critical. And, and now we have a whole new A generation of model that we're even going to get more Model LiveWay hang from. What do you, do you agree with that? Do you think you need to keep building the tools around the model?

A AI engineer exists in the white surface area between the peak capability and deploying it everywhere else, right? So the more model research peaks and spikes capabilities in one domain, but it's not evenly distributed in all products yet, that's where engineers have a job. Forever, basically. Um, so I'm very pro that. Um, I, I, I think capability over overhang will exist for a long time. I think it does keep, have waves of consolidation where you're like, actually all this stuff I built out, I don't need it anymore because the next model has got it from just like a single prompt. So I'm going to throw it out. But like, we do spring cleaning every now and then, like that's normal. And we build it up on previous gen model assumptions that then go away. And I don't think we should feel any attachment to the code. At the end of the day, we're all just trying to like serve customers better, do work cheaper, faster, easier. Yeah.

AI assessment note: “I'm very pro that. Um, I think capability overhang will exist for a long time.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I'm curious, we guys talk for two hours. Anything you can share about, like, surprising, uh, use cases or whatever. It's like, you talk about the tech step, but just like, how does he think about it? And maybe open your mind and say, oh, you know, can you let it both?

A I think for me, his use case really helped for me to kind of crystallize what direction we want to build in. Going back, I would say a month ago, there were these two overlapping, but kind of distinct directions when we're thinking about adoption of claw type of agents, autonomous agents in a business setting in a company. Uh, so there's the one where it's the team manages agents. So it's agents that you build an agent factory or, or agents that automate workflows. I gave a talk about that, the agent factory that we built, and we had started to build that already for ourselves, and we've been building it for a while. It's still kind of under construction. So that's one to see factory. It's the team as a team working together to build the agents and managing them as a team. And then you've got the other side, which is personal agents in, in a work setting, in a business setting, but individual people who have their agent, their assistant that's helping them do their job, and it's more one-to-one. And we were going, working on both of those use cases, both internally, we were using agents in both ways within our team, uh, and started to work with design partners that were interested in both use cases. So we were working with a team where they wanted to give each person in their legal team their own personal assistant. And then we're working with another, uh, design partner that w…

AI assessment note: “his use case really helped for me to kind of crystallize what direction”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q roadmap. Everyone in your category has, has similar things. But I think probably the two of you are leading the two most unique and differentiated initiatives, uh, on, uh, in the landscape. Maybe we'll start with, uh, with, uh, Omnigent, and then we'll, we'll, we'll, we'll go into it. I do think that a lot of People are exploring this sort of meta harness concept. What led you to it?

A Yeah, there were actually a couple of like converging lines, which I think is a good sign that you need something new. So on the one hand, there's all the coding agent info internally. We have a really great, uh, dev info team. Uh, they built something called Isaac that's basically like a wrapper on cloud code and, and codex and, uh, let's you use them either on the web and like, Sandboxes or, uh, just on your dev machine or on your laptop or whatever. And then, you know, they were adding all kinds of stuff there. And, and we saw all the, the sort of more advanced engineers, like, uh, were building their own workflows with tons of agents and they were building their own UIs and stuff on top or even on top of that. And then the other one was like us building agents. We ship this like data science agent called Genie on the research team, which I, I co-lead basically. We, Also build a lot of internal ones for various things. And then we have all the customer ones and all of them running into this thing of like, oh, I need to switch model and harness and so on, uh, you know, every few months. Plus the agent is like completely useless if you can't share sessions with someone and have history and have search and all this like layer on top of it for collaboration. I thought a bit about it from both contexts and, uh, at first people thought it was weird. They're like, why are you doing…

AI assessment note: “there were actually a couple of like converging lines, which I think is a good”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah, you need to, basically people want to take something that works in Forkid, and you might as well have something open source. Yeah, which, which also was another question, which is, Interesting for a Databricks, like what do you choose to open source? What do you choose to make it proprietary? And I mean, this goes back to Spark, right?

A Yeah. One, so, I mean, one of the reasons to open source something is if you think it's a layer that will actually, there'll be some network effect. It'll benefit from many people collaborating, um, on it. So, uh, for example, with Spark, I don't know if you, if you know what, when, when Spark came out, we, we also focused a lot on letting you have libraries on So like they used to be different distributed computing engines for like machine learning and graph computation. We said they should all be libraries that you can compose. And we made it super easy to add connectors to data sources too. And then we benefit because, you know, we, we don't have the time to write like connectors to like, you know, a thousand like different databases and, and file formats, but we can just use the ones people make. And of course they benefit from joining, uh, You know, kind of this, uh, this thing. So that's like one of the reasons. Another way to think about it is like, imagine, you know, uh, we, our thing wasn't open. We had some kind of agent hosting thing, but it's not open. And then there is an open one. If you're, which one's gonna win in the long run? So like here, because there is this benefit from like people writing integrations, it'll be, it'll be that. And then there are other things that like, you just can't, uh, even deliver as open source that are things the company does. Like …

AI assessment note: “one of the reasons to open source something is if you think it's a layer”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you see a future where, you know, small models get good enough? Like, does it cannibalize? It's an interesting position. You have big Gemini, you have Gemma, both get exponentially better over time. Like, current Gemma is much better than what we had closed source a few years ago.

A Yeah, for me, it's quite exciting. I mean, if you look at Gemma, you compare to how we were one year ago, I would say Gemma four is matching state of the art from one, one and a half years ago for most. Things. With local models or models that you can run in your own hardware, you can get capabilities, so you can get agent capabilities, function calling, system instructions, like conversational, and that kind of stuff. Knowledge is much trickier, so for knowledge, you do need a larger model, right? That's why if you compare Gemini to Gemma, Gemini has much better knowledge understanding of the world, right? Like facts, information, and so on. So it really depends. I do think We are heading towards a future in one, two years where imagine like you can run a Gemini three pro powerful model directly in your phone, right? And I think once we get there, things will be quite exciting, uh, from our product integration, from which experiences we can, uh, enable the users. Uh, I wouldn't say it cannibalizes, uh, it's still like two very different things. Like if you want that flagship capabilities, like this super complex, long running task, you would use Gemini if you need factuality and so on. But I don't think for many of these things, we'll get to a point in which we can do very powerful things directly on device.

AI assessment note: “I wouldn't say it cannibalizes, uh, it's still like two very different things.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Are people fine tuning? Outside of, you know, we see a few big companies do, okay, like Cursor has a really good consistent, there's a few that have done fine tuning, but it seems like it's not picking up as, you know.

A Yeah, so there was this period, 20, 24, in which there was like this, maybe 20, 23, like there were all of these fine tuning communities, and I think it's been changing quite a bit over the last two years, because models are getting very good out of the box. So as I was saying, like for Gemma four, we had 50 To 60 partners. And some of them were like, oh yeah, we're going to try and fine tune the 27 B model for this vision task. And they were like, oh, actually the model works too well out of the box. We don't need to fine tune it. Yeah. We saw lots of those things. So I'm seeing this excitement around fine tuning nowadays as general conversational models. Yeah. There is still quite a bit of excitement around fine tuning for specific domains like finance, healthcare, Specific types of data that the model didn't see, but as general conversational, like just changing how the model behaves , you can do most, most of that via prompting nowadays, and in terms of capabilities, the models are very good out of the box. So it's been changing quite a bit. There's still like the onslaught people. I don't know if you know, uh, Daniel Han.

AI assessment note: “it's been changing quite a bit over the last two years, because models are getting very good”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So this is if you go above a hundred percent, right? Like your overflow. If your overflow like spillage or whatever, you probably lose money on it, but it doesn't matter, right?

A Well, you might, you might not, that is a more cost effective way to do it, but it's a slower way to do it. Because basically what you have to do is you have to like queue your requests, spin up these just in time compute, um, get it all ready, provision it, and then get your workload there. And so if the time isn't important that much, that's fine. And you can do that, but if your customer, and especially for, let's say the RL training runs, the reason why a lot of people come to us is because GPUs are more expensive than CPUs, right? So you want your GPU running at what? A hundred percent the entire time. And so when you're running runs on CPUs, when the, when the CPU cycle is like down and spinning up the next one, you want that to be instantaneous so that your GPU doesn't go down.

AI assessment note: “Well, you might, you might not, that is a more cost effective way”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q This is the equivalent of my mom test, right? Like, what have you done that has your solution to this?

A So internally we built a tool called Central Station that allows us to go in and aggregate all the context from all of our users. Um, so every piece of feedback, every piece of customer support, every single thing like that, uh, gets aggregated into, uh, what we call like clusters. If you have an incident brewing or like anything else like that, now we can go and determine how many users are affected, all of those other things, et cetera. And then we can actually break off a discussion based on that. And I think a lot of that is actually a lot more, more helpful and more correct in terms of, Instead of like having just these like long running channels where you're just like, which channel should I put this thing in, right? Like if you can dynamically aggregate that information and dynamically route it to the right person based on the context, right? We know, we know internally like these four people are pretty close on networking, right? And so if we see like, okay, we've got a networking thing, you can roughly like drill it down to like those four people, right? And if you're saying like, oh, okay, cool. It's actually with this part, you can just go and like look at the commits, right? And this is, like, no longer a manual process internally. Like, this is the whole point of why we built, if you go to, like, uh, station or help.railway.com, there's a whole reason we built this…

AI assessment note: “internally we built a tool called Central Station that allows us to go in and aggregate”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. You mentioned the tick stack, Peter. Uh, so I just wanted to give you some reign to just go into it. I'm interested in where by nutrition, uh, starts and ends in, in, in, in some sense, what won't you do? What do you do that's common among all the verticals that you cover?

A There's a few buckets of, of work that we do and, and we've been at this for almost 10 years now, so the technology's pretty broad, but, uh, we got started with a thousand engineers, like you could work on lots. There's lots of stuff you have, especially with AI tools now, but yeah. So we had our start in, in simulation and simulation tooling and infrastructure. And so generally, if you're trying to build a very complex software system that involves moving machines, you need to test that. And the best way to test it is it's a combination of virtual developments, a simulation, and then also obviously real world testing. And then there's a very careful process of that correlation between the simulation results and the real world results and ensuring that the simulator is in fact accurate to that. Simulation is a very deep topic. We have a whole, whole suite of products in that, and we can talk for many, many hours about that specifically. Um, but that, that is one part of what we do as a company. Reinforcement learning as a sub part of that is also super critical. I think a lot of the, a lot of the best advancements happening in, in a lot of these AI systems right now in some way relate to reinforcement learning and With now we have lots of compute, and you can do tons of interesting things in reinforcement learning. The second bucket of work that we do is operating systems techn…

AI assessment note: “There's a few buckets of, of work that we do”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I'm curious about the, um, coding agent adoption, just like since you're mentioning more esoteric languages, like what's the adoption internally? What have you learned?

A Yeah, we, we use everything. Um, I mean, so cursor was, I think the, the hottest tool in the company for a good while. Now, Claude Code, I think, has, has taken the, the reign on that. We have a internal leader, leaderboard that we use just to sort of encourage adoption, uh, with, within the company. And, uh, yeah, they're phenomenally useful. I mean, it's, uh, honestly, we, we take inspiration from, from some of those tools also, and how we're adapting some of that mindset of thinking to the physical realm. Like, if it's so easy to, to build an app for this or that thing that lives just on a screen, We can, we were taking out a lot of the same ideas and, and applying that to, okay, well, if you wanted a physical machine to do something, how easy can we make that, uh, using our own tooling and platform as well?

AI assessment note: “cursor was, I think the, the hottest tool... Now, Claude Code”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q The other thing I was thinking about is just in terms of like AI, uh, adoption, does that change your hiring at least a little bit? Or how do you, how do you sort of manage engineers, um, differently?

A Yeah, absolutely. It does. Um, we, I think like every company in the Valley right now are evolving our, our hiring practices, um, because the, the skills required to be effective are changing so fast, right? I mean, you, you used to really select for just rote implementation ability and, and now it, it is more the AI engineer skillset, right? Where it's like, yeah, you know how to implement, but actually Just banging out code is, is no longer the core job, right? It's, it's actually knowing what questions to ask, knowing how to tie, how to tie together these different AI tools. And so the, the interviews that we give now, I think are way harder than they've ever been, but, but we also allow, right, selective use of AI tools to solve the problems. And I think in that you, you start to see more of a bimodal distribution of engineers, right? You, you start to see like, wow, there's, there's this subset of, of people that they, they really get it. Like they're, They're all in and they've, they've clearly invested the, the hours needed to learn these tools and, and how to be effective. And then there's sort of the, the group of people that haven't done that and that the productivity gap is just enormous. And so we're, we're trying to obviously select for the people that are really, really into this.

AI assessment note: “Yeah, absolutely. It does. Um, we... are evolving our, our hiring practices”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I reached out was because you started, Promoting more sort of internal tooling, uh, primarily Tangled, but also a lot of people have seen and adopted Toby's QMD. Uh, and obviously I think, uh, Shopify has always been sort of leading in terms of engineering. I think more, it's just more recent that you guys have been more vocal about your sort of AI adoption. Is that, is that true?

A Well, I think AI tools in general are fairly recent development, uh, and, uh, we, Shopify, you know, at this stage of its development, we're developing AI in-house and building tools that use AI and, you know, interfacing with the wider AI community, you know, are on the sort of the runaway trajectory. So it just did by sort of natural byproduct. We talk about it more also. We just, uh, just even yesterday, Andrej Karpathy was famous in tweeting about, oh, there's some, uh, ways that you can organize your agents to store the data and then look up the data so that you don't have to research or lose context every time. And a little bit tongue-in-cheek, I tweeted that, hey, we've done it much earlier, and we even have different approaches, Toby and I. Toby, of course, is a big fan of QMD, and I'm More of a SQLite fan, but, yeah, very similar things that we've already done here. The point is, yeah, we're very dynamic, you know, explosively growing company, and we have to be at the forefront of AI adoption, obviously.

AI assessment note: “So it just did by sort of natural byproduct. We talk about it more also.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q think SimAI, I think Yongjun Park who did the Smallville thing. There's a very small cottage industry of people trying to do like the simulated customer thing. I think a lot of people maybe don't super trust this yet because they're like, well, obviously they would just do what you prompt them to do, right? But maybe just think, uh, tell us about the sort of inspiration or origin story.

A That's exactly actually the thing I wanted to cover because if you don't have the historical data, All you can do is prompt agents in the vacuum, and they will do exactly what you prompt them to do. In fact, when I first proposed it, and this is a bit of a, my brainchild initially, if I, I can boast. And then Toby said like, but wouldn't they, they just repeat what, what you tell them? And, uh, but I'm like, yes, except Shopify has decades of history of how people made changes and what there is, uh, what it resulted in terms of sales. So now what we can do is we can, we have this, it's not, it's a noisy data. These are small, usually websites, uh, you know, like things, things are never in isolation. It's almost never AB experiment. It's always AA experiment when there's, has two meanings, but basically, you know, in different time, you run two different things. But if you aggregate in general, uh, like everything together and you apply, uh, denoising and collaborative filtering like approach, you can extract a very clear signal and then you can optimize Your agents, and that's why it took so long. It took almost a year of that optimization of just us sitting and fiddling, and, and we had this internal goals of correlation of heating. Internal goal was to hit 0.7 correlation with, uh, add to cart events, for example, like that, that if we run real A, B test experiment, that it …

AI assessment note: “when I first proposed it, and this is a bit of a, my brainchild”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Uh, okay. Yeah, but, I mean, we, we can, we can discuss the, the, the release briefly, because we'll release this after the, after it's already announced, or whatever. There's a catalog that you guys are doing?

A Yeah, so we are, we are, We are bringing in capabilities of a whole Shopify catalog. Basically, you now, you can search for products. You can do lookups by specific ID. You can do bulk lookups when you need to bring multiple products. You don't need to know in advance what you're trying to show or to sell or check out. Like you can now, you can now have this decided at runtime and this big area for investment for us For both non-personalized and personalized searches, trying to provide basically a window into whole universe of products that are being sold everywhere in the world. And Shopify is really not exactly, but almost like a superset of anything being sold. Now we're bringing it into UCP and, uh, and, uh, identity linking is another big thing for us so that you, you can use, uh, Like Google or whatever, whatever identity you have, uh, they're minimizing a friction.

AI assessment note: “we are bringing in capabilities of a whole Shopify catalog. Basically, you now, you can search”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q worked on. I'm curious about the backlog, right? Like, the, the, the, I actually don't mind a pro-level model taking an hour, two hours to review my PR, because I've dealt with humans who take a week to review my PR, right? And I keep pinging them on Slack, hey, hey, review my PR. So, you know, I think there's some trade-off here where, like, it still doesn't make sense.

A Exactly. That's exactly my point, uh, that on one hand, you can tolerate longer latencies at PR. On the other hand, like right now, the real problem is not in Spending time waiting for PR is real problem is since there's so much more code, then, uh, probability of at least some tests failing, going up, and then you, like, keep failing, then you have to find the offending PR, evict it, retest it without that PR, and so deployment cycle becomes much longer. Uh, so it actually, in terms of the overall time to deploy, it's total time savings if you spend more time on a longer model, like, thinking for an hour, because then, then you, you don't have to spend all that time During testing and rolling, you know, rolling back the deployment.

AI assessment note: “total time savings if you spend more time on a longer model”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Like, is there ever a discussion of, like, we're not going to ship it because we're not able to tie it down? Or are you happy to just, like,

A No, I mean, there are a lot of things where we choose not to use MCP because we want to add more high touch to quality. I think search and agentic find is like the largest instance of that, where we have, um, slack and linear and JIRA search and notion that is not using necessarily the search MCP functionality that is provided by those companies. And that's because it's quite critical. We think to how our agent trajectories work is for us to have, A little bit more control on the functionality of the search journey, and so it usually comes from quality, and there's a long tail of things, and that's why we built an MCP client, or an MCP server, excuse me, so that people can connect whatever they want. There is that long tail, right? But we, for search particularly, I would say that's like the primary entry point, but there are other connections as well that it's a little bit of secret sauce about when we are okay with like MCP functionality and user-driven auth, and when we actually want We want to carry a lot more ourselves.

AI assessment note: “there are a lot of things where we choose not to use MCP”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q you know, we're selling enterprise SaaS. So if we sell credit packs and you get discounts, if you're an enterprise and you buy a certain amount of credit packs and things like that. So it also just helped the sales motion, um, work a little bit easier. So that's the answer on the abstraction of credits to dollars. Now, Was the question how we decide how to price it, or?

A Yeah, like, I mean, I think there's, all tokens are not made equal, but we obviously get charged mostly equal. Like, you can ask Codex to create you a dumb tool for, like, I created one for our StarCraft II LAN for people to, like, find the game, uh, but then people create it to build features in, like, billion-dollar companies, but the token price is the same. Yeah. Like, for you, I can ask this to update my favorite recipes doc. And it'll do it, but I could ask it to like respond to an email from an investor. And like the value is like very different, you know, and you could charge more, but you're not necessarily doing it. So I'm curious if there was any discussion.

AI assessment note: “Yeah, like, I mean, I think there's, all tokens are not made equal”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Just to double click on, on, uh, do you think Alex Blania with worlds, do you think he's got it or is there an alternative?

A Oh, so I mean, there's gonna be, I think there'll be, I think many people will try. We're one of the key, you know, participants. In the world, in the world project. Yeah, so we're partisans. But yeah, I think, so we think world is exactly correct. And the reason is, it has, it has to be, it has to be proof of human. It has, because you can't do proof of not bot, you have to do proof of human. To do proof of human, you need, you need biological validation. You needed to start with, this was actually a person, right? Because otherwise, you have bots signing up as fake people, right? So you have to have like something, you have to have a biometric, and then you have to have cryptographic validation, and then the ability to do, to do the lookup. And then by the way, the other thing you need was that you, you also need selective disclosure. Um, so you need to be able to do proof of human without revealing all the underlying information. By the way, another thing you're gonna need, you're gonna need proof of age, right? Because there's all these laws in all these different countries now around. You need to be 13 or 16 or 18 or whatever to do different things. And so you're gonna, you're gonna need, you know, sort of validated proof of age, um, you know, to be able to legally operate, right? And so that, that's coming. And then you're gonna want like proof of credit score and, you kn…

AI assessment note: “so we think world is exactly correct. And the reason is”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I mean, to me, like, Showing this gives the engineer a complete mental model of what you've done, what you can do with it. For example, the first thing I, I look for a mental checklist of things, right? Like is off in the database. Off looks like it's not, right? So that's a separate layer. That's probably means it's hard to do multi-user apps on the same app, right?

A So you actually, we've solved that. So, um, yes, the platform builds in off. So you as a user sign into the platform. If you're using an agent that was published by someone else, then your identity is, is kind of taken care of by the system. And when you query the database, you're going to get the stuff that is for you, unless the builder specifically said this is public data that everyone should see. So they, they actually get a chance to think about that. And again, sidekick can guide you through building, uh, agents and apps that work that way. So you're right. That's another thing that people have to think about when they're trying to figure out how to build software experiences on dreamer. You, it's built in, you talk to the sidekick as if it were a human being about what you want, and that's what you get. So, you know, my, my big sky app that I just showed you, that was designed for multiple people to use it. And of course the things that we were putting in as expenses were supposed to be visible to everybody. And I just told the sidekick, that's the way I wanted it. Uh, but by default, if I built an app like that, the data from each user would not have been visible to the others.

AI assessment note: “we've solved that. So, um, yes, the platform builds in off.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah, exactly. Um, and then I think, how's hiring changed? Yeah, you've hired plenty of software engineers in your life. I assume something's changed.

A Yeah, absolutely. So one of the main things that I look for now when hiring engineers is how well do you work with coding agents? Our team actually is quite experienced. A good number, everyone at Dreamer, other than, well, I guess I write a lot of code too, everyone's an IC, an individual contributor. Many of the folks that work on the team have previously been managers, and it turns out being an engineering manager, as long as you stay very close to the code and are able to continue to craft it yourself, Is actually a great skill profile for being able to make agents work for you and for your team in this, uh, in this age. And so that's definitely something that we look for quite intently when hiring engineers. And, um, we still have folks write some code like with their fingers. It's just important to know that the kind of core of the craft is there, but the vast majority of what we spend time doing is building quite significant and elaborate stuff together in a fun collaborative environment with coding agents.

AI assessment note: “one of the main things that I look for now when hiring engineers is”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q it's like fundamentally cloud code. We don't want to touch it. There's the cloud app. There's cloud in Chrome. I think you guys do something different in planning, but, uh, I've been talking with Tariq who is on the cloud code team and you guys are, he's like, no, we just exposed planning. Maybe you can clarify what are the major pieces that people should be aware goes into co-work?

A Like, okay, I think you basically have them. So, um, you can take planning more or less out. I think that's a few things that are really valuable in Kovac. Um, the virtual machine is probably the most powerful thing. So we currently run like a, we currently run like a lightweight VM and we put clock code into the VM and we do that for, for, um, a number of reasons. Safety and security is a big one, but even if you, even if you ignore for a second safety and security and you're just like, okay, Yolo, I want this thing to do whatever. It is quite powerful to give cloud its own computer. That is like generally a good idea. And in terms of architecture and UX and everything else that we've been working on Anthropic, it often is quite useful for you to like anthropomorphize, um, cloud aggressively and just be like, this is a person. What would you do if you give, if you had a person, right? And the analogy I've given my dad this morning, who is still like quite insistent on using chat, even for like coding things is if you were a developer, And your employer told you that you don't need a computer. They're just going to like send you emails with the code and you send emails with code back. Like that maybe worked for Petrars in the back, but that is not very effective. Um, so what we can do with the VM is because it's a, it's a Linux system. Cloud code has more or less free reign to …

AI assessment note: “The virtual machine is probably the most powerful thing.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. Have the thread models been updated a lot, or do you feel like you're still using the same thread models as GPT-II of like, you know, paperclip factory, blah, blah, blah, you know, but like, how much are you rising, you know, increasing the bar?

A Yeah, so I'm not an expert in, in the threat modeling piece, more in the, more in the capabilities piece. Um, I, I do think they've been changing to, to some extent. So, something like the autonomous replication threat model, that is being able to set yourself up and, and control resources, something like that, has been deprioritized relative to, um, AR and D acceleration. That is, you know, the possibility there could be some capabilities explosion inside of, inside of a lab, and that could be destabilizing for, for all sorts of reasons that we could talk about. Um, so it's mainly, mainly we're focusing on that latter one, although, although we do think about a, a, a number of trend models.

AI assessment note: “I do think they've been changing to, to some extent.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q of like, is there a way to bring the thousand picojoules down to 50? Like, is it worth designing a new chip to do that? The extreme is like when people say, oh, you should burn the model on the ASIC, and that's kind of like the most extreme thing. How much of it is it worth doing in hardware when things change so quickly? Like, what's the internal discussion?

A Yeah, I mean, we, we have a lot of interaction between, say, the TPU chip design architecture team and the sort of higher level modeling, uh, experts because we really want to take advantage of being able to co-design what should future TPUs look like based on where we think the sort of ML research puck is going, uh, in some sense because, uh, you know, as a hardware designer for ML in particular, you're trying to design a chip starting today And that design might take two years before it even lands in a data center, and then it has to sort of be a reasonable lifetime of the chip to take you three, four, or five years. So you're trying to predict two to six years out where, what ML computations will people want to run two to six years out in a very fast changing field. And so having people with interesting ML research ideas Of things we think will start to work in that time frame, or will be more important in that time frame, uh, really enables us to then get, you know, interesting hardware features put into, you know, TPU N plus two, where TPU N is what we have today.

AI assessment note: “we really want to take advantage of being able to co-design what should future TPUs look like”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And so like, how do we, I guess, do we want to extract that? Can we, can we divorce knowledge from reasoning, you know?

A Yeah, I mean, I think you do want the model to be most effective at reasoning if it can retrieve things, right? Because having the model devote precious parameter space to remembering obscure facts that could be looked up is actually not the best use of that parameter space, right? Like you might prefer something that is more generally useful in more settings than this obscure fact that it has. Um, so I think that's always a tension. At the same time, you also don't want Your model to be kind of completely detached from, you know, knowing stuff about the world, right? Like it's probably useful to know how long the Golden Gate Bridge is just as a general sense of like how long are bridges, right? And, uh, it should have that kind of knowledge. It maybe doesn't need to know how long some teeny little bridge in some other more obscure part of the world is, but, uh, It does help it to have a fair bit of world knowledge, and the bigger your model is, the more you can have. But I do think combining retrieval with sort of reasoning and making the model really good at doing multiple stages of retrieval and reasoning through the intermediate retrieval results is going to be a pretty effective way of making the model seem much more capable. Because if you think about, say, a personal Gemini,

AI assessment note: “having the model devote precious parameter space to remembering obscure facts that could be looked”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. Maybe run people through the Brex agent platform. We'll put the diagram in the video where you had the LLM gateway, you have like the whole MCP layer. We just had David, the creator of MCP. Right before you. So this is very timely. Um, yeah. How did you start building that? What's the architecture?

A Yeah, the architecture, you know, I, I think simple is, uh, is elegant and we, we've had basically an LL gateway and, and, uh, a basic hand rolled platform, uh, from the very early days. In fact, right before being tapped to become CTO, I was leading, uh, like an AI, uh, labs team internally, uh, in the wake of like the announcement of ChatGPT, you know, everybody saw this through technology and said, Hey, what are we going to do with it? And so one of the first things that we did, um, I think January, 20, 23, that would have been, uh, was Try to put together some internal infrastructure that made it possible for us to deploy, deploy, manage version and eval prompts, uh, and then be able to manage, uh, like data egress and model routing and, uh, have some very basic like observability and cost monitoring, uh, in an LLM gateway. So that's, that's infrastructure that we stood up and it still continues to power a lot of those smaller, uh, more, let's say like precise applications of LLM. So like, for instance, we've, uh, we set up a completely automated, uh, Pipeline for, um, evaluating, uh, customer applications to get them onboarded instantly to Brex, which is something that used to require, um, human intervention either for underwriting or KYC. But now we basically have a series of, of agents and, um, and particularly like research agents that will go and do the work that human…

AI assessment note: “put together some internal infrastructure that made it possible for us to deploy, manage version and eval prompts”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q That's true. I never thought about that. I've been in a database data industry prior and there's a lot of shenanigans around benchmarking, right? So I'm just kind of going through the mental laundry list. Did I miss anything else in that, in this category of shenanigans?

A I mean, okay, the, the, the biggest one, like, that I'll bring up, like, is more of a conceptual one, actually, than, like, direct shenanigans. It's that the things that get measured become things that get targeted by labs that they're trying to build, right? Exactly. So that doesn't mean anything that we should really call shenanigans. Like, I'm not talking about training on test set, but if you know that you're going to be great at another particular thing, If you're a researcher, there are a whole bunch of things that you can do to try to get better at that thing that preferably are going to be helpful for a wide range of how actual users want to use the thing that you're building, but will not necessarily do that. So, for instance, the models are exceptional now at answering competition maths problems. There is some, uh, relevance of that type of reasoning, that type of work, um, to, like, how we might use modern coding agents and stuff, um, but it's clearly not one for one. So the thing that we have to be aware of is that once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the next couple of years. There's no, Silver bullet to defeat that other than building new stuff to s…

AI assessment note: “the biggest one, like, that I'll bring up, like, is more of a conceptual one”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Is the dataset public? Or what's, is it, is there a held out set?

A There's a held out set for this one. Um, so we, we have published a public test set, but we, we've only published 10% of it. The reason is that for this one here specifically, it would be very, very easy to, um, like, have data contamination, because it is just factual knowledge questions. Um, we will update it over time to also, um, prevent that, but with, yeah, kept most of it held out so that we can Keep it reliable for a long time. It leads us to a bunch of really cool things, including breakdown quite granularly by topic. And so we've got some of that disclosed on the website publicly right now, and there's lots more coming in terms of our ability to break out very specific topics.

AI assessment note: “There's a held out set for this one... we've only published 10% of it.”

← previous page 5 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.