The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

1,847exchanges match
1,797on raw tape
133redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And you think this is slowing down, right? Like in your, um, blog post, you have a pretty striking sentence where you say, uh, GPUs will no longer Improve meaning fully. We have essentially seen the last generation of significant GPU improvements.

A Yeah. So this has two components. And so one is also sort of a very fundamental thing and it's physical in the sense that, um, I mentioned these two components, uh, memory movement and computation. And so computation can only be useful if you move memory to this sort of local neighborhood where you do this computation. Now, this is a geometric problem. You need to have a large store of information, and then use this large store to move information closer to where you want to do the computation. And we have figured out how to physically do this optimal. We have like a large, slow memory. That's DRAM. Then we move it to a cache. If you look at the geometry, that's how you do it fast. If you have a certain size of computation, this is optimal. If you have a different size of computation, matrix multiplication, then you want to use not a CPU, but more like a GPU, which has higher latency, but more throughput. You can move more data, but more slowly. And, um, yeah, if you look at all of that, you can push around a little bit how you structure everything, like the caches, and how large they are, and how much cores are they shared. But in the end, the fundamental problem Uh, remains the same. You have a geometric problem. You can only fill the space in a certain way. And that means you always have certain access pattern with certain, um, latencies. And the biggest latency is a big blo…

AI assessment note: “Yeah. So this has two components.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q to 15 years is, uh, like all the brains from academia have been sucked into industry, um, and, uh, you sort of doing the opposite, or maybe both at the same time. Curious for the context, is that more of a personal thing, uh, because you always wanted to do academia, or is there something deeper about, like, the kind of work that you can do in academia versus industry?

A Yeah, it's more about the kind of work. Industry is really great at executing on Ideas, uh, and it's maybe not as good at, like, exploring diverse ideas. Even at the scale of Anthropic OpenAI, uh, there is a lot of focus, uh, in the companies, and there isn't a lot of bandwidth to do exploration, and that has been working extremely well so far. We still probably have a lot of low hanging fruit, uh, left, uh, to, like, get the models to be much better, but I personally find it really Exciting to do more exploratory work and to try things that are different. And for that, I feel like having my own lab in the university is just a better tool.

AI assessment note: “Yeah, it's more about the kind of work.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Why do models do that, or why are they able to do this? Is that basically part of the pre-training, and they effectively learned Being deceptive from us by, by being taught all the deceptive ways humans have behaved over the centuries?

A It's a very interesting question. And, um, yeah, it is quite surprising, actually, that the models would, uh, behave that way after going through some of the alignment training. We don't really know what's the source of this type of behaviors, but that's also true for a lot of other behaviors in the models with, like, even the good ones. We don't really, we cannot always pin down, like, where they come from in the pre-training. I think at least part of it is probably The models, uh, seeing descriptions of AI, like in the science fiction literature going rogue, uh, and, uh, like, yeah, that probably affects how the models behave in similar scenarios. Uh, so for example, in that anthropic study, they have this blackmail scenario where it's kind of really well structured so that the model sees some information about like a CEO of a company that, uh, the CEO is involved in some extramarital, uh, Affair. And then like soon after the model observes that it will be shut down. And then the model kind of puts the two things together and it says, okay, I need to use the first information to prevent me from being shut down. So there is, uh, in the blog post, they note that there is this possibility of like a check of scan that like in the text on the internet, probably if two things co-occur close to each other, then it is likely that they are related to each other. And the model is a sta…

AI assessment note: “I think at least part of it is probably The models, uh, seeing descriptions of AI”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Very interesting. All right. Uh, let's switch to, uh, reasoning. Clearly, uh, was a huge year, uh, from that perspective, just massive progress in reasoning. Where do you think we are in that arc? Uh, and what are you excited about on the reasoning front for?

A The biggest step change in the models over the last few years was the reasoning NRL. We have made a lot of progress, and the progress was very fast, uh, in the beginning, where, you know, there was O-one, but then very quickly after that, there was O-three, and on a lot of benchmarks, the, the progress has been extremely dramatic. I remember when we, like, early in the project of the O-one, uh, there was some discussion of, like, will it solve IMO problems, and that seemed kind of Very unlikely to me, but then, yeah, here we are, it can easily solve a lot of IMO problems. So I think we, like, as a community, there was a lot of progress. I think it's, as with many methods, uh, it's starting to be harder to make progress, or at least visible progress. So kind of similar to pre-training, uh, there is still a lot of progress, but it's, the models are already so good that it's kind of harder to see What changes from one to the other as a user of the model? And I think that's also to some extent true for, uh, for the reasoning now, but they're still increasing the scale of the RL, more environments, more, um, more compute spent and, um, yeah, models are still getting more consistent and better. And I think we are at the stage where if we define a benchmark and we can make a relevant RL environment, then we can kind of max it out Uh, pretty quickly, and so we are going through benchma…

AI assessment note: “I think we are at the stage where if we define a benchmark”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Ok, so that's, uh, stage three, long context. Maybe just to bring this to life, um, what's the difference between before and after? Like, if you have a, uh, 40 page PDF, uh, that you fit into the window, it will just, uh, get faster results or better results. What happens?

A At the beginning, you just can't do it. Like, you know, do you pre-train as something like, Four, 8000 tokens. That's what we use for OMO three. That's what Lama is. That's about maybe eight pages. If you use like, you know, double spacing Uline kind of thing. Um, and after that we extend to about 65, um, in industry you have extension of a million token. I think Gemini recently announced like over a million token. At that point, a million token is like 10 books. Um, so you can work with extremely long amount of information. It's nice. You don't have to think about, you know, if you're building an application with this language model, you don't have to think about like, oh, of this amount of information, how the heck I'm going to extract the ones that I need to show the model. You can just give it all and the model will figure it out. So it's, it's really unlocks a lot of opportunities.

AI assessment note: “You can just give it all and the model will figure it out.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q its entirety for free at stateoff.ai. So, um, we're going to riff on some of the most important topics and ideas in the report, but obviously people can go and check out the report directly for more. All right. So, uh, starting from the top, in the world of research, you mentioned that, uh, your reasoning got real. Uh, so, how far have we come in the last 12 months?

A I'd say pretty far. Um, about 12 months ago or so, we had, I think, the very early inklings of it with O.I. Preview, uh, potentially around, like, this time last year. And, uh, that was the first time you had a system that could kind of Show its reasoning, show its stepwise process to get a more complicated answer. And this has generally been the dream in AI for a long time. And, uh, and since then to now, I'd say like the progress is pretty astounding. One of the areas that the progress has kind of unveiled itself is in mathematics and other verifiable domains where you can like explicitly say, yes, the system works or doesn't work. And, you know, we saw gold medals on the International Math Olympiad by a couple of labs, including OpenAI and DeepMind. That area probably with Uh, if you asked the experts again, how long it would have taken? Probably been a decade. Then in areas a bit closer to my heart in biology and science, we've seen reasoning models, uh, kind of be used as a, as an AI co-scientist. So just as a human would be reading lots of papers, planning experiments, running the experiments, and then doing data analysis, and then reformulating their hypothesis as a result. There's examples of, uh, models doing that in lieu of a human, which is exciting because there's way, way too many papers, uh, to read. You know, AI people kind of complain that it's like 50,000 paper…

AI assessment note: “I'd say pretty far. Um, about 12 months ago or so”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q all the, what used to be known as Thin wrapper. So the vendors that happen to be powered by those models, so the cursors, the windsurfs, and all the, whatever, legal, financial, AI startups as examples. The other big debate in the business of AI, of course, is the bubble Uh, question. What's your, what's your take? Are we in an AI bubble? Are we not in an AI bubble?

A Yeah. I think like with most things in, in markets, there are probably localized bubbles all over the place. And I think at a, at a high level, what's interesting in terms of vibes and who's calling bubbles and who's not like the finance crowd in New York is definitely talking about bubbles a lot more than what we're talking about in San Francisco, where they're, Their view is like, this is the golden era of AI, and a lot of things are working. We have so much more to, to do, uh, you know, compute build outs are enabling us to experiment a lot faster. Uh, you know, this huge flood of, like, talent that's built the consumer internet and cloud computing is moving into AI, and with that is bringing a lot of optimization techniques and knowledge that AI researchers didn't have when they built the first generations of ChatGPT, et cetera. But I think you, you can't ignore the fact that the, The sums of money going into this industry are truly gargantuan. Um, you know, like, five hundred billion to build, uh, Stargate, and then, uh, you know, a couple hundred billion here, a couple hundred billion there, like, pretty soon it's real money. And then the, um, and then, like, the circularity of these deals is, like, is interesting. Uh, of course, NVIDIA is at the center of this, and it has incentives to use its, its balance sheet to, sort of, spin the wheel faster. Uh, and then perhaps mo…

AI assessment note: “there are probably localized bubbles all over the place.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And by the way, just for the lore of it, uh, did you guys have any, uh, sense that, uh, AlphaGo was gonna crush Lee Seedal, so the, the famous goal player that you mentioned earlier in the conversation, was it, was it obvious before, or was it a surprise?

A We thought we had a pretty good chance. But we were very nervous about, like, you know, are we gonna win, are we not gonna win, are we gonna lose? Yeah, we actually had some bets beforehand of, like, how many games are we gonna win or lose? Like, I think it was very ambitious to put the match as early as we did. If we had wanted to be a bit more safe, we may have, like, tried to do a few months later. And I think if we had done it a few months earlier, we would have probably lost. So it was a very knife edge of, I guess, which also made it much more interesting for us, right? Because it really means that each game is like a nail biter of, oh, what's going to happen? Are we going to win? Are we going to play a dumb move? What's going to happen? So that was very exciting.

AI assessment note: “We thought we had a pretty good chance. But we were very nervous”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q and you were the lead author on mu zero, uh, which, um, In the world of AI. I'm sure you're going to be very humble about it, but like in the world of AI is like as, as big a deal as it gets. So, um, I'll say it so you don't have to say it. Um, so MuZero, what was the next, um, what was, how was that different?

A So the main motivation I had for making MuZero was that if you want to solve many real world tasks, you have no way of perfectly simulating what's going to happen. You know, if you play a board game, obviously, you know, if you make this move, You know what's going to happen. It's like the piece is going to go there, it's going to take a piece, whatever, right? But if you actually want to solve something like a robotics task, or anything more complicated, it's impossible for you to simulate what's going to happen accurately. And also, we as a human, we don't do this, right? We just imagine in our head of, oh, I'm going to say this, then he's probably going to respond in that way. This meant that alpha zero, as it was, Could not be applied to such problems because it required some way of, you know, simulating the game, scoring the outcomes. And the idea with mu zero was that, well, we already have a deep neural network, right? These networks can learn a lot of things. So why not let it, why not teach it to predict the future of the environment, the future of the world? Why not make the model be able to learn for itself? What is going to happen after each action it takes.

AI assessment note: “The idea with mu zero was that... teach it to predict the future of the environment”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q the very beginning of this conversation, this, this, this concept of External benchmark. And then you quoted your piece of Goodhart Law. Uh, so what, what, what, first of all, what is Goodhart Law? And then, uh, how should labs compare results, uh, so that it doesn't end up, uh, with this kind of leaderboard theater that we've seen a little bit in the last, you know, couple of years?

A Yeah, so Goodhart's law basically says that any measure that becomes a target stops being a good measure. And, you know, you can think of that intuitively that If you start paying, for example, programmers based on how many lines of code they write, well, suddenly they will discover many ways to add more lines of comments, which is, you know, completely useless. And this is a very general effect that obviously, right, if you give people an incentive that they should optimize, they will try very hard. Yes. And we also see this with language model benchmarks. Of course, people want to get promoted, they want to launch their model. So any benchmark that is too easily measured or that has a lot of attention on it, people will optimize very hard for it, which means that probably the model will look very good at that benchmark, but if you then use it for your own task, you might get different performance. Yeah, you asked about like how, what do we do about this? It's very hard to prevent people from optimizing on the benchmark. So one possibility is just periodically create completely new held out benchmarks. That nobody has seen before. And that gives you a fairly, you know, good estimate of model performance. I know, for example, like a lot of researchers have their own toy problems that they use to test all the models precisely for that reason. So that, you know, this is a problem…

AI assessment note: “periodically create completely new held out benchmarks”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So, um, so Google noticed you, and then what happened?

A Um, so Google noticed me, uh, I started at Google, I think like a month before ChatGPT or something like this, so it was actually, it was a fantastic time to start at Google, because the entire company was suddenly forced to react, um, instantaneously and, and, and compete with the Gemini program. Um, so, uh, it meant that there was this Gap of, I guess, like the typical, uh, command structures and everything were not well suited for that particular battle. You know, it wasn't a pre-existing org. You know, Gemini was sort of forged out of the foundation, out of the merging of, uh, of brain and deep mind. Um, it, it meant that There was just a huge gap in terms of agency, really, of figuring out what we needed to do, doing it as fast as possible, organizing people together to, to work on important things. Um, and so I ended up, uh, one, getting the chance to develop a lot of taste by working closely with people, um, in those like early months of Gemini. Um, But two, also quickly got the opportunity to step up, uh, and, and to lead various parts of this. So one example of this is we just didn't have an inference stack that was, uh, at all, like, sensible for the modern world of LLMs. Um, and so we had to notice that, design one from scratch. A lot of the things you now see, uh, in, you know, like, the sort of SGLangs and stuff of the world are, like, things that we, um, had to de…

AI assessment note: “quickly got the opportunity to step up, uh, and, and to lead various parts”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Is this your haiku somewhere? Maybe walk us through the differences between those models.

A Yeah, so we, Release models along three categories, uh, three tiers. So Opus, which is the smartest model, Sonnet, which is, uh, the mid-tier model, and Haiku, which is the, the fastest, cheapest model. Um, one of the interesting things about, uh, this most recent release is actually Sonnet is smarter than Opus. And this has happened before. In fact, this happens last year. It's a reflection of fast progress because, uh, it, it is cheaper to train, uh, you know, mid-tier models than large models. And so what happens is that you end up doing a lot of progress on smaller models. Eventually you need to choose when to scale up and, and sort of get the benefits of scale in a model. Often you make progress fast enough that your mid-tier model is, is, is super, is like great anyway. Um, and, and it's actually better than the large scale up model that you did previously. Uh, and I think you, I think this is also a little bit of a reflection of the reinforcement learning, um, paradigm where you can, you can take a model and you can train it, uh, and, you know, it is, uh, and extend it with reinforcement learning, basically. So that allows you to take a, a mid-tier model and make it as good as a larger-tier model of, of six months ago or three months ago.

AI assessment note: “Opus, which is the smartest model, Sonnet, which is, uh, the mid-tier model, and Haiku”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yeah. Let's unpack the 30 hour aspect, which is, which is fascinating. So, uh, first of all, to, to just ground it, um, for people. So this is a computer use, uh, just coding, I think. Just, just, just, um, coding. So what does the agent do for 30 hours? It's just clicking on stuff?

A Yeah. Uh, it is there. It's, it's reading files, and it's writing code, um, and running tests. Uh, so what, in exactly the same way that a human would, um, it, uh, Basically, you can think of the model as running in a loop where it can constantly decide what to do. People often mention something called tool use, and tool use is the ability to, uh, well, I mean, it's in the name, but in this case, it can use things like tools, like read file, write file, et cetera, um, or run code in the terminal, uh, and it is sitting there in a terminal on a computer in a loop, just constantly looking at the current code, deciding, oh, well, It can't quite do this yet, so I'm going to work on that next. It's often making plans, um, particularly to run for, uh, you know, 30 hours. One of the things that we're pretty happy with about the recent launches, we've finally taught the models to, uh, to use what's called memory. Um, and so, and we've built that into the agentic, uh, harness. So it's able to create a markdown file of to-dos and things that it thinks are important to do, uh, check them off and work on them and check whether they've been completed. There's almost this, like, Self-verification loop. One of the things that people were worried about, um, with language models over, like, I think a year ago or so, was that they would fall off track. Like, they wouldn't be able to self-correct,…

AI assessment note: “It's reading files, and it's writing code, um, and running tests.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, so maybe let's talk about, uh, progress at a more abstract level, uh, but grounded in, in, in, in 25. So, a big part of the discussion seems to have been the evolution from a focus on pre-training, Uh, to RL, which we touched upon a couple of times. Talk about, uh, the impact of RL, and why is RL such a big part of the conversation today?

A For those listening, a good way to, like, understand at a high level of pre-training in RL, pre-training is like skim reading every textbook in existence, and RL is like doing the work problems and getting feedback on whether you were wrong or right. And there are actually a lot of things that you can only learn via RL. And a good example of this is the skill to say, I don't know, in response to a question. Because in pre-training, remember, you're modeling the, you know, you're trying to predict what, what text is going to come next in the, you know, all of these, you know, textbooks, the entire internet in the world. Um, and so, The only reason you would say, I don't know, on, you know, as a pre-trained model, is if you think the character that you're modeling in the text would say, I don't know. Like, if it's a likely completion, right? Not whether you, in fact, don't know, but whether you think that the, like, the sort of player that you've pulled from this cast of characters that you could model, um, would say, I don't know. Whereas in reinforcement learning, you could, in theory, set up a battery of tests where there are things the model knows and things the model doesn't know, and you could reward it for correctly answering, uh, things it should know and penalize it for, uh, for, for falsely answering when it doesn't know. And what it will then learn to do is it will lea…

AI assessment note: “there are actually a lot of things that you can only learn via RL.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay. How do you think about keeping a consistent user experience in a world precisely where those models, one, evolve all the time, two, behave differently, and three, as we all know, are stochastic, not deterministic. Do you expect the users, as you just said, to like be Be smart about it, and that's sort of theirs to figure out, or is that something that you abstract away for them?

A We abstracted away from, from people like the, the general design. At first we didn't, for the longest time, we didn't let users choose their model. Right. And we only let users choose their model on chat. On the note generation side, we completely abstract it away. And, and the reason we do that is every time a new model comes out, we have to completely change or tweak the prompts that we use for, for note generation to provide consistency of experience and an improvement of experience. And there's significant work that goes into that. Um, it's, I think one of, it's one of the value adds that Granola brings as opposed to just working with base models is that we take care of that and we make sure you get, Granola feeling or sounding notes consistently, um, and that they keep getting better over time.

AI assessment note: “On the note generation side, we completely abstract it away.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q on that topic that, uh, caught my attention, not having the bot first experience was actually a trade off in terms of like virality because you don't have the built in, um, expo product exposure because there's no bot. Showing up. Uh, so what have you done to, um, uh, sort of overcome that? And what are the, you know, viral sort of growth mechanics built into the product today?

A We haven't focused on growth. Basically what we really focus on is making the product really good for people. And turns out that that's actually, uh, led to a lot of viral growth, but that viral growth is from people telling each other. Uh, like, so like an interesting story that this, uh, This is, I never imagined that this could have happened, but I hear a lot now is, um, if you, if you have a one-on like, you know, you're meeting with someone on a zoom call and your AI bot shows up and you basically are told like, Hey, what are you, what are you doing with an AI bot? Like, why aren't you on granola yet? You know? And it's like, Oh, wow. The AI bot is now like a conversation start. Yeah. It's the weird thing. It's a conversation starter for a human to bring up granola. And to vouch for it, which is incredible. I never would have sat down and imagined that, that world. Um, so we always start from like a, like a value standpoint, like what is valuable to the users? Like, oh, we, we could email your notes to everybody in the meeting, like all the other companies do. But again, is that, is, is that a tool that you want to use? You know, is that acting like a tool for you or is that acting like, You know, like a growth engine. I, what, what we do have is, um, we do let people share granola notes. Basically, you can share notes on a link, and you can send that link to people, and w…

AI assessment note: “we do let people share granola notes. Basically, you can share notes on a link”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So a little bit on that note, uh, to, to close and, and, and, and zoom out, anything that you can talk about in terms of the roadmap or the futures? You mentioned a couple of times the idea of, um, searching through the history of, of meetings. So that's one thing, maybe double click on that, uh, and, or anything else that, uh, you can talk about.

A Yeah, absolutely. So I, I think that the world we're moving towards is, um, You have a bucket of context, and then you, you generate documents, um, or artifacts on, on a per need basis on the fly. And things, the types of things that we're working on are, uh, given my entire history of meetings, can you pull out really? So for example, a good question, um, that you could ask could be, um, that we could ask could be like, okay, out of everyone we've met in the last two years, like who are the firms who are most likely Uh, candidates to lead our series C, right? That's a question that I can't ask anywhere else in the world right now, but I have a version of granola that will go through my 2500 meetings and spit out a remarkably intelligent answer to that in 20 seconds. It's like a deep research mode, right? Um, the other thing that we've played around with is like, if you're, if you're dynamically generating Artifacts or UIs on the fly. Can those be shared, right? So this idea of we have a folder where all our sales calls, um, get put into and they're shared within the company. And then we have this artifact that you go to the URL. It doesn't have to be in the granola app. And it'll tell you, here are the most important things our, um, enterprise customers are telling us like today. And every time you reload it, it's up to date, but it's like a, it's, it's like a memo. Um, so the…

AI assessment note: “Granola that will go through my 2500 meetings and spit out a remarkably intelligent answer”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And, uh, how does the MCP, uh, part work? So MCP, the model context protocol being something that, that you guys at Anthropic, uh, defined, how does that work? You just leverage the, the protocol to connect to any tool?

A Yeah, exactly. Quad code is a MCP client and an MCP server. And what this means is if you give it tools to use, so maybe at your company, you have a bunch of MCP tools that you build for kind of all your systems to integrate with. Like I said, maybe there's one to integrate with Jira, maybe there's another one to read Slack, and maybe write messages to Slack. Maybe there's another one to, um, you know, fetch some internal knowledge base or something like this. You're going to plug this into a bunch of your tools. You can plug it into Quad AI, into Quad Desktop. You can also plug it into Quad Code. So it gets all the same tools that you do.

AI assessment note: “Quad code is a MCP client and an MCP server.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Fascinating. And do you have to declaratively add to the memory, or whether today or in the future, the memory will automatically pull from the context and sort of improve itself?

A Yeah, you have to add to it manually today. We, we've actually had a bunch of internal experiments to do automatic memory so Quad can automatically remember things. The problem is there's kind of two ways in which it fails. One is that it remembers things that it shouldn't. So for example, if I say make the button blue, it might remember the user always wants the button to be blue. And this is, you know, maybe that's the case for this button. That's not the case for every button. And then sometimes it doesn't remember very important things that it should remember. And so for the last few months, we've been doing a lot of experiments to try to get this performance really good. And it's something we've been using internally. And at some point when we're happy with it, it's something we're, we're going to release for everyone. Um, but generally our bar is if we find ourselves really happy with it and we find ourselves using it every day, then we release it to everyone. And this one's not quite there yet.

AI assessment note: “Yeah, you have to add to it manually today.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And then the, the key question for, uh, agents is always, uh, the, the level of autonomy, uh, versus, uh, human in the loop. How, how do you guys think about that and where do I get pinged as the human coder, uh, in my cloud code workflow?

A The default behavior is there's always a human in the loop. This is, this is super important because this is, in the, in the end, this is a model and it's not predictable and, You want to make sure that it doesn't do anything dangerous. Um, so yeah, there's, there's always a human loop. So for actions that we know can't have any, um, kind of dangerous repercussions. So for example, reading a file, we know this is inherently safe. We just let the model do this in the folder that you let it do this in. But for other actions like, uh, editing a file or running a command or using the internet, this always needs a human in the loop and it always needs a human to approve it. There's ways to reduce this burden a little bit. So, for example, if you find yourself always approving edits to the same file, or always approving the same command, there's a, there's a settings file that you can configure across your team, and you can use this to essentially allow list or block list certain commands, or certain files that you always want the model to be able to edit without human approval, or you never want it to be able to run.

AI assessment note: “editing a file or running a command or using the internet, this always needs a human”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So in the, in the same vein, what do you find yourself, uh, using cloud code for as a, as, as a leader and engineer, what's your, uh, you know, daily use case?

A Yeah, I use it all day for all, all sorts of stuff. The, so obviously, like I said, code-based research, If I'm working on a piece of code I'm not familiar with, I'll just start by asking quad code to tell me about it. Whenever I'm working on a small feature, I'll usually use quad code in GitHub Actions. So I'll just say add quad, I'll make a new GitHub issue, and then I'll say add quad, implement this feature for me. And it'll just do it usually in one shot. Um, and sometimes I'll do this on the command line too. So I'll just say, you know, implement this feature and make a pull request and, you know, I'll come back a few minutes later and it's done. Then there's this kind of other work where it's a little bit more complex. You can't really do it in one shot. It's, it's not as simple as changing a piece of text or changing a button or building a small feature. Maybe it's like something more involved. There's probably two workflows I have here. One is for really complex stuff. I'll prototype it a bunch. And this is something that I did even before a quad code. You know, when you, when you write a complex piece of code or a complex feature, often engineers will write it a few times. Because you don't actually know the right way to do it, and so you'll try one approach, you'll try a second approach, you'll try a third approach, and you'll kind of figure out the edge cases for eac…

AI assessment note: “Yeah, I use it all day for all, all sorts of stuff.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, great. So that's, uh, the actions part of the agentic workflow. Uh, let's talk about the, uh, awareness and memory of this. One of the exciting features is that, uh, Cloud Code can, can connect with the existing sort of code knowledge, uh, in, uh, the company. How does that work?

A There's a few different ways to pull in context and this kind of knowledge from, from the company. The simplest one is just looking at files. There was this, ah, there's a, there's a few different approaches actually to reading files. So I'll, I'll go a little bit into depth into the way that actually happens. In the past, the thing that people use the most is this thing called RAG, and essentially this is a technique where you take the whole code base, and this actually works for any document, set of documents, it's not necessarily code, but you take a set of documents like, uh, like all the files in the code base, you do this kind of indexing step, and then you store essentially this database of all the knowledge that's in these files in a very, very particular form that makes it really easy for the model to search. There's a lot of trade-offs to doing this. The indexing takes time. It's pretty expensive to maintain this database. It's quite tricky practically to make sure that security is really good and privacy is really good because, uh, it's just a, it's, it's a very sensitive information like your code base. And so you want to keep it really safe. And so quad code actually doesn't use this technique called rag. Instead, what it does is it just searches files the same way that a human would. You can think of it like, um, you know, at the engineering level, it uses the too…

AI assessment note: “Instead, what it does is it just searches files the same way that a human would.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And as a quick segue on that note, like, it was super interesting to see the Nobel Prizes a few months ago, or maybe that was last year at this point, like, everything converting towards AI. Do you think AI is eating all those, uh, other scientific fields?

A I think it's augmenting, like it's, um, and in some sense, I think that part of what happened was that AI's impact in the world had clearly become, ah, hard to ignore, but there's no Nobel Prize for computer science, which is a field that, you know, so, yeah, the Turing Awards have gone to AI for, you know, a number of years now, and so I think that the Nobel Committee, I'd imagine, felt like it needed to somehow Shoehorn this, uh, and so obviously there were like very impactful, uh, AI breakthroughs with AlphaFold that, um, resulted in a prize. Um, but what's interesting is that the Physics Nobel Prize was given to something that has not really had that much impact in physics, but it is, um, but I still buy it because it's kind of, uh, there's a physics smell to the breakthroughs that led to, uh, you know, these systems like called Boltzmann machines and Hopfield networks that, uh, That Jeff Hinton and Hopfield got the prize for. Um, they're very physics-y, and they, they, they look, they look like the same exact objects that physicists, physicists, physicists study.

AI assessment note: “I think it's augmenting, like it's, um, and in some sense”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And the next step after that was DeepMind?

A And then I went to DeepMind. At the time, I was really interested in- In Toronto? Yeah, Toronto and New York. Um, so I joined the group of, um, this researcher named Vlad Mani, who was, uh, largely credited with starting the field of DeepRL. He was the first author of the DeepQ Networks paper, which was the paper that got, um, neural networks to play Atari. And he, like, his first set of papers actually largely defined deep reinforcement learning as, uh, as a field then. Were basically DeepMind's Claim to fame for a very long time. Uh, and so I joined this group to study the problem of what we called, I mean, we started, we, the team we built together is called the general agents team, and so the whole point was to do research that enabled, you know, for us to figure out how do we build general agents. I think that, um, it was much more opaque, I think, then than now, and the big problem we were trying to solve is, um, what People called, and still do, but it, uh, called unsupervised reinforcement learning, which is really, how do you train reinforcement learning systems that are capable of, um, assigning their own rewards? Like if you don't have rewards without supervision, like in the same way that, you know, you can give things some rewards, but kids and animals, when you look at them, they, they learn a lot in an unsupervised way. Like they interact with their environments …

AI assessment note: “And then I went to DeepMind.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q unique vantage point, uh, into, uh, what companies do and, um, their growth and all the things, and you release from time to time, uh, really interesting stats, and perhaps we'll put some of those, uh, as a link in the, in the show notes. Uh, To start at a high level, what do you see, uh, that's, um, different in this generation of AI companies from your vantage point?

A One of the things, you know, from my, from my economist hat that I love about working here is just kind of this front row seat to, hey, what's the growth trajectory of each successive wave of startups in particular? And, you know, the current wave of course is AI. Um, we work with AI companies across the stack. So when I talk about the AI companies, uh, on Stripe, you should think of this as everything from like infrastructure and modeling to full-blown applications, open AI, Anthropics, you know, perplexity, cognition, 11 Labs. Um, Decagon, Sierra, right? And like a long tail of, of others. Um, we, we recently looked at the, the Forbes AI-Fifty and 78% of them are Stripe users. That 78% reflects 100% of the Forbes AI-Fifty that accept online payments. Um, and you know, I think there's, there's a lot of hype around AI tech and I think fair questions around the monetization. And so we took a look at, hey, with this current wave of AI startups, What do we see in, in their monetization trends and, and in their growth trajectories? The long and short of it is like they are monetizing super fast. They are monetizing faster than any previous generation of startups that we've seen. Um, we focused in, um, just for concreteness on the top hundred highest grossing AI companies on Stripe. Um, and we asked, okay, for the median in that cohort, how long did it take them to hit various reven…

AI assessment note: “They are monetizing faster than any previous generation of startups that we've seen.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q All right. So everyone, uh, in tech obviously knows Stripe, uh, which is a monster of a company, but, um, maybe for context, uh, what is the latest and greatest way of describing the full breadth of what the company does and maybe the latest stats?

A Well, Stripe builds programmable financial infrastructure. So put kind of less buzzwordy, we are giving any business, whether it's A twenty-year-old selling a Figma template, or, you know, now more than half of the Fortune 100, um, the rails and the intelligence to move money online, um, and to grow faster. Uh, you asked about the numbers. Last year, companies processed about 1.4 trillion dollars, uh, on Stripe. Um, to put that in perspective, because it's a lot of zeros, that's about 1.3% of Global GDP. Um, and that number grew 38% year over year, uh, in what many experienced as kind of a rocky macro climate. Um, Stripe's network handles on average about 50,000 new transactions every minute. So those are the transactions that are adding up to, uh, 1.4 trillion in, in payments volumes processed annually. Um, and every one of those transactions is training data for some of the AI systems Um, that we will, that we will talk about today. I just say because of the flywheel, like Stripe is no longer the payments API. If we were talking 10 years ago, we'd be talking about a payments company, but in practice, we're optimizing now the entire payments life cycle, the checkout, user experience, fraud prevention, bank routing, automatic card update retries, even how you handle disputes as a business. Um, and that's all in service of Merchants profits, right? Growing their revenue, um, and…

AI assessment note: “Last year, companies processed about 1.4 trillion dollars, uh, on Stripe.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q In particular around the response time, idle time. Okay. Um, what about sandbox?

A Yeah. So one of the fascinating things in the world of agents is that fluid needs to run the code that humans wrote to power your applications. But what's happening is that increasingly more and more code is being created by agents on demand. A good example is when you're working on things like deep research. Let's say you're a big pharmaceutical company or a bank, and you want to create your own deep research system. So first of all, you're going to need this sort of like high efficiency compute, like fluid. You're going to need obviously some of the foundational infrastructure we talked about, like CDNs, firewalls, et cetera. But now you need a new thing that actually quite doesn't exist in the world, which is you need a sandbox, meaning a, a secure place for compute to run that is generated by the model. So when agents are doing research, they might decide, Hey, I need to run a quick Python script to compute some numbers or to produce a visualization of the data or to make up my mind or whatever. So you can think of sandbox as the, uh, Amazon easy two of AI. It's not for the code that your developers wrote is for the code that gets emitted by the LMS and it allows you to create this incredible new products. I mean, you could create your own V zero and lovable. With this sandbox primitive, but also you can create systems where the code runs behind the scenes in the service of…

AI assessment note: “you need a sandbox, meaning a, a secure place for compute to run”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yes. Uh, so you launched it in the, in the fall of, of 24 and you've been on a tear since, uh, are there any, uh, metrics or any qualitative information, uh, you can share to help us get a sense for traction?

A Yeah, it's pretty insane. We've passed over a hundred million generations of applications. For those that don't know, VZero allows you to convert text to application, idea to application. It speaks to an audience that Vercel has typically not spoken to. It's basically everybody with a job or even everybody without a job. Whereas Vercel required any new engineering skills. This requires that you have an idea. I think it's the embodiment of what people have been calling vibe coding. And to give you an idea of the difference between the two personas, every single second, there is seven app generations happening on V zero. VZero has more than doubled the entire user base of Vercel, and Vercel has been around for almost 10 years. Uh, VZero has doubled our number of users in less than a year. And so what's happening really is that coding is being automated. More people can code with the ease that you could type into chat GPT. And I think that's creating more software. Some people call it personal software or hyper specialized applications. But the two things that are driving this is one, I call it. Everybody can cook. You and I have an idea. We can start bringing it to life. As individuals, right? Like during the weekend, maybe with our kids. The other thing that's happening is that within companies, teams, and enterprises, people have the same fundamental need. They need to prototyp…

AI assessment note: “We've passed over a hundred million generations of applications.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So that's very interesting. So it's 200, uh, machine learning slash AI folks, um, and so you're saying there is a central lab that does R&D research?

A Yeah, so we take a hybrid approach there. We certainly have, uh, teams that are dedicated to more foundational research, but we also have embedded, um, uh, folks embedded in product teams who are doing research. And, uh, they tend to be on shorter product cycles. Uh, and they tend to, you know, their output is usually a prototype, which a product will look at and make judgments about, and then we'll see if we want to turn that into something that's a production feature, and we'll try to turn those around very, very quickly. We've invested enormously in putting a platform abstraction over a lot of the third party AI that we use, uh, whether that's kind of like base layer infrastructure, Uh, like Bedrock or SageMaker, or whether it's third party AI providers like OpenAI, and that allows us to rapidly switch out models and experiment with different models. It allows us to augment models with our own proprietary tech, uh, and it allows us to really iterate very, very quickly when new models come onto the market to do something exciting. We can get them in front of users In really cohesive experiences very, very quickly. So we were able to, for instance, you know, Canva code is, is an interesting feature that we've just, uh, launched just back in April. That's a feature that allows users to build little intelligent, um, interactive widgets in their designs. That went from kind of fi…

AI assessment note: “Yeah, so we take a hybrid approach there. We certainly have, uh, teams that are dedicated”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q number of companies that we all work with as, as VCs, early stage, um, startups that, uh, at some point graduate from that, uh, Early, uh, kind of PLG inbound motion to being outbound and enterprise focused. From an engineering standpoint and product standpoint, any sort of tips and tricks about, um, what took longer than you thought or was harder than you, you would have thought, would have imagined?

A Catering to an enterprise market, um, it certainly pulls the product teams in different directions. Um, enterprise customers want sophisticated admin controls. They want sophisticated auth. They want data residency. These kinds of considerations, if you haven't baked them into your product early, they become quite expensive to, to retrofit. Very early in Canva's, um, architectural journey, we put a, a quite good abstraction in over our, um, AWS infrastructure, but I do think back now and, and, and wonder, oof, we could have just spent a little bit more time Really abstraction, abstracting region-based storage. Uh, it would have been cheap to do then. Uh, we have got it now, but it was a, a massive engineering effort. Um, you know, back when it was 10 of us and a few services and a few databases, uh, it would have been relatively cheap to kind of build that culture in there, build that kind of tax, if you like, on every product that you have to think about what, what region shard You're gonna store the data in. Uh, having to retrofit it across hundreds of teams, um, hundreds of services was a, was a monumental lift.

AI assessment note: “having to retrofit it across hundreds of teams, hundreds of services was a monumental lift”

← previous page 3 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.