The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

640exchanges match
640on raw tape
31redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Like, uh, how do you think that changes?

A I think, um, so I'm actually bullish on engineers in terms of their kind of long-term economic value. Um, Not despite all the movements in Cogen and all the things that we're, you know, already seeing, but because of it, uh, because what's going to happen as a result of AI, and people have talked about this, um, in, um, even other disciplines, we're going to be able to solve many more problems. So my math guy in me is like, okay, so we always say, oh, well now, you know, agents are going to be doing code or whatever. And so there's going to be a million software engineer, uh, you know, virtual digital software engineers out there. And so the value per engineer is going to go down because I'm just in that, that same mix I as an engineer. What they don't recognize is that it's not just about the denominator. There's a numerator as well, which is what's a total economic value that's possible. And I would argue that's growing faster than the kind of denominator is that the actual economic value that's possible as a result of software and what engineers can produce, you know, with the tools that they will have at hand. Um, so I think the value of an engineer actually goes up. They're gonna have the power tools are gonna be able to solve a larger base of problems that are gonna need to be solved.

AI assessment note: “I'm actually bullish on engineers in terms of their kind of long-term economic value.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. I think like something I struggled with, with this conviction, you said you pursue things to conviction, but like, You start out not knowing anything. Yeah. And so how do you develop a conviction when there's, you, you find it along the way where you stumble along the way, then you lose conviction and then you stop working on it. You know, like how do you keep going?

A The way I've sort of approached it is that, um, so I don't generally tend to have conviction around a solution or a product. I have conviction around a problem, uh, that says this is an actual real problem. That needs to be solved. And I may have an idea for how to be solved, uh, you know, right now and that I may be get dissuaded. It's like, ah, I'm not smart enough. Technology's not good enough, whatever the constraints are, but it's the problem I have conviction around. It's like, oh, that problem still hasn't gone away. Uh, so like I sort of filed away in the back of my brain and I'll revisit it's like, okay, well, you know, the kind of board changes, uh, and it changes really fast now with AI, like things that weren't Possible before are now possible. So you kind of go back to your roster of things that you believe or believed and say, maybe now, uh, now is the time. Maybe then wasn't the time. Uh, but I'm a big believer in kind of attaching yourself passionately with conviction to problems that matter. Um, that, and there are some that are just too highfalutin for me that I'm not going to ever be able to kind of take on. I have the humility to recognize that.

AI assessment note: “I don't generally tend to have conviction around a solution... I have conviction around a problem”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q in any reasonable quantity. Like, you know, you know, Claude 3.5 Opus, like if it, if it does exist, still not like, you know, the, the thing that we actually use is Sonnet, right? So like, it's almost like a deployment strategy. Like you, you train the large model, the, the teacher model in order to distill. Like, you don't actually expect to use the teacher model for inference anymore.

A Yeah, I think it is really the, the marginal benefit versus cost argument, right? Like, if the marginal benefit and capability is not worth the additional cost to you, then people don't want to use that bigger model. And it is very, very task dependent. Maybe you really want that capability, but sometimes, especially when you're doing things at bulk, maybe it's okay that you're slightly worse off, but it's much cheaper. And that's maybe one of the reasons, at least this graph, that's why you want to be at the Pareto frontier. And I think the other thing is the capability which we have right now, maybe next year will be much cheaper to have that same thing. And that likely is the result of distillation, right? Like that's, that's, it's not just because we are doing or figured out something magical. It is because distillation can help you bring the cost of any existing capability.

AI assessment note: “distillation can help you bring the cost of any existing capability”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. Um, so, you know, I think the, this is interesting. Um, the, the, how do you compare it versus the other benchmarks that you see out there, like the suite bench. And then, you know, more recently now people are starting to compare benchmarks with The, like, like actual money-making projects, like Sui Lancer, something that OpenPlay recently released. Yeah, just any thoughts on the meta game of benchmarking?

A Yeah, I mean, I think, like, the This is, you know, I guess like in terms of experimental design, right? Like the Sweebench category of like starting from real GitHub data and like working towards that does feel like a really nice like end-to-end test. It's almost kind of like in my mind, you know, feels like the same type of test as like an integration test. This test is definitely a lot more targeted, right? It's like saying for this very specific skill of Taking a front end app and taking a specification of it. Can a agent then perform in this way that is easy to grade, right? It's like easy to say that given a front end app and given a back end implementation, like just, you know, test the functionality that it works because then also the front end is kept standard, right? So I think like, you know, the loss of generality here is in my mind, it's the trade off is that versus like, uh, power to grade, right? Like here we can have a, without having a gold standard of what the answer should be, we still can, uh, evaluate the solutions quite effectively. Right. Um, so I think like exploring different points on that trade off space of like generality versus like focused and Easy to grade is something that we're pretty interested in. We have another suite of evals for, um, we just call them convex evals that are more coding, like kind of small units of code focused, and they're j…

AI assessment note: “the Sweebench category of like starting from real GitHub data... feels like a really nice like end-to-end test”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q just my, uh, I think we were talking about, and then if you, if you don't have anything in this feature, we can, we can, uh, stop it as well. But, uh, it's more just about like, uh, I guess Running Convex in, in the, in the age of AI, right? Like, like, I think this is something that you were saying in pre-chat about building developer tooling in general.

A Yeah, totally. So we've been at Convex since 20, 21 and, you know, a long, no time we've like, especially because we're kind of in this, like think from first principles, like redesign things that have been the same for a while. We've, you know, thought a lot about ergonomics for humans and like what helps people who may not have a lot of backend experience write really effective Convex code. And I think, you know, in the past year and certainly in the past few months, it's just been very apparent that You know, people are writing code really differently and, you know, if someone's a dev tool startup and they're writing something that is like in some set of APIs or systems that are targets for AI coding, they, you know, it's silly not to strongly consider how AI coders perform on one's platform. So that it was like the, you know, motivations for kind of investing in making some of this analysis a bit more rigorous. And then also, yeah, like, you know, users kept on telling us they're like, Hey, like Convex works really well for this, but I noticed that there's some knowledge gaps and there's like consistent hallucinations. And I think shaping one's platform to be a great coding target is like, you know, something that's just table stakes in 20, 25.

AI assessment note: “shaping one's platform to be a great coding target is like, you know, something that's just table stakes”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay, okay. Let's talk about Snit. What is Snipt? And, you know, then we'll talk about your origin story, but I just, let's, let's get a crisp. What is Snipt?

A Yeah. I always see two definitions of Snipt. So I'll give you one really simple, straightforward one, and then a second more nuanced, um, which I think will be valuable for the rest of our conversation. So the most simple one is just to say, look, we are an AI powered podcast app. So if you listen to podcasts, we're now providing this AI enhanced experience. But if you look at the more nuanced, uh, perspective, It's actually, we, we have a very big focus on people who, like your audience, who listen to podcasts to learn something new. Like your audience, you want, they want to learn about AI, what's happening, what's, what's, what's the latest research, what's going on. And we want to provide a, a spoken audio platform where you can do that most effectively. And AI is basically the way that we can achieve that.

AI assessment note: “we are an AI powered podcast app”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q you, they had, uh, from Twitter about how ChatGPT interleaves code of text, and it was basically just imitation learning on contractor data. Is that what you're seeing? I think, like, as you mentioned, like, a lot of approaches typically end up with some version of self-play as well that scales up yourself. How, how are you approaching your data, or like, what are your strong beliefs on your data?

A Well, you kind of, yeah, you kind of have to have both. So it's, it's really important to both get the right initial data mixture for, from supervised fine-tuning, however it is that you gather it, that feeds the correct behaviors that you want to elicit out of the model. And then you kind of amplify those behaviors with reinforcement learning. So the key thing to reinforcement learning is that it only works. I mean, It can potentially work otherwise, but practically it only works when the agent has interacted with a reward, right? It's received a positive reward for what it's done. Maybe one out of 10 times, one out of 50 times, but if it's getting zero reward, then you don't really know what behavior to amplify. And so that's why it's really important in the kind of initial mixture for it to already be doing things out of the box that are somewhat sensible, even if they're unreliable.

AI assessment note: “Well, you kind of, yeah, you kind of have to have both.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q kind of model access. Are you going to, uh, you know, like poolside AI is in a similar sort of category of code-focused model labs, and they focus on their product being more of like a VS code extension that people can use to access their models. Like, what is the shape of the product's thinking that you have so far? You know, I know you're not launching it yet.

A Yeah, so the kinds of problems we want to solve are maybe first just thinking about what, there are multiple form factors for a coding product, and the form factor that is most common today is, you know, this kind of cruise control form factor, which is the Copilot. Github Copilot or Cursor, which are incredible products. We use them internally. We're very happy with them. But again, this is, these are products where the engineer is driving most of the work. Like, even in agent mode, really the engineer is, like, driving most of the work, stepping the agent through, and so forth, and really supervising it. Uh, what we're building is more of the autonomous vehicle. So, right, the thing that you give it a task and it takes you from point A to point B. The form factor for that is, uh, is a bit different. I mean, one way to instantiate that form factor is through an IDE. Like, that's, that's definitely something that people have been doing. But another way is to provide access to it as an API. And what I mean by API is different than an API that just streams tokens, right? Because if you need a, you need a model that's kind of coupled to a computer, right? That it's able to read, write, and run code. So it's kind of, um, an API that rather than taking in tokens and outputting tokens, it takes in a task and a code base, and it outputs, it does some work on its own, the background, a…

AI assessment note: “what we're building is more of the autonomous vehicle”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Like sometimes it gets confused. Who is the actual playable character in the, in the scene? Like how, how do you steer that?

A Uh, yeah, I think, like, sometimes it gets confused, uh, it can be applied to many things in quad-playing Pokémon. In particular, when it's trying to, like, look at the screen and understand what's going on. So... I have, like, attempted to prompt it all sorts of ways. Like, you are at this exact coordinate, and you're in the middle of the screen, and you're wearing a red hat, and things like that. And, like, that's all neat, but Quad doesn't particularly understand, like, the middle of a Game Boy screen and a whole bunch of concepts like that, which means, like, you can prompt all around everywhere, but, like, this kind of, like, spatial awareness and where something is with respect to something else is something that Quad's still just, like, not great at in its current incarnation, so. One of the side effects is that sometimes loses track of who it is on the screen and thinks there's something else there. I'll keep trucking through this, so. I hinted at this, like, other tool that I give it called Navigator, and this is just, like, the only other patch that I have for the, the vision issue. So Navigator, basically what it does is, like, Claude can say it wants to go to one of these coordinates, uh, that we provide in the screenshot, and then we, like, automatically press the buttons to get there. Uh, it has to be something on the screen, like, I'm not trying to let Claude jus…

AI assessment note: “I hinted at this, like, other tool that I give it called Navigator”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q wow. It was, like, very dramatic, and I was talking about the game. How, I, yeah, is there any thought being put into, like, trying to have it more, like, do you prompt it to be more rational, to let it know that it's not a real life, that it's a game? It's like, it feels like it gets very distressed when they're actually, the Pokemons are actually gonna die.

A It's funny. They, um, it knows it's Pokemon. It's like, you're playing Pokemon Red, like, it does know that, and it has a sense of that, but it clearly gives us some attachment. I'll tell you a fun story. We tell it to nickname its Pokemon now. It will occasionally do without it, but it's, like, more fun if it nicknames its Pokemon. So that's, like, in the prompt is, like, it's fun if you nickname Pokemon, you should consider it. And one thing we found when we started doing that is it got more protective of the Pokemon it nicknamed. Like, it's pretty obvious, like, when it catches a Pokemon, Now that it has a nickname, it will, like, go heal it right away if it's hurt, and that did not ever happen before, which is pretty, like, so there, there's some cute little things, cute quirks about Quad who really wants to protect its, uh, precious nicknamed Pokemon, which is great.

AI assessment note: “it knows it's Pokemon. It's like, you're playing Pokemon Red”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And you're basically transpiling, so to speak, all of these extensions to work with every model. Uh, I think open AI and entropic are pretty kind of like, you know, backward compatible, so to speak, but did you have any issues with some of the, Alternative models that you offer, or?

A Definitely. Like, I feel like every model has their little nuances. Finish reasons are different. How they're calling tools are a bit different. Behaviors are a bit different. Formatting is a bit different. So there was, like, one thing which we basically had to go over all of the different models and patch them. And we're not using, like, a proxy, like a router or something like this, so we built that internally. But like, yeah, that's, that's one of the knally things you have to do. Like, even though all models say they support open AI compatible APIs, they're always in nuances which are slightly different and then put you off when you see it for the first time.

AI assessment note: “Definitely. Like, I feel like every model has their little nuances. Finish reasons are different.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Um, do we want to start a query so that it runs in the meantime and then we can chat over it?

A Okay, here's one query that, that we, like, we love to test, like, super niche random things, like, things where there's, like, No Wikipedia page already about this topic or something like that, right? Because that's where you'll see the most lift from, from a feature like this. So for this one, I've come, I've come, come up with a square. Uh, this is actually Morgan Square that he's, he loves to test is help me understand how milk and meat regulations differ between the U S and Europe. What's nice is the first step is actually where it puts together a research plan that you can review. And so this is sort of, it's Guide for how it's going to go about and carry out the research, right? And so this was like a pretty decently well-specified query, but like, let's say you came to Gemini and were like, tell me about batteries, right? That query, you could mean so many different things. You might want to know about the like latest innovations in battery tech. You might want to know about like a specific type of battery chemistry. And if we're going to spend like five to even 10 minutes researching something, we want to one, understand what exactly are you trying to Accomplish here. And two, give you an opportunity, like, to steer where the research goes, right? Because, like, if you had an intern and you asked them this question, the first thing they'd do is ask you, like, a bunch o…

AI assessment note: “Okay, here's one query that, that we, like, we love to test”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What's the initial ranking of the websites? So when you first started it, there were 36. How do you decide where to start? Since it sounds like, you know, the initial websites kind of carry a lot of weight too, because then they inform the following.

A Yes. So what happens in the initial turns, again, this is not like a, it's not something we enforce. It's mostly the model making these choices. But typically we see the model exploring all the different aspects in the, in the research plan that was presented. So we kind of get like a breadth first idea of what are the different topics to explore. And in terms of which ones to double click on, I think it really comes down to every time you search, the model gets some idea of what the page is. And then depending on what pieces of it, sometimes there's inconsistency, sometimes there's just like partial information. Those are the ones that double clicks on. And, uh, yeah, you can continually like iteratively search and, and, and browse until it feels like it's done.

AI assessment note: “it's not something we enforce. It's mostly the model making these choices.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I think we all, we all get that. I was more curious, like from a product perspective, when you decide to do rag versus shit like this, you didn't need to, You know, do you get better performance just putting everything in context or?

A The tricky thing for RAG, it really works well because a lot of these things are doing like cosine distance, like a dot product kind of a thing, and that kind of gets challenging when your query side has multiple different attributes. Uh, the dot product doesn't really work as well. I would say, at least for me, that's, that's my guiding principle on Uh, when to avoid drag, that's one. The second one is, I think, every generation of these models are, uh, like, the initial generations, even though they offered, like, long context, their performance as the context kept growing was, you would see some kind of a decline, but I think, uh, as the newer generation models came out, uh, they were really good even if you kept filling in the context in being able to piece out, uh, like, these really fine-term information. So I think these two, at least for me, are like guiding principles on Movento.

AI assessment note: “at least for me, that's, that's my guiding principle on Uh, when to avoid drag”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And so from a user perspective, is it better to just start a new research instead of like extending the context?

A Yeah. I think that's a good question. I think if it's a related topic, I think there's benefit to continue with this thread, uh, because you could The model, since it has this in memory, could figure out, oh, I've found this niche thing about, I don't know, milk regulation in this case. In the US, let me check if your, you know, follow-up country or place also has something like that. So these kind of things you might have not caught if you started a new thread. So I think it really depends on, on the use case. If there's a natural progression and you feel like this is like part of one cohesive kind of a project, you should just continue using it. My follow-up tone is going to be like, oh, I'm just going to look for summer camps or something. Then, yeah, I don't think it should make a difference, but we haven't really, uh, you know, pushed that to see, uh, and, and, and tested that, that aspect of it for us. Most of our tests are like more natural transitions.

AI assessment note: “if it's a related topic, I think there's benefit to continue with this thread”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Um, what was your experience then? Did you like the model right away? Like, did you kind of, uh, you know, did it grow new?

A Yeah, so I think my experience with O-one, where it's been different from other models for me, is that O-one is actually the first model where I'm getting more impressed by it the more I use it. So, like, when ChatGPT first came out, right, I think it was GPT-III. And at first it seemed like, oh wow, this is amazing, it can actually create text that sounds like a human, and it can Write poems, but kind of as you started to use it more, you would discover more of the things that, where, you know, there was error cases, or there's areas where it couldn't, it wasn't capable. And with, with O-one, it's been, it's been the opposite experience. So the more I've used it, the more impressed I become by it, and the more I realize it can do. So it's kind of like, by the nature of it is so capable, it's sort of been like pulling me towards using it more.

AI assessment note: “O-one is actually the first model where I'm getting more impressed by it”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q know, one of the things they said was we haven't seen anything that's kind of eyeopening to see people going to 200 dollar tier on this sort of thing. Haven't seen anything else like that in the space. Cause I, I think this is very new because of the new model capabilities, right? Where people, you know, it makes sense. Like you're willing to pay more money for this stuff.

A So this is something I've talked about before in terms of matching the dollar amount of spend to the capabilities of the AIs. The chart that I published in the past was, you know, OpenAI has like five levels of AGI-ness and level, level one is sort of like a chat boss. Level two is reasoning. Level three is agents. Four is organizations. Five is something super, superhuman. I don't remember what the exact levels are, but each, each, you can sort of each match each of them with like tiers. Like, uh, 20 dollars is like the ChatGPT tier. 200 dollars is where you're at. 2000 dollars is higher to 20,200 thousand, right? Like, you can see levels where it makes sense. I think Brightwave is also there, by the way. Like, I don't know what Brightwave charges, but it's higher, right, than a ChatGPT. And like, you have to deliver more value for that. But you, you can do it now.

AI assessment note: “matching the dollar amount of spend to the capabilities of the AIs”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah, yeah, exactly. I saw there's no test writer tool. Is it because it generates the code and then you're running it against Sweebench anyway, so it doesn't really need to write the test or?

A Yeah, so this is, this is one of the interesting things about Sweebench is that The tests that the model's output is graded on are hidden from it. That's basically so that the model can't cheat by looking at the tests and writing the exact solution. But I'd say typically the model, the first thing it does is it usually writes a little script to reproduce the error. Uh, and again, most sweet bench tasks are like, Hey, here's a bug that I found. I run this and I get this error. So the first thing the model does is try to reproduce that. And so it's kind of in rerunning that script as a mini test. But yeah, sometimes the model will like accidentally introduce a bug that breaks some other tests and it doesn't know about that.

AI assessment note: “the first thing it does is it usually writes a little script to reproduce”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. And what about the hardware? A lot of my friends that work in robotics, one of their Big issues. Like sometimes you just have a servo that fails and then you gotta fail and it takes a bunch of time to like fix that. Is that holding back things or is the software still anyway?

A I think both. I think there's, there's been a lot more progress in the software in the last few years. And I think a lot of the humanoid robot companies now are really trying to build amazing hardware. Hardware is just so hard. Um, it's something where classic, you know, you build your first robot and it works, you know, great. Then you build 10 of them. Five of them work, three of them work half the time, two of them don't work, and you built them all the same, and you don't know why. And it's just like, the real world has, like, this level of detail and differences that software doesn't have. Like, imagine if every for loop you wrote, some of them just didn't work. Some of them were slower than others. Like, how do you deal with that? Like, imagine if every binary that you shipped to a customer, each of those for loops was a little bit differently, was a little different. It becomes just so hard to scale and sort of maintain quality. Uh, of these things. And I think that's like, that's what makes hardware really hard. It's not building one of something, but repeatedly building something and making it work reliably. Where again, like you'll, you'll buy a batch of a hundred motors and each of those motors will behave a little bit differently to the same input command.

AI assessment note: “I think both. I think there's, there's been a lot more progress”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What was it like when you, when you tried it for the first time? Was it, was it obvious that Claude had reached that stage where you could do computer use?

A It was somewhat of a surprise to me. Like, I think I actually, I had been on vacation and I came back and everyone's like, computer use works. Um, and so it was kind of this very exciting moment. I mean, after the first, just like, you know, go to Google, I think I tried to have it play Minecraft or something and it actually like installed and like opened Minecraft. I was like, wow, this is pretty cool. So I was like, wow, yeah, this thing can actually use a computer. And certainly it is still beta. You know, there's certain things that it's, it's not very good at yet, but I'm, I'm really excited. I think most broadly, not just for like new things that weren't possible before, but as a much lower friction way to implement tool use. One anecdote from my days at Cobalt Robotics. We wanted our robots to be able to ride elevators, to go between floors and fully cover a building. The first way that we did this was doing API integrations with the elevator companies. And some of them actually had APIs. We could send a request and it would move the elevator. Each new company we did took like six months to do, because they were, they were very slow. They didn't really care.

AI assessment note: “It was somewhat of a surprise to me. Like, I think I actually, I had been on vacation”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q model, or whether it's like a chain of models, and Gnome and basically everyone on the Strawberry team was very insistent that what they did for reinforcement learning, train of thought, cannot be replicated by a whole bunch of open source model calls. Do you think that that is, they're wrong? Have you done the same amount of work on RL as they have, or was it a different direction?

A I think they take a very specific approach where I do, the caliber of team is very high, right? So I do think they are the domain expert in doing the things they are doing, but I, I don't think there's only one way to achieve the same goal. We are on the same direction in the sense that the quality scaling law is shifting from training to inference. We are definitely honest. For that, I fully agree with them, but we are taking a completely different approach to the problem. All of that is because, of course, we didn't train the model from scratch. All of that is because we built on the show of giants, right? So the current model available we have access to is getting better and better. The future trend is the gap between the open source model, the open source model, it's just going to shrink to the point there's not much difference. And then we are on the same level field. That's why I kind of, I think, our early investment in inference and all the work we do around balancing across quality, latency, and cost pay off because we have accumulated a lot of experience there, and that empowers us to, to build and release this new model that is approaching OpenAI's quality.

AI assessment note: “I fully agree with them, but we are taking a completely different approach to the”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q think about, you know, people are like, hey, Lama, 2.2 is X on MMLU, but maybe, you know, using speculative decoding, you go down a different path. Maybe some providers run a quantized model. How should people think about How much they should care about how you're actually running the model? You know, like what's the delta between all the magic that you do and like what a raw model.

A Okay. So there are two big development cycle. One is experimentation where they need fast situation. They don't want to think about quality and they just kind of want to experiment with product experience and so on. Right. So that's one. And then it looks good and they want to kind of postpartum market by scaling. And the quality is really important and latency and all the other things are becoming important. During the experimentation phase is just pick a good model. Don't worry about anything else. Make sure even like JNI is the right solution to your product. And that's the focus. And then postmodern market fit, then that's kind of the three dimensional optimization curve start to kick in. Across quality, latency, cost, where you should land. And to me, it's a purely a product decision. To many products, if you choose a lower quality, but better speed and lower cost, but it doesn't make a difference to the product experience, then you should do it. So that's why I think inference is part of the validation. The validation doesn't stop at offline well. The validation is kind of, will go through A-B testing through inference. And that's why we kind of offer various different configurations for you to test which is the best setting. So this is the, like, traditional product evaluation. So product evaluation should also include your new model versions, um, and different model set…

AI assessment note: “to me, it's a purely a product decision. To many products, if you choose”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q That kind of personal automation, would you say it's kind of like, um, LLM Zapier type of thing? Like if this, then that, and then, you know, do this, then this, this is so very, you're programming with English.

A So you're programming with English. So you're just saying, oh, you do this. And then that you can even create some, some, some form of API. As you say, when I give you the command X do this, when I give you the command Y do this and you describe the workflow, but you don't have to create boxes and create the workflow explicitly. It's just need to describe what's Are the tasks supposed to be and make the tool available to the agent? Tool can be a semantic search. The tool can be querying into a structured database. The tool can be searching on the web. Um, and obviously the interesting tools that we only starting to scratch are actually creating external actions like reimbursing something on Stripe, sending an email, clicking on a button in the admin or something like that.

AI assessment note: “you describe the workflow, but you don't have to create boxes”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And I noticed that in your language, you're very much focused on non-technical users. You don't really mention API here. You mention instruction instead of system prompt, right? That's very conscious.

A Yeah, it's very conscious. It's a mark of our designer, Ed, who kind of pushed us to create a friendly product. I was knee deep into AI when I started, obviously, and my co-founder Gabriel was, uh, was at Stripe as well. Uh, we started a company Glazer that got acquired by Stripe 15 years ago, was at Alain, a healthcare company in, in Paris after that. It was a little bit, uh, less so knee deep in AI, uh, but, uh, really focused on product. And I didn't realize how important it is to make that technology not scary to end users. It didn't feel scary to me. But it was really seen by head, our designer, that it was feeling scary to the users, and so we were very proactive and very deliberate about creating a brand that feels not too scary, and creating a wording and a language, as you say, that, that really tried to communicate the fact that it's gonna be fine, it's gonna be easy, you're gonna make it.

AI assessment note: “Yeah, it's very conscious. It's a mark of our designer, Ed”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. Do you see in the future products offering kind of like a simulation environment the same way all SaaS now kind of offers APIs to build programmatically? Like in cybersecurity, there are A lot of companies working on building simulative environments so that then you can use agents like Red Team, but I haven't really seen that.

A Yeah, no, me neither. Uh, that's a super interesting question. I think it really going to depend on how much, uh, because you need to simulate to generate data, you need to train data to train models. And the questions at the end is, are we going to be training models or are we just going to be using frontier models as they are? On that question, I don't have a strong opinion. It might be the case that we'll be training models because in all of those AI-first products, the model is so close to the product surface that as you get big and you want to really own your product, you're going to have to own the model as well. Owning the model doesn't mean doing the pre-training. That would be crazy. But at least having an internal post-training realignment loop makes a lot of sense. And so if we see many companies going towards that over time, Then there might be incentives for the SASSs of the world to provide assistance in getting there. But at the same time, there's a tension because those SASSs, they don't want to be interacted by assistance. They want it there by agents. They want, they want the human to click on the button. So that's an interesting thing.

AI assessment note: “Yeah, no, me neither. Uh, that's a super interesting question.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q benchmarks are really not what we need to evaluate whether or not a model is good. Why did you not make a benchmark? Maybe at the time, you know, it was just like, Hey, let's just put together a whole bunch of data again, run a, Make a score that seems much easier than coming out with a whole website where, like, users need to vote. Any thoughts behind that?

A I think it's more like fundamentally we don't know how to automate this kind of benchmarks when it's more like, you know, conversational, multi-turn, and more open-ended tasks that may not come with a ground truth in. So let's say if you ask a model to help you write an email for you for whatever purpose, there's no ground truth in how you score them. Or write a story, or a creative story, or many other things. Like, how we use strategy these days, it's oftentimes more like open-ended. You know, we need human in a loop to give us feedback. Which one is better? And I think nuance here is like, sometimes it's also hard for human to give the absolute rating. So that's why we have this kind of pairwise comparison, easier for people to choose which one is better. So, from that we, you know, use these pairwise comparisons and those to, to calculate the leadable.

AI assessment note: “fundamentally we don't know how to automate this kind of benchmarks”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And then what was the discussion from there to where we are today? Is there any maybe intermediate step of the product that people missed, uh, between this was launch or?

A It was interesting because every step of the way, I think we had, we hit like some pretty critical milestones. So I think from the initial demo, I think there was so much excitement of like, wow, what is this thing that Google is launching? And so we capitalized on that. We built the wait list. That's actually when we also launched the Discord server, which has been huge for us because for us in particular, one of the things that I really wanted to do was to be able to launch features and get feedback ASAP. Like the moment somebody tries it, like I want to hear what they think right now. And I want to ask follow-up questions and the discord has just been so great for that. But then we basically took the feedback from IO. We continued to refine the product. So we added more features. We added sort of like the ability to save notes, write notes. We generate follow-up questions. So there's a bunch of stuff in the product that shows like a lot of that research, but it was really the rolling out of things. Like we removed the waitlist. So rolled out to all of the United States. We rolled out to, uh, over 200 countries and territories. We started supporting more languages, both in the UI and, like, the actual source stuff. We experienced, like, in terms of milestones, there was, like, an explosion of, like, users in Japan. This was super interesting in terms of just, like, unexpected…

AI assessment note: “every step of the way, I think we had, we hit like some pretty critical milestones.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So did he just commit to using Notebook LM for everything? Or did you just model his existing workflow?

A Both, right? Like in the beginning, there was no product for him to use. And so he just kept describing the thing that he wanted. And then eventually, like we started building the thing and then I would start watching him use it. One of the things that I love about Steven, Is he uses the product in ways where it kind of does it, but doesn't quite like he's always using it at like the absolute max limit of this thing. But the way that he describes it is so full of promise where he's like, I can see it going here. And all I have to do is sort of like meet him there and sort of pressure test whether or not, you know, everyday people want it and we just have to build it.

AI assessment note: “Both, right? Like in the beginning, there was no product for him to use.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q there. How should people think about that? Like, is this, and also like the future of the product as far as monetization too, you know, like, is it going to be The voice thing going to be a core to it? Is it just going to be one output modality and like you're still looking to build like a broader kind of like a interface with data and documents platform?

A I mean, that's such a, that's such a good question that I think the answer it's, I'm waiting to get more data. I think because we are still in the period where everyone's really excited about it. Everyone's trying it. I think I'm getting a lot of sort of like positive feedback on the audio. We have some early signal that says it's a really good hook. But people stay for the other features. So that's really good too. I was making a joke yesterday. I was like, it'd be really nice, you know, if it was just the audio, because then I could just like simplify the train, right? I don't have to think about all this other functionality. But I think the reality is that the framework kind of like what we were talking about earlier that we had laid out, which is like, you bring your own sources, there's something you do in the middle, and then there's an output is a really extensible one. And it's a really interesting one. And I think like, Particularly when we think about what a big business looks like, especially when we think about commercialization, audio is just one such modality, but the editor itself, like the space in which you're able to do these things is like, that's the business, right? Like maybe the audio by itself, not so much, but like in this big package, like, oh, I can see that. I can see that being like a, a really big business.

AI assessment note: “audio is just one such modality, but the editor itself... that's the business”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q of your previous interviews, and even the choice of the name development in the ministry name was very specific. You mentioned naming is your ethos. Can you explain maybe a bit about what the ministry does, which is not simply funding R&D, but it's also thinking about how to apply these technologies and industry, and just maybe give people an overview, since there's not really an equivalent in the US?

A Yeah, so when people talk about our smart nation efforts, It was helpful in articulating a few key pillars. We talked about one pillar being a vibrant digital economy. We also talk about a stable digital society because digital technologies, the way in which they are used, can sometimes cause divisions in society or entrench polarization. They can also have the potential of causing social upheaval. So when we talked about Stable digital society. That was what we had in mind. How do you preserve cohesion? Then we said that in this domain, government has to be progressive too. You can't expect, you know, the rest of Singapore to digitalize, and yet the government is falling behind. So a progressive digital government is another very important pillar. And underpinning all of this has to be comprehensive digital security. There is, of course, cybersecurity, but there is also how individuals feel safe in the digital domain. Whether as users on social media, or if they're using devices and they're using services that are delivered digitally. So when we talk about these four pillars of a smart nation, people get it. And when we then asked ourselves, what is the appropriate way to think of the ministry? We used to be known as the Ministry of Communications and Information, and we had been Doing all this digital stuff without actually putting it into our name. So when we eventually deci…

AI assessment note: “ultimately we decided on digital development because it wasn't the technologies... impact to society”

← previous page 8 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.