Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What's the high level take for people kind of out of the loop?
A Yes. The high level take. So yes, basically sure. A lot of your listeners already know this, but these inference providers are companies that are, you know, helping developers like run and train and, Customize open source AI models more easily than they would be able to otherwise. And so, yeah, there's just so many companies in the space. You mentioned some of them, Fireworks, Modal, there's Together, there's Base 10. There's just a ton of these companies in the space. And so, yeah, FAL is another one. Yes. Um, and so, yeah, it's, it's very interesting because a lot of these companies have gotten so much funding, but there's still a lot of like skepticism around this space and like whether these companies are actually going to end up being like, you know, like Billions or trillions of dollars, like trillion dollar company. Um, so I think a lot of people like think, basically just call these companies like GPU resellers. So they're just like, oh, all you're doing is just being like a cloud basically. You know, helping developers get access to chips to run models on. So you're like a knockoff Amazon or like a knockoff like Azure or Google cloud, um, which is like a bit derogatory. But I think that's kind of like what the con argument, that's like the argument against them. Um, because also unlike software companies, they have to like spend so much money getting the, you know, get…
AI assessment note: “these inference providers are companies that are, you know, helping developers like run and train”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Okay, cool. Before we move on from meta, I just want to, you know, leave the door open. Any other themes that you're watching, uh, on, on the meta side?
A Yeah, I mean, I think we definitely covered a lot of the really interesting ones. I think, I mean, one thing that I think, like, Mark has even kind of touched on recently is whether they'll continue open sourcing models, whether they're going to start closed sourcing stuff, and, like, what does that mean for its business model? Because obviously, like, like right now, you know, Meta doesn't make money from an API the way that, like, OpenAI and Anthropet does, but if they start closed sourcing models, maybe that might change, or That could change kind of the shape of the way that they think about how to make money from this, and so, yeah, I think that's very interesting, and I don't know, I think for me as a consumer, like, I'm always really interested in, like, consumer AI, because I feel like there's not really as much coverage on it as I would like, and, like, I personally, I use a lot of AI stuff, but nothing that's, like, truly kind of transformed my, like, non-work-related life, so I'm very curious, like, what meta will come out with, with consumer AI.
AI assessment note: “whether they'll continue open sourcing models, whether they're going to start closed sourcing stuff”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q at, um, at Google, but like, you know, there was some kind of revesting thing going on. Uh, yeah, cool. Any other sort of, uh, uh, coding agents commentaries while we're still on the sort of coding topic? Anything else you're watching? How do you cover Uh, I, I guess like coding in general, if, if you're, you know, out in, in New York and, you know, we're over here.
A Yeah. If I'm not, if I'm not like a developer and I'm like trying to understand what the heck is going on with coding startups. Yeah. Yeah. I mean, I think a lot of it is just like finding sources and developers that I really trust and like being very straight up with them and being like, Hey, I, like, I'm not a developer, but, like, show me a couple examples of things that you built with this, like, tell me your opinions about what they're good for or not good for. I think, like, one thing I'm interested in is, like, the rise of some of these, like, open source coding assistants, like Klein, which I, I know was on your podcast.
AI assessment note: “finding sources and developers that I really trust and like being very straight up”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q hits that you're, um, that, you know, you, you want to talk about? Like, I don't get a sense of like what your biggest hit is or your general sort of Thematic overview, if that, if that makes sense. You know, there's like robots, there's XAI, there's perplexity, there's all these other things. Uh, what, what stands out to you as something that you really want to, uh, chat about?
A Yeah. I mean, I guess two things. I think first, like, the first thing is like, I think this is what I'm most interested in, but, um, I did kind of touch on it with like our conversation about GPT-Five, but just like the trajectory of AI progress and kind of like, you know, First, there was, like, pre-training scaling, and then now there's, like, reasoning models and reinforcement learning. Then, like, what comes after that? Like, is it RL environments for agents? Is it something that's different than the transformer? Like, I think we're always curious just, like, what is coming next? And I mean, I like, I think that's why I'd love to talk to more people at the big AI labs, because I feel like they're really on the forefront of that.
AI assessment note: “just like the trajectory of AI progress and kind of like, you know”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Uh, I mean, there's, there's a, there's a lot, like, do you, um, do you run evals, um, you know, in a, in a way that Aether does?
A Yeah, so we looked at all the different benchmarks we could find, and it's almost comical how none of them look like my day-to-day work. Most of these evals are like, You know, solve this maze. And I'm just like, I'm never solving a maze. Like it's never anything that I've ever asked an LM to do. I'm usually asking you to do something like, Hey, I'm implementing this feature. The backend's pretty easy. I just need you to go like implement a new data model in the database, create a migration, expose an API. And I'll work on the front end while you do that. Almost no benchmarks have anything that resemble anything close to that. So one of the people on our team, Frank, he's working right now on developing our own set of benchmarks. We don't really, I don't really care too much about being like, Hey, open code's like the best on his benchmark. I more care about, I made a change to how our editor tool works. Did I make things worse? Did I make things better? Um, this is another funny thing you see right now because nobody really has a great way to do this. So every new feature is positioned as we made things better. But I've seen myself make changes where I make it better in one dimension. And it's just like horrible. I realized it's actually horrible. And like the dimensions that matter two weeks down the road. So I just feel lost trying to improve this stuff. So at first we're tr…
AI assessment note: “Frank, he's working right now on developing our own set of benchmarks.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Not, not like, you know, is there a why now? Obviously you have some secret sauce, but like, I'm just trying to figure out like, is there anything fundamental?
A I mean, it depends how you define fundamental. I think there was some math that developed in terms of like thinking about, I like to think of diffusion models in terms of like learning score functions and, and kind of like gradients basically of probability density functions, which are not really well defined in the discrete world, but there are kind of like mathematical analog goes objects that you can kind of train using Denoising, like objective, like score matching, like objective. So there was some pretty fundamental math that had to be developed. Then there's also a lot of engineering and, and, and tricks that do to build on top. And then I would also add the fact that, you know, once somebody shows that it's possible, you know, suddenly other people are much more incentivized to, to get their own thing to work. And, and I think everybody progresses more rapidly once there is kind of like a proof of concept that, that something is possible. And so, yeah, we've seen, you know, that the, the Google guys also have a, have a model. So it's exciting to see that the, the field is really starting to, to pick up speed.
AI assessment note: “So there was some pretty fundamental math that had to be developed.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q something that people should not be using the fusion models for, or do you, this is just kind of In your mind, it's like a superior, just swap everything over to these. And then as we scaled them, it's funny because you're seeing image generation go towards autoregressive models, and then you're seeing the text models go to diffusion. How do you kind of see the model use case match?
A So the API is the same. It's like text in, text out. So in principle, any use cases that is supported by an autoregressive model is also supported by a diffusion language model. In terms of capabilities, yeah, as I mentioned, like we're not Frontier level. And so there's certainly use cases where you probably wouldn't want to use a diffusion language model, or at least not the current, uh, existing diffusion language models. And so, yeah, I think at the moment it's more a function of kind of like what's the intelligence level, how much do you care about latency or cost? So there's definitely a bunch of use cases where people, where diffusion language models are already competitive. They're already the best solution for those use cases. Uh, there are others that are still yet to come, and I think that will require, uh, scaling up and making further improvements to, to the, to the training and the data and various kinds of things. I'm pretty optimistic about a future where, um, diffusion models Can become the dominant solution. I've seen it happen before with GANs a few years ago, so I wouldn't be surprised if that's the case also here. Um, I really like this idea of using context to the left and to the right. The fact that you have error correction that is built in. So you think about an autoregressive model. Once you output something, you can never take it back. And so if you w…
AI assessment note: “there's certainly use cases where you probably wouldn't want to use a diffusion language model”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q engineer, I think that's a win for us doing this, right? Okay. So I want to move on a little bit in terms of like we, we, let's say we have production and high quality LLMs. You made an interesting controversial statement earlier that you said most frontier models or most language models will use diffusion in the future. I don't think that's consensus. Uh, could you just elaborate why?
A I think that there is a chance. I mean, I, I don't know if that's gonna happen, but I feel that is, there is a world where that could happen. Um, and, and I think it could be driven. If it happens, it's gonna be driven by efficiency. Like we're all constrained by essentially power. Uh, and, uh, and, uh, if you have, I mean, at the end of the day, it's all an inference game, right? Okay. Training is expensive, but then the thing that matters is being able to serve these models cheaply and without consuming more energy that we have access to. And you only have so many gigawatts of Power and so many data centers, and there is a growing demand for tokens that you need to be able to serve to users and all kinds of apps that are being built. And so there is a real need to find a solution that can give us maybe a 10 X improvement on the, on the efficiency side. And there's a really good chance that diffusion models are the thing. So if indeed it's possible to match the quality of frontier models in terms of quality, if you have a win on the, on the inference side, that is going to be the solution that will be Uh, that will be deployed.
AI assessment note: “If it happens, it's gonna be driven by efficiency. Like we're all constrained”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And you have this technique called dynamic multiplexing, which is basically instead of having a one-to-one relationship, you have one GPU for multiple. Clients. And I saw one of your customers, they went from. 30 clients to just one single GPU and they cut costs by 97%. What were some of those learning, seeing hardware usage inefficiencies and how that then played into what you're building now?
A Yeah, I think it basically showed that there was probably a gap with even very sophisticated teams making good use of the hardware is just not an easy problem. I think that was the main, I, it's not that these teams were like not good at what they were doing. It's just that they were trying to solve a completely separate problem. They had a model that was trained in-house and their goal was to just run it. And it, that should be an easy, easy thing to do, but surprisingly still, it's not that easy. And that problem compounds in complexity with the fact that there are more accelerators now in the cloud. There's like TPUs, Inferentia, and there's a lot of decisions that users need to make, even in terms of GPU types. And I guess sort of what we had was we had internal expertise on what the right way to run the workload was. And we were basically able to build infrastructure to make it so that companies could do that without thinking.
AI assessment note: “we were basically able to build infrastructure to make it so that companies could do that”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And you said, look, what we care about is building things on top of code completion. How did you decide to, like, just not focus on, like, short-term kind of, like, growth monetization of, like, the individual developer and, like, build some of this? Because the alternative would have been, hey, all these people are using it, so we're gonna make this other, like, five bucks a month plan, monetize.
A I think, I think this might be a little bit of, like, commercial instinct, That the company has, and unclear if the commercial instinct is right. I think that right now optimizing for making money off of individual developers is probably the wrong, actually, strategy. Largely because I think individual developers can switch off of products, like, very quickly, and unless we have, like, a very large lead trying to optimize for making a lot of profit off of individual developers, it's probably something that someone else could just vaporize very quickly, and then, and then they move, they move to another product. And I'm going to say this very honestly, right? Like, when you use a product, like, Podium on the individual, on the individual side. There's not much thing, not much that prevents you to switch onto another product. I think that will change with time as the products get better and better and deeper and deeper. I constantly say this, like there's a book in business called like seven powers. And I think one of the powers that a business like ours need to have is like real switching costs. But like you first need something in the product that makes people switch on and stay on before you think about how do you make people switch off. And I think for us, We believe that there's probably much more differentiation we can derive in the enterprise by working with these large co…
AI assessment note: “optimizing for making money off of individual developers is probably the wrong, actually, strategy”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q but I wanted to sort of give the opportunity for folks to, like, get to know the new cognition and understand, like, how did this deal come together? Both of you have, like, relayed a little bit in your, in your tweets and stuff, uh, but I, I don't know, you, you want to take it from here as to, like, how the, the relationship between the two companies started?
A Yeah, I got a, actually, there was a couple of ways. Scott, Scott tried multiple paths to get to me. So there was a text message through a friend that, that put us together. And then there was also an email that had a subject matter of like, it's just a chat with a question mark. And, uh, I mean, it, it caught our eye. And I think when we were looking at like Friday was like a frantic day of telephone calls, I'll say, I think we, I probably called at least 70 different people on Friday. Um, and like when Scott, when I talked to Scott, it was just like, Hey, this kind of makes sense. Like there was other companies. That were interested. There was actually VCs that wanted to invest. Um, but Scott, as, as a person with, with Russell, and then I was with Graham, uh, just talking about like, what are the companies actually missing from each other? And then also what are the, what are the product synergies as well?
AI assessment note: “Scott tried multiple paths to get to me. So there was a text message”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And was this the first shape of the product or did you get to the plan act iteratively? And maybe was this the first idea of the company itself or were you exploring other stuff?
A It was a lot of Especially in the early days of the client, it was a lot of experimenting and talking to our users and seeing what kind of workflows came up that they found that were useful for them and translating them into the product. So plan and act was really a byproduct of just talking to people in our Discord, just asking them what would be useful to them, what kind of prompt shortcuts we could add into the UI. I mean, that's really all plan and act mode is. It's, is essentially a shortcut for the user to save them the trouble of having to Type out, you know, I want you to ask me questions and put together a plan, um, the way that you might have to and, you know, some of the other tools you'd have to, like, be explicit about. I want you to come up with a plan before, you know, acting on it or editing files. Incorporating that into the UI just saves the user the trouble of having to type that out themselves.
AI assessment note: “plan and act was really a byproduct of just talking to people in our Discord”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I'm happy to move on. I would say like the last thing that's kind of curious is like, if Anthropic hasn't, hadn't come along and made MCP, what would have happened? What's the alternative history, right? Like, would you have come with MCP?
A So we saw some of, uh, our competitors who have been kind of working on their own version of Plug and play tools into these agents. They kind of had to natively create these tools and integrations themselves directly into their product. And so I think anybody in the space would have had to just do the laborious work of having to recreate these tools and integrations for. So I think Anthropic just saved us all a lot of trouble and tapped into the power of open source and community driven development and allowed, you know, individual contributors to make an MCP for anything people could think of. I'm gonna really take advantage of people's imagination in a way that I think is, like, necessary right now for us to really tap into full potential of, of this sort of thing, so.
AI assessment note: “anybody in the space would have had to just do the laborious work”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah. Uh, double clicking on the AST mention. That's very verbose. When do you use that?
A Right now it's a tool. The way that it works is when Klein, once, when Klein is doing sort of the agentic exploration of trying to pull in relevant context, and it wants to sort of get an idea of what's going on in a certain directory, for example, there's a tool that lets it pull in all the sort of language from a directory. So it could be the names of classes, the names of functions, And that gives it some idea of, okay, here's what's going on in this, in this folder. And if it's, if it seems relevant to whatever the task is trying to accomplish is, then it sort of like zooms in and starts to actually read those entire files into context. So it's, it's essentially a way to help it kind of figure out how to navigate through large code bases.
AI assessment note: “when Klein is doing sort of the agentic exploration of trying to pull in relevant context”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Any thought on CloudMD versus AgentsMD versus AgentMD? I built an open source tool called Agents-Nine-to-seven, like the XKCD that just copy-pastes it across all the different file names, so all of them have access to it. Do you think there should be a single file? Like, there's also like the IDE rules versus the agent rules. There's kind of like a lot of issues.
A I actually think it's fine that each of these different tools have their own specific instructions, because I find myself using A cursor rules and a client rules separately. When I want client, the agent, I want him to work, you know, a certain way that's different than how I might want, you know, cursor to interact with my code base. So I think each tool is specific to the kind of work that I do and I have different instructions for how I want these things to operate. So I think I've seen like a lot of people complain about it and I get that it could make code bases look a little bit ugly, but for me, it's been like incredibly helpful for them to be separated.
AI assessment note: “I actually think it's fine that each of these different tools have their own specific instructions”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So, so Taubench, uh, I actually didn't know it was from Sierra, uh, but it's, it's actually very, very, uh, influential. Yeah, can you talk more about Taubench?
A Oh yes. So essentially tower bench when it was released, uh, I mean, yes, the top bench is for two, uh, two domains, the test agent capabilities for two domains, which is airline and retail. And the setup is that they have very long prompts in the beginning with lot of requests. And then you have these tools defined and the tools can interact with certain Databases for read and write operations, and it is predefined that once the whole agent tick task is done, then they will validate on the final, uh, edits made on, and they will check for the expected value in the data sets, uh, in the database, which are basically CSV style of files. Uh, and in that fashion, they will validate whether the task was done or not. So it's, it's, it's pretty well defined, such that You get a deterministic evaluation all the times. That's the beauty of this. And it's, it's domain specific, which I think is missing from other evaluation data sets. Uh, and that is also one of the themes for our V two, um, because I think people don't just come from a general mindset, right? People in the industry come from, let's say I'm from healthcare. I want to see if my, if this models work for healthcare or not, people are coming from investment, finance, insurance industry, and they want to know. That is, are the models ready for my domain or not? Because we all know that things work on the general domain side,…
AI assessment note: “top bench is for two, uh, two domains, the test agent capabilities”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q obviously only, only, you know, like those kinds of people, you are now CTO co-founder of Speak. Uh, I would say from a very early stage, like one of the most successful and prominent open AI partners that like anyone would know is like doing, doing well and like teaching English to Koreans is like your, your, your rough, um, remit at the time. How did that all come about?
A It's funny that you say that because despite our current sort of revenue scale and Objectively, I think how successful we are. We've always operated in a market, at least initially on the other side of the world and been much, much more popular in the sort of Eastern world and a bunch of Asian markets and relatively unknown in the West. So it hasn't really felt like we've had that sort of awareness until, you know, the past few years really. But brief story is that my co-founder and I back in Fascinated by the promise of AI, and we spent a year sabbatical basically learning everything we could. We talked to Karpathy back then, actually, when he was, like, just finishing grad school, and did a lot of sort of self-study research, and we were just so convinced, I think fundamentally, that speech models were going like this, language models were going like this, and in the five to 10 year span, they would become superhuman. And We were utterly convinced of this future, and we saw that the way people learn things, and specifically learn languages, which was a very sort of human-based thing, if you really care about fluency, that would completely change, and we'd be able to build language shooters that were pure software, pure AI. So that was kind of the genesis story of Speak. It took much, much longer Then we expected to build a great product and find good PMF. The first few years …
AI assessment note: “So that was kind of the genesis story of Speak.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q lot of Americans learn Korean because of K-pop. That's a, that's a side thing. But like, you could have done Taiwan. You could have done China. I saw, I remember starting a documentary about how China was crazy about English or mad about English. I think that was the title of the documentary. Was it obvious? Were you sure when you went into Korea or was it just a test?
A We visited a bunch of Asian countries when we were thinking about how do we relaunch things? How do we focus in? And we almost chose Taiwan actually. Um, but I think it was a little bit serendipitous. So our first employee is Korean and was my co-founder's college roommate, actually. When my co-founder visited Seoul, To check out the market. He asked SJ to come along as essentially a translator and to like, you know, facilitate. And I think that just went really well. And it was just very obvious from being on the ground in the market that Korea is pretty obsessed with learning English. And there is every human-based solution Possible, right? You know, like English academies, classes, skyscrapers full of classrooms, stuff like that. And our logic was basically, if we can really make headway and win this market that is chock full of these human competitor products and all these people that fundamentally care about fluency, then we probably have something pretty real and strong PMF that we could win other markets with. So that was the original logic. And, you know, so far it's been working.
AI assessment note: “it was just very obvious from being on the ground in the market”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I'm curious to sort of double click on to just the tech side. We talked a little bit about the content that you, that you own and develop in-house. And we talked a little bit about the onboarding memory. I assume that you have conversational memory as you, as you go, right? And any other major pieces of the puzzle that really unlocked it for you?
A So there's a few things I can talk about. I think one thing is In order to go from teaching English to teaching a bunch more languages, we needed to really figure out more direct AI content generation. That was a pretty, right, because it's hard to scale like our little studio in LA where we shoot a lot of the video lessons. Um, all of the scripts were written manually before by our content team, but we want like a hundred X more content, right? And 10 X more languages, eventually a hundred X more language pairs, which is how we think about it. It's like, what's your native language and then what language are you learning? And really the only way to do that is to make it more AI generated and, you know, very much like a AI native company. We want to be on a frontier here. We want to keep a small team and to have as much leverage as possible through these types of tools. So that's a big active area where we're building out I think using, you know, people overuse the word agent, but we have a tutor agent, we have a curriculum writing agent, we have a giant LM based pipeline that creates curriculum, scaffolds it in the right way, writes the lessons themselves. That's a big active area that will basically help us to scale to a lot more markets and a lot more languages. So that's like one big thing. Another big thing is we care a lot about Fluency, obviously. Specifically, we want t…
AI assessment note: “we have a tutor agent, we have a curriculum writing agent, we have a giant LM based pipeline”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q time looking at this like, uh, you know, like video tree where you do video plus audio at the same time on how you can tweak the audio part versus the video part? Because I can imagine you might work on a video part, And then you want to change the audio generation model. I don't actually know how the model works inside on like how much you can tweak.
A We haven't really looked at the video stuff much. We basically think that we're very bandwidth constrained, right? So we're just scaling and trying to hire as fast as possible, like everyone else is. And as a result, we're really focusing on just like the most in reach, highest impact things. I do think that the barriers are coming down very fast for all of this sort of stuff. I'm just so excited about multimodality and where things are going here, because imagine if you're learning Spanish, being able to look at an image that the model generates for you and then doing Q and A on it, right? Like a beach scene. And then the model will ask you like, oh, how many people are running on the beach? And then you have to sort of respond in the target language that you're learning. Very traditional language learning exercise, but you can imagine it being fully generative, which is really cool.
AI assessment note: “We haven't really looked at the video stuff much.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q this all ends up, that's great. But like people aren't, are not doing that. Instead we're building, You know, five hundred billion dollar data centers in the middle of Texas and like, you know, all hail the, the, the God cluster, uh, that just will, you know, eventually wrap around the sun and consume solar energy because that's, that's what we need. Do we finish out the universal geometry thing?
A Let me finish the kind of, uh, methodological description. So, so we had this goal. So, so yeah, back to the embedding universality, we started with going from embeddings to text. We know about this platonic representation hypothesis, and maybe I'll skip over the details, but basically we had total inspiration from computer vision and this model from 20 17 called CycleGAN, which is among other things, uh, it's a way to map between two different distributions without any underlying notion of like which thing should be mapped where. It's just based on some kind of Idea of closeness. So like the cool thing about this, if you look at the top left, so I guess the, the top left is Monet. So impressionist paintings and this picture on the right is a photograph. So like it's learning this kind of like semantic notion of what content goes, where just by mapping a distribution of Monet pictures to a distribution of photographs without actually telling it which Monet picture should map to which photograph. It's kind of a subtle point I'm making. It takes a little bit of time to wrap your head around, or maybe like go to the middle one, if you don't mind the zebras and the horses. So like, it's clearly learning like what an animal is and what legs are and sort of like more abstract stuff, like what, uh, the camera position should be and, and what grass is and stuff like that. And it's lear…
AI assessment note: “Let me finish the kind of, uh, methodological description. So, so we had this goal.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Okay. Do you think this is a hard limit? Do you think someone can come up with a better algorithm, but better architecture, and then sort of just change the slope?
A There are two axes here. One is the ability of the model to store data, and I think we can definitely improve that. I think, like, maybe even if we tested this with LALAMA architecture, like, there's sort of like a GPT++ architecture, like, I would guess that can store better data just because the kind of numerical flow is a little bit better, the nonlinearities are maybe, like, A little bit more suitable to training. Like that will probably raise the bound a little bit. And then the second axis is that our measurement tools are just not that good. Like this is, you know, me, I'm a grad student. I'm running all these hyperparameter sweeps and sort of like where we draw conclusions from them. But even that being said, like there are probably ways to measure this better. And, but all that would do is push the number up. So it's possible. Like there is a way to store five bits per parameter. If you have like a better optimization technique or If you were a super genius and you could just perfectly set the weights to store the data, then maybe you can do better. And this is just sort of like what we can reach through optimization is this 3.6 bits per parameter. But I would be happy if someone came along with a much better measurement tool. Like, uh, this is just sort of like the first measurement. I mean, I would, I would guess in the future, like, you know, people will look back a…
AI assessment note: “There are two axes here. One is the ability of the model to store data”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q well as you could. I think it, It's, it could be cleaner, but you know, sometimes you just have to compete where you compete. I think that the challenge for you is just the sheer amount of firepower the big labs are putting behind their agents, right? And you're coming out saying you got a better agent than them. Good for you, but also like, you know, you're fighting them.
A So it's an interesting position because we are competing with them. And we're also their customer. And it's, it's interesting question to me from their perspective. How much do they want to own the end user relationship? How much do they want people building on their APIs? There's this interesting dynamic where I think Anthropic is like pretty clearly the leader in the, in the coding models right now, but I don't know if that's going to be perpetuity. Like they have a competitive dynamic with each other also. Right. Whereas we have this interesting advantage of like, If Gemini became the best coding model in two weeks, everyone who had, like, gone all in on cloud code would be, like, in a little bit of a weird position, whereas we're just like, okay, we'll make Gemini our default model. So there is, like, some advantage to being where we are at a layer above. You know, I'm sure that the model companies are thinking, like, how do we lock people in? Because they don't, like, that's a very big risk for, for Anthropic, in my opinion. Uh, to just be like, okay, like, Gemini's best next week. Cursor moves. Like, cursor I think is some very significant portion of their revenue. I don't totally know, like, is the short answer, but the certain north star for us is like, can we build the best user experience around these, like, around the models? I do think there's a bunch of alpha in te…
AI assessment note: “So it's an interesting position because we are competing with them.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q you know, exists at a different part of the stack than, than Cloud Code specifically. Cloud Code as a CLI, like, you could use it in any environment, so it's up to you to compose it together. Should we talk about how, how and when models fail? Because I think that was another hot topic for you. I'll just leave it open. Like, how do you observe Cloud Code failing?
A There's definitely a lot of room for improvement in the models, which I think is very exciting. Most of our research team actually uses quad code day to day, and so it's been a great way for them to be very hands-on and, like, experience the model failures, which makes it a lot easier for us to target these in model training and to actually provide better models, not just for quad code, but for, like, all of our coding customers. I think one of the things about The latest Sonnet three seven is, it's a very persistent model. It's like very, very motivated to accomplish the user's goal, but it sometimes takes the user's goal very literally, and so it doesn't always fulfill what, like, the implied parts of the request are, because it's just so narrowed in on, like, I must, like, get x done. And so we're trying to figure out, okay, how do we give it a bit more common sense? So that it, it knows the line between trying very hard and like, no, the user definitely doesn't want that.
AI assessment note: “it sometimes takes the user's goal very literally, and so it doesn't always fulfill”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Just to zoom out, you obviously do not have a separate Cloud Code subscription. I'm curious what the roadmap is. Like, is this just going to be a research review for much longer? Are you going to turn it into an actual product? I know you were talking to a lot of CTOs, MVPs, or is there going to be Cloud Code Enterprise? What's the, what's the vision?
A Yeah, so, um, we have a permanent team on Cloud Code. Uh, we're growing the team. We're really excited to support Cloud Code in the long run. And so, yeah, uh, well, we plan to be around for a while. In terms of subscription itself, it's something that we've talked about. It depends a lot on whether or not most users would prefer that over pay as you go. Um, so far pay as you go has made it really easy for people to Start experiencing the product because there's no upfront commitment. And it also makes a lot more sense with a more autonomous world in which people are scripting cloud code a lot more. But we also hear the concern around, hey, I want more price predictability if this is going to be my go-to tool. So we're very much still in the stages of figuring that out. I think for enterprises, given that cloud code is very much like a productivity multiplier for ICs and most ICs can adopt it directly. We've been just, like, supporting enterprises as they have questions around security and productivity monitoring, and so, yeah, we've found that a lot of folks see the announcement and they want to learn more, and so we've been just engaging in those.
AI assessment note: “we have a permanent team on Cloud Code... In terms of subscription itself”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q So why move to SF? Do you think that everybody in your similar, kind of similar situation should? Any pros and cons that you're experiencing?
A I think it's, ah, you can definitely build a DevTool company from Europe. I think it's a lot about question of how easy you want it to be in the earlier days, especially if you're building like a, you know, kind of like red ocean versus blue ocean waters. If you are building in a field that already exists and you have like large competitors, you're building something that's 10 times, hundred times better. It's probably, you probably don't need to be NSF. Uh, you eventually probably will need some kind of US base because for sales and customers, but You can very well build this, uh, from Europe. I know great companies doing that, uh, because all it's, the knowledge is already among all the developers. But I would maybe argue that it might be harder to find people that are comfortable with a fast iteration loop and changing things early on, you know, almost we are pivoting every week, every month. But the main motivation for us was we just wanted to be very close to our users. And it was clear after a few weeks that SF is becoming this AI hub. And what we, uh, used to do, and we, we still, we still do it sometimes, but slightly less because of, like, not, uh, having that much time and resources, but we just met with the customers that had problems, aren't users, and we just, like, implemented it to be for them, like, next to them, like, we made a PR.
AI assessment note: “the main motivation for us was we just wanted to be very close to our users.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q the, I posted a tweet because I found it in the dashboard that you can just opt in and, and like, there's, there's basically 16 days left for this program where you can just get free inference. Um, and like, so I'm just kind of curious, like, um, What you found from that kind of IF eval that, that might be different from the normal IF eval that people have?
A Yeah, totally. A lot of the instruction following evals that are open sourced are open sourced, you know, or, or crafted in a way that are easy to craft. So for example, like, Graph Walks is, is somewhat easy to craft. Like, you can create this graph and verify it easily, but it is not exactly aligned with what the users are doing. And this is true for some of the instruction following evals where you ask the model to output exactly four words or, you know, three paragraphs. Or stuff like that. Things that you can verify easily in code. Um, and these are useful instructions, but we find that many of the really interesting instructions are actually challenging to grade. And so the open source evals often don't have them. And so getting this, like, real world, uh, diverse set of data actually helps us find, like, what are the commonalities in what developers are doing? What is a really good example of, like, a negative instruction? And then we can go from there and figure out how to, how to evaluate it.
AI assessment note: “many of the really interesting instructions are actually challenging to grade”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q We were kind of talking about it before, and there's this weird thing where One week is more expensive of both one day and one month. What are like some of the market pricing dynamics? What are things that like this to somebody that is not in the business? This looks really weird, but I'm curious, like if you have an explanation for it, if that looks normal to you.
A Yeah. So the, the simple answer is preemptible pricing is cheaper than non-preemptible pricing. And the same economic principle is the reason why that's the case right now. That's not entirely true on SF Compute. SF Compute doesn't really have the concept of preemptible. Instead, what it has is very short reservations. So, you know, you go to a traditional cloud provider and you can say, hey, I want to reserve contract for a year. We will let you do a reserve contract for one hour, which is the part of SFC. Um, but what you can do is you can just buy every single hour continuously, um, and you're reserving just for that hour. And then the next hour you reserve just for that next hour. And this is obviously like a built-in, this is like an automation that you can use. But what you're seeing when you see the cheap price is you're seeing somebody who's buying the next hour, but maybe not necessarily buying an hour after that. So if the price goes Up too much. They might not get that next hour. And the underlying part of this, of where that's coming from in the market, is you can imagine like day-old milk, or like milk that's about to be old, it might drop its price until it's expired, um, because nobody wants to buy the milk that's in the past, or maybe you can't legally sell it. Compute is the same way. No, you can't sell a block of compute that is not, that is in the past. And s…
AI assessment note: “preemptible pricing is cheaper than non-preemptible pricing”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Like even the, you know, when you mentioned in the, I'm reading from the blog post, you decided not to investigate this direction because of significant extra costs. Like if it costs 10 times more, how much more of a result would we get? You know, like if the, if the models start going down in costs, are we just going to do more of that?
A Yes, I expect we will. I think it will become less of over time as cost will come down. I expect it will become less of a cost question and more of a UX question because with ensembling, well, what we found is that users really want to see what the agent is doing and follow along. And if you're doing ensembling, you don't have one trajectory anymore. You have several. And then the question of how do users supervise that? As it's happening becomes pretty tricky. Like you can present the results after it's done, but typically users don't, don't like waiting that long to see what's happening. They want to make sure the agent is on the right track. So I think even if the cost ends up making sense and it might, the UX will become a problem in terms of how much more you can squeeze out of this. Yeah, I would guess you can get a few more percentage points, at least with better ensembling, because what we did is very simple. And I know others have been experimenting with this as well. I don't think it's It's, ensembling is usually not the kind of thing that just solves a benchmark, just also in other domains, you typically can get like, yeah, and this, and first we bench probably a few percentage points is my guess.
AI assessment note: “Yes, I expect we will. I think it will become less of over time”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What's the plan for moving it out of the ID? Do you also see a world in which just in linear, you can say, please implement, and then you don't have to actually open VS code and be on top of it?
A Yeah. So that's, that's a great question that we, we ask ourselves almost on a weekly basis. These things are getting powerful quickly. Um, and there are like, if you're doing zero to one development, It seems it's already at a point where you don't need a lot of the ID or maybe not even an ID at all to get to something that works and iterating on it. What we found so far is that when you're working in complex code bases, it's still quite often that you need to use the ID to make changes. And so for now, our product is going to be in the ID because of this reason. And also when we add support for multiple agents, we expect that it's still going to be in there. And this is also based on, on feedback from early users. I think in the future, at some point, my guess is that the IDE is going to become less of the focal point and more like an app that you can launch when you need to dig in deeper, but you spend most of your time away from it. So maybe 80% of your time is in a website or an app where you control the agents and then 20% of your time, maybe you go into the IDE. But I feel like for, for the developers that we're targeting, we're not Quite there yet.
AI assessment note: “for now, our product is going to be in the ID because of this reason.”