Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q we had a nice launch of the book. Very successful. But before we get into all that, I want to start off with a fun question for you. Ok, you're expert inference engineer. What happens when I send a long query, say, 200,000 tokens into base tens inference? What's the process of query through GPU model routing, balancing, all that? What is all the stuff that we don't think about?
A With a long query specifically, the first thing that I'm going to ask is, have you sent me this query before, or at least part of it? Um, and I really hope you have, because it's going to be a lot easier for me and a lot cheaper for you. So the first thing that we're going to look at is some kind of cache away routing where we're going to see, we probably have a number of instances, a number of replicas up serving whatever model you're hitting. Uh, we want to send this one to Something with number one, available pre-fill workers, and number two, ideally, some cached input already there so that we can skip pre-fill on at least part of these 200,000 tokens. Um, if you're doing 200,000 tokens, it's probably coding or a multi-tone agent or something where you would expect to have that cached. Um, if you don't, we're gonna have to send it to a pre-fill worker. Um, we've, at least on certain models, disaggregated pre-fill and decode. Um, so You're going to have one set of GPUs that's solely going to process the input, create that KV cache, and get you your first token, and then that's going to be passed over to a separate set of GPUs, which is going to run decode. We're going to iteratively make those tokens. Um, we're probably going to have some kind of speculator model in front of that. I'm going to assume that you're doing coding, and because of that, our speculator model, which a…
AI assessment note: “the first thing that we're going to look at is some kind of cache away routing”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q know, it's the same thing, it's just a model. Can we not just distill a lot of the weights, quantize the speculator, but out of my domain. Um, I guess the, the question that also comes up is, this is all for big server workloads, right? How much of this applies to, say I have this MacBook, I want to run Gemma really efficiently, um, Similar problems? Not the same?
A Pretty different. I talked to Cero, um, about this on, on his podcast a couple weeks ago. The difference between inference engineering for the data center and for production workloads versus inference engineering for local AI is that we start with fundamentally like different constraints and different goals. With local AI, it's how do I fit this model onto my hardware? And then make it less dumb. And with data center inference, it's how do I load this model and then make it less slow? And obviously, you know, we care about less dumb and they care about less slow, but the local AI inference engineering ecosystem, I think actually has a lot for us to learn from in the data center space. They are experts in various forms of quantization, including dynamic quantization that we just kind of don't touch. In the pruning, in the distillation, in the, uh, you know, layer removal. There's removal letters less. Yeah. Yeah, but, but,
AI assessment note: “Pretty different. The difference between inference engineering for the data center and local AI”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q What goes into shipping a mainline model like this? Like, what's the behind the scenes?
A It's complex. The Gemma team is actually relatively small. We have, like, two or three PMs. We have one marketing person, and then there are, like, engineers and researchers working on shipping this. Of course, there's, like, default training part. How do we do the post-training, distillation, post-training techniques, and so on. What is quite exciting is that once we have the model, then we collaborate with a bunch of open source partners, right? So for example, we work with Lama CPP, Olama, MLX, Hogan Faces, BLM, NVIDIA, AMD. So we have almost 50 external partners for every, well, for the Gemma for launch, which has been the most complex launch. And also internally, we collaborate with a bunch of different teams. So think of Google Cloud, Vertex, Vertex Models as a Service, ADK, uh, and then Android as well, right? So we work, for example, with the Android team, and with the launch of Gemma IV, we released an integration with Android Studio. So in Android Studio, there is this agent mode where you can have a model helping you buy code and do things within Android Studio. And they should say integration with offline models using Lama CPP or BLM or any OpenAI compatible endpoint. So now you can use Gemma IV to also buy code Android applications in Android Studio.
AI assessment note: “The Gemma team is actually relatively small. We have, like, two or three PMs.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q You use Windsor for an example, but use that to tell the story of- Of KP? Of, yeah.
A When I joined six years ago, my charter was kind of twofold at KP. Number one was, hey, we have this group of CIOs and customer networks. Can you help us manage it? Right. Number two was, Hey, our founders need a lot of help on sales and distribution where you can, can you help them there? That was like the core charter, right? Then I realized that in order to help founders with go to market, like we needed to help them hire. So that was the excuse for the podcast, right? It was like, all right, I need an excuse to get to know these people so I can help these founders hire great CROs. Uh, then That all started to work, and we were like, great, let's double down on helping founders with sales. So we hired somebody on my team, Liam. Then we were like, great. Let's double down on helping folks like Varun get access to world-class customers. So we doubled down on that and hired, uh, hired somebody else. And we're like, great. Let's help founders with building their demand gen funnels and a bunch of stuff on the marketing side.
AI assessment note: “When I joined six years ago, my charter was kind of twofold at KP.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Let's run through maybe outside of, I think the, this is the prompt to code kind of thing. And then I think what people like about clock code is, uh, the plan mode, kind of like the thinking mode, there's kind of like all these different ways to use it. What does this look like here?
A Yeah. So we're still exploring exactly how that's gonna look. Uh, so plan mode is one of the most requested features. We're probably gonna ship it this week, but we're not implementing it as An explicit plan mode. We're implementing the concept of modes. So you'll be able to hit tab and cycle through different modes. We'll have a built in like plan mode and implement mode, but you'll be able to extend this with your own types of modes. And all a mode is a combination of a custom system prompt, optionally a different model and optionally a different set of tools. So you can make a mode where you're like, I want, I like talking to Gemini when I'm doing like larger code based things. So you can have a good Gemini mode. That you can tab into. It'll switch to using Gemini. It'll take advantage of the larger context. It'll maybe only pass it read only tools. Um, and then once you have a plan there, you can, you can tab back and get to implement mode, which might be done what's on it and actually do the implementation.
AI assessment note: “We're implementing the concept of modes. So you'll be able to hit tab”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Definitely, they do. Um, yeah, what, what else is, um, interesting among all this? Like, is this, is this sort of a PipeCat centric. Are we going to spend a lot of time on, I guess this like guardrails scripting? I think, I think really like building smart voice is a really quick demo, super easy, but then adding capabilities to it, making it reliable is really hard.
A That's a lot of the genesis of the course is like, how do you get from a demo to something that you could actually use in production? And there'll be a lot of PipeCat just cause I'm doing a lot of the work to kind of glue the course together, but it's not PipeCat. Centric or pipe cat only. There are a bunch of folks who are participating in the course who have their own stuff going on. Like Freddie from hugging face is going to talk about fast RTC and hang out in the discord and kind of be around. I really love that work. Uh, Sean from open AI is going to talk about their web RTC and one, 800 chat GPT layer. And so as Dominic can talk about the open AI agents framework, which they did a really elegant, like voice layer on top of that original agents framework in the most recent release. We have like Vappy, which is a kind of higher up the stack, like everything together, a platform that I recommend often to a lot of people who start with pipe cat, but kind of want more batteries included. There's a brand new competitor Vappy. I think it's fair to say called layer code, uh, that has a thousand dollars of credits for students in the course that I got an early demo of. I think, you know, those guys and we were like, oh, well you should, you should hang out in the course too. You've got good ideas.
AI assessment note: “there'll be a lot of PipeCat... but it's not PipeCat. Centric or pipe cat only.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q it's not just evaluations. I mean, evaluations, you know, you mentioned your own internal evals, but also, um, do you, do you just run a battery of tests every single time? Do you have a well-defined process? Do you have tooling that you, that you built that might be interesting to dig into, right? Like I think as a founder of any, any AI company, basically every founder needs this.
A Yeah, absolutely. So I would say the process that I like to follow is in the beginning when developing a feature, come up with a curated set of samples. Could be as small as 10, 10 samples that you run through, and then yeah, it's all notebooks basically. You run through the samples. The data set is small enough that you can keep it in your head. You become kind of very familiar with those samples, and you see how well iterations improve or regress on them beyond a certain point that's not good enough anymore, and then you go to Hundreds or sometimes maybe thousands or more of samples. So we have infrastructure for doing these kinds of evaluations, both with and without code execution. So code execution, like checking the correctness of solutions and running tests automatically can help, although not, although you can get a lot of mileage out of evals that don't have code execution in them that just compare against ground truth. The other thing that we found to be very useful is there is often a gap between what you do in research And what you have in production, those systems are not always the same. And so we've worked to, to make sure that at least some of our evaluations that we use in research, we can also run the production system against them. And so whenever we deploy a new model, if it's one of our own models, we can, we'll automatically run those evals and then we wil…
AI assessment note: “So I would say the process that I like to follow is in the beginning”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Since you spend so much time on the tool design, so you have this added tool that can make changes and whatnot. Any learnings from that, that you wish like the AI IDEs would take in? Is there Some special way to like look at files, feed them in.
A I would say the core of that tool is string replace. And so we did a few different experiments with like different ways to specify how to edit a file and string replace. Basically the model has to write out the existing version of the string and then a new version, and that just gets swapped in. We found that to be the most reliable way to do these edits. Other things that we tried were like having the model directly, like write a diff, having the model fully regenerate files. That one is actually the most accurate, but it takes so many tokens. And if you're in a very big file, it's cost prohibitive. There's basically a lot of different ways to sort of represent the same task. And they actually have pretty big differences in terms of like model accuracy. I think either, they have a really good blog where they, uh, they explore some of these different methods for editing files and they post results about them. Um, which I think is interesting, but I think this is like a really good example of the broader idea that like, You need to iterate on tools rather than just a prompt. And I think a lot of people, when they make tools for an LLM, they kind of treat it like they're just writing an API for a computer and it's sort of very minimal. It's sort of just the bare bones of what you'd need. And honestly, like it's so hard for the models to use those. I really, again, I come back to …
AI assessment note: “we did a few different experiments with like different ways to specify how to edit”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q that helps you with deciding what action to take next, like TreeSearch, can kind of help you with reasoning. Any learnings from like, kind of going through all the decomposition ones? Are there state of the art ones? Are there ones that are like I don't know what skeleton of thought is, you know, there's a lot of funny names. Uh, what's the state of the art in the composition?
A Yeah. So the skeleton of thought is actually a bit of a different technique. It has to deal with how to parallelize and improve efficiency of prompts. So not very related to the other ones, but in terms of state of the art, I think something like true thought is state of the art on a number of tasks. Of course, the complexity of implementation and the time it takes can be restrictive. My My favorite simple things to do here are just like in a, let's think, step-by-step, say, like, make sure to break the problem down into sub-problems and then solve each of those sub-problems individually. Something like that, which is just like a zero-shot decomposition prompt, often works pretty well. It becomes more clear how to build a more complicated system, which you could bring in API calls to solve each sub-problem individually and then put them all back in the main prompt. Stuff like that, but starting off simple with decomposition is always good. The other thing that I think is quite notable is the similarity between decomposition and thought generation, because they're kind of both generating intermediate reasoning, and actually over the course of this research paper process, I would sometimes come back to the paper like a couple days later, and someone would have moved all of the decomposition techniques into the thought generation, Section. At some point, I did not agree with this,…
AI assessment note: “in terms of state of the art, I think something like true thought is state of the art”
Answered raw tape
D 5 · C 5 · P 5 · Cm 5 5.00
Q Cool. We cannot leave this section without talking a little bit about automatic prompt engineering. You have some sections in here, but I don't think it's like a big focus of prompts, the prompt report. DSPy is up and coming sort of approach. You explored that in your self study or case study. What do you think about Ape and DSPy?
A Yeah. Before this paper, I thought it's really going to keep being a human thing for quite a while. And that like any optimized prompting approach is just sort of too difficult. And then I spent 20 hours prompt engineering for a task, and Dyspy beat me in 10 minutes, and that's when I changed my mind. I would absolutely recommend using these, Dyspy in particular, because it's just so easy to set up. Really great Python library experience. One limitation, I guess, is that you really need ground truth labels, so it's harder, if not impossible currently, to optimize open generation tasks, so like Writing, writing newsletters, I suppose. It's harder to automatically optimize those, and I'm actually not aware of any approaches that do other than sort of meta-prompting where you go and you say to ChatGPG, here's my prompt, improve it for me. I've seen those. I don't know how well those work. Do you do that?
AI assessment note: “Dyspy beat me in 10 minutes, and that's when I changed my mind.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I think on the customer side, they have a lot of data about their users who has bought, they have the action data. Can you kind of walk us through an example of what does someone come to you for? What questions would they want solved in the process of, do you customize a model for them? Do you have something off the shelf? What does that look like?
A Today, uh, when people leverage our models, it's often to better understand the population of their interest. So usually the start of the relationship Uh, we basically come together and hear about what population they want us to model, right? Um, so it might be that if you're a CPG company that's selling to all of the US, that might be fairly straightforward. You want to model the gem pop of the US, but at the same time, there is a vertical, or if there's a market that they're trying to go into, imagine, ah, they want to better understand, let's say people in their twenties and thirties living in California. That's a much more specific population. So we hear about these population and we go recruit these People, uh, with consent, uh, and with incentives, and we basically collect some of their data and create a model of these people. And then what our product allows you to do is basically query them, uh, so it can take us input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. Environment can literally be a survey questions. It can be behavioral experiments. It can be A-B testing. Oftentimes, the core use cases are things like concept testing, uh, to start with, but also, you know, people sometimes want to do focus group, or one of the sort of fun use cases that we also serve is actually even modeli…
AI assessment note: “when people leverage our models, it's often to better understand the population of their interest”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q your own solution to this, but how far off are we, and how do you check if it's grounded? Uh, you have some interesting stuff on your site that actually points to how you run real evals, but If you could take us through that side, you know, I think that's one of the big concerns that people have. They're like LLM's hallucinate. You're just hallucinating layer after layer, right?
A The way we do this, and this is actually the, the paper that we worked on after the Generative Agents paper that really became the, at least for a simile and also the field of simulation and synthetic panels, really became the foundation. Yeah, this is the paper. The paper is called Generative Agents Simulations of Thousand People. Here's what we've done. For this paper, we actually brought thousand people that's representatively sample from the US to a virtual app. And what we basically have done was we spent two hours collecting fairly wide-ranging data. In this particular study, we focused a lot on this interview data, uh, that was, uh, whose script was taken from this project called American Voices Project. And then we would also pair it up with a lot of behavior data and so forth, whatever we can collect within two hours. And then we would actually send these people away for a couple of weeks. And during that time, I would use this data to create their digital twins. And I would bring the humans participants back after two weeks and have them complete a battery of surveys, experiments, behavior studies. So we actually have the list here, which basically included things like the behavior economic games. We would run literally like big five personality tests, general social survey. We would also go ahead and run the randomized control trials that were published on PNAS. And …
AI assessment note: “we basically could replicate people's behaviors and attitudes, 85% as accurately”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Tell us about the company. You guys just raised a lot. You're half a research lab, half a company. I guess you're hiring. Are you based?
A Yeah. So we're based in Mission Rock. Uh, so not too far away from where we are right now. So we're in SF. Uh, but we are also bi-coastal. So we have our, uh, team, uh, I would say our headquarters in NSF, and we have a lot of our technical talent in NSF, and we do have a smaller office that just opened up actually in New York. We are, as a company, an interesting one in that today, obviously, there are AI Neolabs, and then there are AI product companies. Similarly, it truly is both. So this is a company that was founded by four co-founders, myself, Michael Bernstein, Percy Leung, Laney Yellen, Uh, Michael Percy and I are all researchers. So, of course, Michael was one of the co-authors of the ImageNet, kicks her the AI revolution back in 2013, has been instrumental in human-centered AI. Percy coined the term foundation model, and obviously, you know, one of the, the greats of the AI researchers today. And Lainey is my business counterpart, where she led some of the fastest growing AI native companies from their C to A and B. But we have this DNA at the company where The vision of the technology that we're creating is continuously developing that we are getting people who are basically my lab mates. We are right now about 60 or so people. 15%, almost 20% of the company population actually are just my lab mates. And we are, it's actually quite fun because many of them then had g…
AI assessment note: “we're based in Mission Rock... founded by four co-founders”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Do we want to keep going on the paper, uh, routes?
A Yeah, for sure. Uh, so the last one, uh, was sort of an interesting one. So this, uh, paper was the follow-up paper that we had, uh, to the Thousand Agents paper, where basically the idea was now can we augment the models even further and actually post-train the model based on a lot of randomized controlled trials. So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there's this, and there's this platform called Open Science Foundation. So some, uh, the audience might be familiar with this, and there has been, especially in the social sciences over the past five years or so, there has been this concern around replicability of studies, right? So it was a bit of a crisis that scientists acknowledged where we rerun the study and we don't actually see the same finding. It's, it's tough. And the reason why that was often the case was there's basically this survival bias where the papers that get published often need to maintain what we call the p-value of less than 0.05 in the experiments that we ran. That basically suggests that only, there's only five percent chance that the results that we saw is false positive. But the tricky part was all the papers that were not published, and there's still a five percent chance That whatever we publish is actually totally just randomly generated. Like there's a fi…
AI assessment note: “Yeah, for sure. Uh, so the last one, uh, was sort of an interesting one.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q talked about all these steps of, okay, you gotta do quantization, train your own speculative decoder, run on different hardware. Um, looking at other model providers, Okay, you kicked off a inference speed race. On the consumer end, um, what goes into keeping quality the same across them, right? Sure, you can run benchmarks, but like, how do you determine how much quantization are the standards? What actually goes into?
A There's a few things on quality. Most inference optimizations are lossless. KV caching, for example, you are just recomputing or preventing recomputing the same values. Speculation, of course, If a draft token is wrong, it gets rejected. The main lossy optimization is quantization, um, and that really comes down to number one, data format, uh, number two, uh, which parts of the model you choose to quantize, which layers, and number three, like, doing a lot of calibration on the quantized weights, uh, to ensure that you're sort of preserving all the outliers. There's other sort of tricks that you can do, though. A big one is long context, because one thing you asked right at the beginning is, oh, what's going to happen if I send a 200,000 token request in? So obviously with a long input sequence, you need to, you know, store a lot more information, you need to process a lot more tokens, and so even if a model has a context of a certain length, You might, as an inference writer, choose to build an API with a shorter context length, and of course a full length one as well, because if someone doesn't need the full million token context, for example, you can get them better performance. I don't know if that's exactly like quality of the model. The way that I think about quality is to what degree are we faithfully serving the original model? If you think of a sort of golden implement…
AI assessment note: “The main lossy optimization is quantization, um, and that really comes down to”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And how far does that get you? And how, how easy is that for the average person to do? So say right now, I want to throw the weights of GLM five two on a node of B 200. How easy is it to find speculative decoder, decoder model or already quantized model? How much work goes into it?
A If you're doing it up front, it's quite a lot of work. If you're doing it today, there's going to be people who have published Things that you can just, you can just grab some NVFP four weights. You can grab a speculator. Uh, yeah. If we're thinking about like, what are the two X's we're stacking? Going from, uh, BF-sixteen to NVFP four is, it's not quite a two X, right? It's like, I think it's about like 30, 30 to 40%, um, from 16 to eight, and then another 30 to 40% multiplied from, um, eight to four. So that doesn't quite get you a two X, but like roughly a two X. Speculator, roughly a two X. Disag on top of that, if you're able to get enough hardware and put enough traffic through it, another roughly a two X. And then you add in some, you know, double digit percent increase from having just a, a better run time with, you know, the, the latest kernels and stuff behind it. Um, and that, that's kind of how it stacks up. Um, so building each of those, like building the, Um, quantized weights is, for someone who really knows what they're doing, hours to days of work, um, building the speculator, again, like, hours to days of work, and the, um, dis-ag setup, hours to days. Well, uh, ok, once, once you have it, yeah,
AI assessment note: “If you're doing it up front, it's quite a lot of work.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Anything else on the support side, when you, when you say, like, get it to fully production ready?
A Yeah, I think that, There's also a question of just, you know, we can test a model to a pretty extensive degree, but we're trying to get it out quickly, and then you see a bunch of other people test it, and you get interesting results. There was an issue with, um, GLM briefly, where we had some, like, mode collapses where it would just output the same token over and over again for certain prompts on certain temperatures, like, Once you expose an endpoint to, to the real world, there's going to be, you know, so many more varieties of, of things given to it that, that you're able to, you know, discover and, and patch things. So it's not just a, you know, day zero process. It's then like for the first week, for the first month, if a model remains popular, like how do you both fix bugs and then continue to push the envelope on performance?
AI assessment note: “There was an issue with, um, GLM briefly, where we had some, like, mode collapses”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q People used to say that you would also do Franken merges. Where you would take, like, layers from each model. Does anyone do that anymore?
A Well, to your point previously, when you were mentioning, like, um, the work that goes into supporting a model when it first comes out, like JLM-VII or Minimax M-III or whatever the case is, sometimes you do have to, like, you do have to switch out some things. Like, for instance, the Minimax M-III head uses full attention. And with full attention, you end up with this, like, insane bottleneck inspector because you're doing autoregressive token generation for three tokens, and you're doing this, like, like O of N squared over all of the tokens that are in your sequence. Um, your KV cache is, like, very large because it's not sparse, it's not OK. So we find it better to, like, okay, we're going to replace this, you know, we're going to replace this layer with a layer from another model that's using, like, GQA, for instance. And then just for the right training, you can get it to have the same acceptance rate. So it is, it is very possible to, to retrofit layers from other models, and very much needed, actually. If a layer is, like, inefficient, the training just becomes the challenge. Like, how do you ensure that you train it properly? Which again, to earlier points, is like, the mesh between training and inference. As in, like, you need very good training in order to do fast inference. That's, like, I feel like more and more becoming true.
AI assessment note: “we're going to replace this layer with a layer from another model”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Uh, okay. One more thing while this chart is still up. All gather, all reduce is expensive. One of the things that is a movement in Silicon Valley is mega kernels. Just keep fusing kernels. I don't know. Is it that simple?
A Well, I mean, like a fused kernel can't save you. Like, like here with the tensor parallelism, you're the half, the matrix is one GPU and the other half is on another. And if I need the entire matrix in order to do like a non-linear operation on the next step, which is for instance, like if I'm doing attention, I need the soft max or I need to like, like exponentiation, I need to have the entire row. So I need to know what the partial result was from GPU two and what the partial result was from GPU one in order to be able to do the soft class in the next stage. So I, I like, I have to make them communicate with each other, even if I had a fused kernel because of the nonlinearities within each one. Also with, with like mega kernels, like honestly, I'm, I'm, I'm very bearish on, I'll be honest.
AI assessment note: “Well, I mean, like a fused kernel can't save you.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q in video compression where basically frame by frame, there's not that much difference. So actually you don't have to regenerate or resave the whole frame, right? Um, I think MP four compression or something else like that. Is it tempting to use that? Or as far as I can tell, everyone just treats it as, no, we will just generate every frame. Is that roughly the state of the art?
A There are a few different approaches. Let's say first, like you, you want to just directly use MP four compression and you use that as the tokens for The transformers to train, right? So people actually have tried that, but the, the main challenge is the latent space for the MP four tokens are not, we're not very comprehensible for the models. It's extremely hard to train on that. And there's a. So that's why they created VAEs, which creates more continuous latent space. So the models can understand that latent space and learn from it much easier. Even within the VAEs, there are different difficulties of the latent space. So you, you can imagine something that the simplest, the most naive VAE is like you, you have an image and you just shuffle all of the images into a Into a vector. So you don't need to train any of these, right? But that latent space is extremely hard for models to train on top of. So that, that's why there's some debate on like, how do you compress the, the tokens? So, so you mentioned like you can compress frame by frame. Also you can compress, uh, the temporal dimension. Yes. The difference is if you compress the temporal dimension, you, you get a much higher compression rate. Because there is temporal redundancy between frames, because this frame and the last frame, likely they are mostly similar. So there's only some small difference. Uh, for example, lik…
AI assessment note: “people actually have tried that, but the, the main challenge is the latent space”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q How many people do people spit up consecutively? So we have currently, I guess is the concurrency.
A So there's, there's three metrics that we look at. And so one is like time to spin up one. And so our time to spin up one is 60 milliseconds with network agency. So requests, spin up, reply, 60, the whole thing, 60 milliseconds. That is one. But if you want to spin up 50,000 at once, we are now at about 75 seconds. So it takes about 75 seconds to spin up concurrently 50,000. Some others, there's public data around this, like take 2000 seconds, which is 30 minutes. Like there's different variations of that. And then there is that. So it is speed of one, speed of like multiple, and then how many can you consistently have up and running? And so we basically have right now no limit to how much we can add because we basically own our own metal. But the biggest customer of ours does like about 850,000 every single day is sort of where they were there just shy of a million every single day that they're running. Um, we do have a request for half a million concurrent, which is literally half a million CPUs somewhere running. So that's an interesting.
AI assessment note: “takes about 75 seconds to spin up concurrently 50,000”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah, weird pockets. So actually, but that helps you to distribute your load through all time, you know?
A Yeah, so the interesting thing is that we have those kinds of loads, but if you look at the researcher loads, they're quite different. So what they are is like, if you give them concurrency of 10,000, or 50,000, or a 100,000 CPUs, whatever it may be, when they fire off a, a run, it's just a hundred percent, and then just runs, runs, runs, and then it stops. So it's very, The usage pattern is squares basically, right? And it's also not follow the sun because people will fire it off at midnight before they go to sleep, but then wake up and so it's very unpredictable. So you don't know where that is. So the shapes of the usage are quite different than we have had before. And also what's interesting is when it's sort of a fall of the sun, even if you have a high growth company, you can sort of predict your usage patterns and, and have enough Capacity for that because it's sort of, it grows in a, in a way you can project when you have companies doing sort of like evals and RL, they're super spiky. So they're going to come in. It's like, we're going to use nothing. Then can we have a 100,000, right? And then go back down and then have a thousand and get back down. So it's very, very different. Right.
AI assessment note: “we have those kinds of loads, but if you look at the researcher loads”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. So for, for listeners, uh, what do you guys do just to situate you guys in the, in the company?
A Uh, Bridge is a clinical intelligence layer for health systems. We really started with documentation and building for clinicians, and we think that, you know, as we think about reducing the burden that clinicians have, they're spending 10 to 20 hours a week on documentation. There's a massive doctor shortage in the country. We also think that conversations between patients and clinicians are probably the most important workflow in healthcare. It's, Obviously, where care is given and received, but if you think about the 20% of our GDP that goes towards healthcare, almost everything is a derivative of that conversation, whether it's the claim, the payment, the actual diagnosis given the treatment, and we've started with a conversation to reduce the burden for doctors on documentation, but we're really excited about the path ahead as we become this broader clinical intelligence layer.
AI assessment note: “Bridge is a clinical intelligence layer for health systems. We really started with documentation”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Do you have, um, failing evals that you're just hoping that will have success eventually when a good model comes out?
A I mean, yeah, so I think, I mean, I could talk about this for 60 minutes, so I will limit myself. I think it's a real issue when people say evals, and it's just like, that's quality, that's like unit, I mean, it's like saying testing, it's not just unit tests, right? So we have the equivalent of unit tests, regression tests, those live in CI, those have to pass a certain percent, you know, within some stochastic error rate. Then we have, as you're building a product evals of these aren't passing right now, and this is launch quality. So we have a report card, and we need to, on these categories, you know, be at 80 or 90% of all of these user journeys to launch. And then what we have, what we call Frontier Headroom evals, where we actively want to be at 30% pass rate. And that's actually been a effort that we took in partnership with Anthropic and OpenAI in the past maybe two or three months, because we actually hit a point where our evals were saturated, and we weren't able to really give insightful feedback other than it wasn't worse. And not only is that not helpful for our partners, it's not helpful for us to understand where the stream is going, you know, going back to that analogy. And so we spent a lot of time thinking about what Notion's last exam looks like, right? Not just Humanity's last exam, but Notion's last exam. And, um, there's a lot of, you know, dreams about w…
AI assessment note: “we have, what we call Frontier Headroom evals, where we actively want to be at 30% pass rate”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Talking about integrations, you prompted me, so I got to ask MCP, CLI, what's going on? What's the opinion?
A I think, I mean, I'm, I'm definitely bullish and excited about CLIs. I think there's a few really cool things about CLI. So one really cool thing is like, um, is that it's in the terminal environment, so it gets a bunch of extra power. So, you know, for example, I can like, like paginating and cursor through like long outputs. Um, and it has a progressive disclosure inherently. Uh, so, you know, You don't see all the tools at once. It's just, you see the CLI wrapper and you can like use the help commands and, and, and read files. And then I think the most important thing that's, that's super cool is that there it's also inherently a bootstrapped. So if there's an issue of the agent can debug and fix itself within the same environment that it uses the tool, right? Like, you know, I think I saw a tweet this morning. Someone said, you know, my agent didn't have a browser. So I asked it to make it a little browser tool. And within a hundred lines of code, it gave us a little browser, like, like wrapping the, the, the Chromium API. Um, that's pretty incredible. And then if there was a bug, it would just immediately try to fix it.
AI assessment note: “I'm definitely bullish and excited about CLIs. I think there's a few really cool”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q kind of like the multiplayer mode? I think subagents is like, Single players split up the context. Yeah. And the multiplayer coworker is like, my colleague has some file on their machine that I want to know about, or I want to know how their task is going to then update my thing. Like, is that interesting? Is that something that makes sense for you to build or for like?
A It's like super interesting to me. It, it almost goes back to like some of the scaffolding where I'm like, okay, are we going to be, end up, are we, will we end up building scaffolding that will just go away? And like a question I have here is, At what point do we just assign these things like their own Gmail account? We'll just give them their own like Slack handle, and then they will just like use the same tools we humans use to interact with each other. You mentioned our finance people. They've been working pretty hard on very good office integrations. And I think for a while we've been like, we built so much tech around Claude leaving useful comments inside a Google Doc, and now it just does it, just like leaves a comment in your Google Doc and that's how you interact with it. Maybe like the similar thing where I still have open questions around what is the best interaction mode? Is it for us to be something super custom for co-work agents to talk to each other? Or is it, okay, let's just jump straight to the finish line and say, well, we're just going to give this thing, if you use Slack at work, we're just going to give this thing a Slack handle, and that's going to be the way it's like multiplayer capable.
AI assessment note: “It's like super interesting to me. It, it almost goes back to like”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Aww. Yeah, yeah, yeah. No, well, she, she made a good choice there. Uh, was that, like, basically the origin story for Launchables? Is that, we, and maybe we should explain what Breve is.
A Yeah, I mean, Breve is just, it's a developer tool that makes it really easy to get a GPU. We connect a bunch of different GPU sources. So the basics of it is like, how quickly can we SSH you into a, into a GPU? And, um, whenever we would talk to users, they wanted a GPU, they wanted an A-one hundred. And if you go to like any cloud provisioning page, usually it's like three pages of forms or in the form somewhere, there's a dropdown and in the dropdown, there's some weird code that, you know, to translate to an A-one hundred. And I remember just thinking like every time someone says they want an A-one hundred, like the piece of text that they're telling me that they want is like stuffed away in the corner. And so we were like, what if the biggest piece of text was what the user's asking for? And so when you go to Brev, it's just big GPU chips with the type that you want.
AI assessment note: “Breve is just, it's a developer tool that makes it really easy to get a GPU.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. I mean, do you want to tell that story? I don't know if you've told the story. It's like a random intro, right?
A Well, it was just, he used to have house parties. Uh, TechCrunch had, had these house parties and And it was, um, probably no different than somebody's doing a house party in SF. Uh, you just go and you meet the VCs and founders and like, I'm going to make up examples. So I don't want to like, you know, there'd be like Chad Hurley over there pitching his, you know, YouTube to people. And like, oh, like that's just like how it worked. And it was just like, wow. Like that was this era where all these new companies were, were emerging. And I met our first investor, uh, in Silicon Valley at one of these house parties, Emily Melton, who then brought us into DFJ.
AI assessment note: “I met our first investor, uh, in Silicon Valley at one of these house parties”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q found out my friend from OpenAI took a break a year off to bike through Japan. How could you? Like, you're gonna miss it. You're gonna miss everything. But he's like, I'm good, you know, like, like, I'm, you know, having kids, whatever. You did a sabbatical as well, and like, it was pre-AI, but, uh, it was interesting. I, I, you did, you did the Appalachian Trail. Which one?
A Uh, so there's three big ones in the United States. It's the Appalachian, the Pacific Crest Trail, and then there's a Continental Divide Trail. So I did the Continental Divide Trail, which is the longest and most remote of the three. Sometimes considered like the, the, the older, bad, whatever. But like, honestly, the PC, they're all different trails. Like I'm, I'm pretty steeped in hiking culture. Um, I think mile for mile AT is actually the hardest, but I did the CDT as my first trail, as my first through hike. Um, You know, you learn a little bit about the three when I was choosing which one I wanted to do, and the CDT was the one that scared me the most. Uh, I was like, hey, this would be the hardest, biggest accomplishment I could possibly imagine, and I thought, if I never have an opportunity to ever do this ever again, which so far seems to be pretty correct, um, which one I'm going to do to feel the most like, hey, I did the thing that I really wanted to do, because I've always wanted to do a long-distance hike, and so I, I chose the Content of Divide Trail. I did that in 2021, Pre-AI, and.
AI assessment note: “So I did the Continental Divide Trail, which is the longest and most remote”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q much different for an LLM because you can expect a person to look at maybe the Five, six links in a Google search versus for an LLM, should you expect to have 20 links that are highly relevant? Like, how do you internally figure out, you know, how do we build the AI mode that is like maybe like much broader search and span versus like the more human one?
A Yeah, I mean, I think even pre-language model based work, you know, our ranking systems would be built to start with a giant number of web pages in our index. Many of them are not relevant. So you identify a subset of them that are relevant with very lightweight kinds of methods. You know, you're down to like, 30,000 documents or something. And then you gradually refine that to apply more and more sophisticated algorithms and more and more sophisticated Sort of signals of various kinds in order to get down to ultimately what you show, which is, you know, the final 10 results or, you know, 10 results plus other kinds of information. And I think an LLM based system is not going to be that dissimilar, right? You're going to tend to trillions of tokens, but you're going to want to identify, you know, what are the 30,000 ish documents that are with the, you know, uh, maybe Thirty million interesting tokens. And then how do you go from that into what are the 117 documents I really should be paying attention to in order to carry out the tasks that the user has asked me to do? Um, and I think, you know, you can, you can imagine systems where you have, you know, a lot of, uh, highly parallel processing to identify those initial 30,000 candidates, maybe with very lightweight kinds of models. Um, then you have some system that sort of helps you narrow down from 30,000 to the 117, uh, with…
AI assessment note: “an LLM based system is not going to be that dissimilar”