The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

1,847exchanges match
1,797on raw tape
133redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Great. What's the current state of the Neumontron family? You got Nano, you got Super, you got Ultra. What do those models do and what are the use cases for them?

A So Nano is a thirty billion, um, uh, total three billion active parameter model. Super is one 20 and 12, and Ultra is five 50 and 55. Um, they're designed, um, really to fit, you know, it's kind of small, medium, and large, um, uh, deployment scenarios. Um, uh, you know, nano can be really capable for things that, um, you know, don't require nearly as much knowledge or reasoning, but obviously for the, for the most, um, capable model, you go for ultra. Um, super in a lot of ways is our most popular model because it represents kind of a great balance between, Um, cost and, and intelligence. So we, we kind of like, um, having this small, medium, and large, um, approach to building a family just because our customers, um, seem to respond to that, um, pretty well. But, um, you know, uh, the most important thing from NVIDIA's point of view that people are doing with, uh, uh, with LLMs is agents, right, is, um, building agentic workflows Having it, having an agent working on your behalf, solving problems for you night and day, um, is such an exciting way of approaching the problems that we have to solve. Um, and, um, it's our dream to make NemoTron amazing for that purpose. That's, that's our goal.

AI assessment note: “Nano is a thirty billion, um, uh, total three billion active parameter model”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Another important characteristic of Numatron, free ultra is a one million token context, the, the long context window. How important is that in the overall mix and what does it enable the model to do?

A The longer the context length, the more challenging problems we can solve with a language model, um, that allows us to do things like append all sorts of information to a query, which could be a code base. It could be instructions. Um, you know, uh, in, in the longterm, I'm hoping that I have my own personal LLM that's able to read all of my emails, you know, and help me answer questions about that. You know, the more information that we can attach to a particular query, Um, the, the more useful the model can be. Um, now, uh, it can get more and more expensive, right, to reason over large amounts of, of input data. And, um, and so that's one of the, the reasons why there's usually a limit on how big the context length can be. But with Nemotron III, we, we tried to push it as far as we could go. Uh, we think a million tokens is a lot of tokens, um, and you can do a lot of things with that.

AI assessment note: “allows us to do things like append all sorts of information to a query”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q big question is, can they become great at, uh, Law and consulting, and then, you know, all sorts of different domains, and part of the black box of closed models is like how people go about doing all of this, where do they get the data from? To the extent that you can talk about all of this, I'd be very curious about how you guys have gone about it.

A It's not an easy question to answer because it is quite complex, but I would say, um, we rely on a number of things. One is that we do purchase data from, uh, companies that, um, uh, that, you know, are, are building data sets that you can purchase. Um, and to the extent that, you know, we have the rights to redistribute, uh, or to, to, to open up that data, we do as part of, um, our, our, uh, Mnemotron data effort. Um, you know, with, with Nemo Tron, we are trying to be maximally open with the data that we release because our goal is to support the ecosystem, right? Our goal is, is not to be the only model out there, and we love it when we hear of other models around the industry that are using our data sets, um, to, to make their AI stronger, because that means we're succeeding in our job to keep the ecosystem thriving and growing. Um, now, uh, We also are big believers in synthetic data generation. Um, we use an enormous amount of compute, um, uh, running language models on our own systems to create synthetic data that then helps our models be better at, uh, solving problems in specific domains, and we release a lot of that data as well. Now it's, of course, not very straightforward to do this. Like, you know, AI is always garbage in, garbage out. So you have to work really hard to make sure that any synthetic data that you create is actually adding value. That's actually he…

AI assessment note: “we rely on a number of things. One is that we do purchase data”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. And what, what does that change from, I guess, an infrastructure perspective? Like what you just described, uh, 5000 sites versus five. In some ways that sounds Inefficient? I mean, it's efficient for me as a user, ultimately, because I get the best answer. What does that change in terms of requirements for the infrastructure?

A Yeah, I think, I mean, first of all, there, there, if, if this trend continues, there's just going to be an enormous amount of additional demand on, on the internet. And I go back to, you know, during COVID, where over the course of two weeks, we saw internet traffic double. Um, you know, if, if, if our projections are right, Um, that's gonna seem quaint very soon. Um, you know, in five years you might have a thousand times as much traffic on the internet as you do today, and that's gonna mean more, you know, servers, more network infrastructure, more CPUs, more GPUs, more memory, all of the things that have to serve these bots are, are gonna be very important. And it's also gonna mean we're gonna have to figure out some new business model in order to fund that. Someone has to pay for it. And, um, and, you know, the, the business model of the internet Historically has been ads and bots don't click on ads. So it's, it's going to be something different going forward. And, um, you know, I, I think the most interesting question, uh, that in the world today is over the next five years, the business model of the internet is going to change radically and what it changes to, um, is, is, is totally undefined at this point.

AI assessment note: “more servers, more network infrastructure, more CPUs, more GPUs, more memory”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And do you think that there is a painful transition period from the current world to the next world? So do you think everybody's going to have to go through this as organizations recalibrate?

A We laid off more than 20% of our, of our team, uh, and not because the business was struggling, but, and not because they weren't great team members. Um, but we just need fewer middle managers. We need fewer measurers. Um, and, um, And those roles were going away. And as I talked to peers at other companies, they're all seeing the same thing and they're all saying, yeah, we're gonna have to do the same thing at some point. And you're, and you're seeing these sort of trickle out, but there's a real fear in leaders out there that they don't want to be the first ones, um, to, to, to do that. And, um, you know, pardon the vernacular, but I think that's chicken shit because The cruelest thing that you can do is wait. Um, I, I do think in the next six to 12 months, almost every company is going to go through some exercise like this where they're going to cut a bunch of their, their team. Uh, and, and I think a lot of people are wait, a lot of, a lot of CEOs are waiting around to do it because they're afraid that they're going to look bad in the process. What really struck me Was the realization that once you know that that change is going to happen, the kindest thing that you can do for your, your team is to do it as soon as possible, because it's way easier to get a job today than it's going to be in six to 12 months because the market's going to get flooded. And so we put together …

AI assessment note: “in the next six to 12 months, almost every company is going to go through”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you need for performance reasons to be close to the customer the way you, you need to have regions in cloud?

A You know, it's super interesting. A lot. I get this question a lot and people they're like, well, does latency matter? Does, so I'll tell you what, what matters and what doesn't matter. You can look at your own utilization of whether it's ChatGPT or Claude or Grok or Gemini, and you can see, hey, a lot of the things that I'm doing, I kind of shoot it off. I come back later and there's a research report for me. Maybe it's a long running agent workflow. In those cases, latency doesn't matter at all. The only thing that matters is your cost per token. That's all that matters. And, um, So that's been a really interesting change. I think that, you know, the old school traditional legacy cloud business was so latency focused because of some of the applications, but this new fleet of AI applications are far less latency sensitive. So that's one, but there is the caveat, which is this governance and data governance is becoming an important thing. And a lot of countries are wanting to have the AI compute that their citizens are using Be run out of their own country so that they can, you know, at least have their own, their perception of control or whatever. And the, you know, that is, that is another, that is an element to it, but I'd say that from the latency, there's no technical reasons.

AI assessment note: “from the latency, there's no technical reasons.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Are some of the people that were there at the beginning still around? I think you started the company with your brother. Is that right? And your brother is still at the company.

A Yeah. And so in terms of like the early people, um, basically It's not, I mean, not even basically of the four people who are making DreamScope, me, Michael Balaban, my co-founder and fraternal twin brother, Chuan Li, who's our chief scientific officer, and then Steve Clarkson, who's, um, an engineering leader at the company and, you know, has a bunch of folks reporting into him. Uh, now, you know, they're all still at the company. Um, the next hire, one of those, uh, gentlemen named Mitesh, Uh, Agrawal, who's one of the, the next hires in that team. Um, he was with the company for maybe eight years or something like this. Um, yeah, something like eight years. And, uh, then he eventually, uh, left and, and joined another former Lambda team member, Thomas Summers to start Positron, which is, um, uh, an accelerator company. And they're like now valued at over a billion dollars. And, uh, so, uh, not only has like the original team stuck around, but we've already started to kind of see what like a Lambda, uh, alumni, a Lambda mafia network looks like in, in the world, Lambda lab member alumni.

AI assessment note: “they're all still at the company.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you worry about model training and model inference becoming, I don't know, 10 X more compute efficient and what that would mean in terms of the buildup?

A I think that generally speaking, What you're seeing is if, let's say you do become 10 times more efficient, I think that that just means that everybody is able to process 10 times more tokens, and there's, there's still the same fixed amount of compute in the world at any given point in time, and so in the early days, it's funny, we used to talk a lot about this back in, let's say, 2017. Oh, well, maybe there's gonna be some new type of model, let's say, that will look more like a random forest model, which The audience might some, some members of the audience might know you can kind of train a random forest model on a MacBook, right? And there was, there was this concern that was kind of persistently raised around like, well, okay, what happens if you have this sort of like, um, adjacent disruption on the model side of things? And so far we haven't seen that. And again, everything that we're building towards is, Sort of based on these scaling laws, which is really about scaling up this architecture. So, um, I don't really foresee a very likely outcome where we have this huge model disruption that would cause a decline in the demand for compute.

AI assessment note: “I don't really foresee a very likely outcome where we have this huge model disruption”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Since you mentioned, uh, test time compute, I think there's, um, something that, that still puzzles people, which is the The whole chain of thought thing which is so magical from a user perspective, whatever you can see, what actually happens during test time compute that creates those artifacts? What does the model actually do?

A I think it does what you see it do. We lightly rewrite it or summarize it, but it just, it just produces tokens, and those tokens are like a running thought process, just like You might have, or, or maybe it's more akin to, if you're solving a math problem, the, the scratch pad, the collection of notes that, that you have, but it, it just keeps generating. The cool thing about generating is that, you know, it, it, it, it's, uh, a forward pass to the model. So we're using a bunch of computation. So we're, we're, you know, it's a way of leveraging a lot more computation on a problem. Then, then you would before. So my colleague, Noam Brown likes to talk about the Riemann hypothesis a lot. And, you know, wouldn't you want to have a model that runs for years that, that, that can, um, resolve, resolve that, prove that if you present it and you want it to produce an answer, then it only has the number of flops in a, in a single forward pass to produce one, one token if it's forced to answer right away. But if it gets to answer after thing, you know, after a long time, it can, it can Re, reuse its weights, you know, produce, um, a final answer that is a function of a much, much larger amount of computation, and the, like, the natural way it thinks is in language. It's a language model, and so that's sort of this key insight that, that you can, um, cause it to do better just by produci…

AI assessment note: “it just produces tokens, and those tokens are like a running thought process”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you think there could be an equivalent in AI to thermodynamics, meaning, uh, you know, a compact theory that, uh, predicts behavior without tracking every individual bit?

A Yeah. Kaplan McCandlish scaling, open AI, uh, scaling laws work originally is, is a version of this where you throw away, you know, all you know about the network is how many parameters it is and how, how much you've, how much data you've trained it on it. And you, you can predict like the, the final loss. I think the, the missing piece Is going from all the individual weights and biases and, and how does that add up to the scaling law? I have some very like initial work and there's some other initial work about like trying to bridge that connection. But like, I think that's, that's the missing piece, like the sort of statistical mechanics to thermodynamics of how do we like, how do these things emerge? But there's definitely a lot of useful effective descriptions of how these systems behave. I think the other part of your question is like, Is it, is enough to characterize everything that we care about, right? There's probably a lot, there's a lot that we care about other than just the final loss function, and so there's, there's more thermodynamics to be worked out. In addition to like, how does the thermodynamics arise from the microscopic description?

AI assessment note: “Yeah. Kaplan McCandlish scaling, open AI, uh, scaling laws work originally is, is a version of this”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Thank you for that. Where do you think we are in the evolution of AI being increasingly able to solve difficult scientific problems? I mean, certainly something that we've been talking as an industry about for a while now, but it seems to be accelerating, perhaps just like everything else in AI. But where do you think we are?

A I think one of the interesting things is that this process is smooth. The, there's no sharp point, or I don't think there will be a sharp point where we'll say that systems didn't, weren't able to be useful for scientific, the scientific process to their fully fledged scientists. There'll be sort of a gradual shift. If you had to point to one moment, maybe it would be the release of O-one and by OpenAI and, and the sort of paradigm of test time compute and, and, and reasoning. But I'm sure if I tried to make that claim, You could go and look at GPT-IV and, and see that there's, um, glimpses of that sort of useful behavior for, for the scientific process were, were already present. As a general point, you know, the, the models are very good at certain types of things that clearly are amenable to, to making progress in math. They're not open loop, fully fledged scientists in, in any domain, although, you know, neither am I. It seems like it's just this really nice gradual process.

AI assessment note: “There'll be sort of a gradual shift.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And do you think it's necessarily humans have a seat and agents have consumption, or would there be an argument for saying that, uh, agents in some way are not that dissimilar from, from humans, although they, they'll be doing a lot more with a lot more, uh, volume of data, and therefore there should be some kind of, like, seat-based pricing for agents?

A I think, I think this was, this is sort of a tougher category because, um, it all depends on the agentic use case. So, um, like I can totally see a world where we already have some customers playing around this idea of like, should agents have a box seat? Because, because why? Because they actually need to store data, uh, that gets retained and governed over a long period of time. And you want to be able to track it and manage it just like a person, but it's gotta be stateful. And so that kind of makes sense as like, we have to give it a name and a thing in our system to make that work. Do we charge the same as a regular end user seat? Probably not. Probably it's gotta be cheaper. Um, but then there's a lot of situations where the agent doesn't need an ongoing seat. They just need to be doing a lot of operations, in which case it's, it's probably just pure consumption. So I think it really depends on where does your software category land on? Is there a reason why you'd have an agent be stateful in that organization? Um, and, and, and kind of take on an identity and take on ongoing work. Uh, versus it's a thing that just every employee calls on demand. And that, that would probably determine, you know, what that business model looks like.

AI assessment note: “it all depends on the agentic use case.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q and then we did a pilot and that never really worked out. So now you're coming back to me like two or three years later and you say, no, no, no, no, no, no. Agent is the thing. And like this time, if you don't do it, you're going to, you're going to die. In the spectrum between skeptical and enthusiastic, where would you put the mood in large enterprises?

A I would say if I did like the broadest sample, I think the mood would veer statistically more optimistic than maybe the framing that I think maybe you landed on like a few extra cynical CIOs in that. I think because what's happening is the CIO and our main audience is the CIO. Uh, when we talk to the CIO, they know that their engineering teams are using cloud code and codex and cursor, and they're seeing the productivity gains come out of those teams. And they're like, yeah, like my teams are just building, you know, apps way faster. They're, they're being able to tackle IT projects much more quickly. Like we're, we're, we're doing security reviews faster. Like they're seeing the productivity gains in their, in their function. And I think they're often saying, well, how do I bring those same IT, you know, productivity gains To the non-IT parts of the organization, and they're having the business pull them and say, I want access to, to co-work. I want access to codex. I want access to these tools as well. So there's, there's actually a certain kind of sex appeal to, to these tools right now where the business is sort of demanding. I want to be on the agentic train because I'm seeing all these, these great use cases. And so I, I think the tone is actually remarkably optimistic and excited and positive as opposed to, you know, there, there's a sort of, You know, typical trough of …

AI assessment note: “I think the mood would veer statistically more optimistic”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q IPOs and we've seen companies that are compounding faster than ever. Where do you think that leaves startups, including vertical startups? Where do you see the opportunities? Are we in a, in a world where everybody is ultimately either an open AI or an anthropic employee? Or, you know, service, uh, industry, uh, supporting them? Or is there room for lots of people to do, uh, lots of different things?

A I remain pretty, pretty confident and optimistic on the, the need for a kind of a bridge layer from the AI capability to the end user workflow. And, and some might sort of say that this gets kind of bitter lessened out, um, which is, which is, you know, oh, these things are wrappers on the model. And, and at some point there's a training run where it just like is the final training run That makes the, the renders the, the kind of vertical app or, or function specific, you know, app, uh, you know, not as useful. And, um, and I think that is a little bit too much of an accelerationist view of what people are doing with the tool, which is like, it's not just like what the model is spitting out or the model's ability to review information. It is how was the thing wired up into the business workflow? How did it get the context that it needed to be useful? I think if you're in a, in an, in an industry or a line of business There's a heavy amount of, of kind of integration with data sets, heavy amount of, of kind of bespoke workflows that that company does. That usually means that there's going to be a need for change management, implementation, ongoing support, ongoing expertise. And unless the labs build out literally the equivalent of hundreds or thousands of people for every single vertical and every single line of business, that means that there's actually a lot of opportunity in…

AI assessment note: “I remain pretty, pretty confident and optimistic on the, the need for a kind of a bridge layer”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q and I'm, I'm curious, What reasoning means in 2026 that's any different from, you know, a conversation we could have had about oh one or three. Um, in particular, one of the claims, uh, of 5.5 and, and also my experience as a user is that it's particularly good with, with messy data, which seems to imply that, um, it needs to reason through ambiguity more. Um, what has changed?

A What I would say is that oh one and oh one preview, uh, we're really. Really breakthroughs, um, in, in the research community about having a model that can think, and the longer they thought for, the more, like, the higher likelihood they would be of being correct. Um, so that was really a breakthrough, but initially, and if you look at, like, old blog posts, you would mostly see, like, math, uh, math evals, and also, like, Maybe coding competitions, but things that are really easy to test whether you're correct or whether you're not. Um, and it also gives you like some suggestion about like how we were training some of these models. Um, and how I see maybe all of last year and especially the end of last year and the beginning of this year is that we were able to take these algorithms that work with, uh, verify rewards, like things where we can say you're correct or you're not, uh, to the messy real world. Um, and really optimize for the utility that we provide to users, and like making them more productive. Uh, so I think that's what really changed.

AI assessment note: “we were able to take these algorithms that work with, uh, verify rewards... to the messy real world”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q The 5.5 was particularly good, um, on genetic coding, computer use, knowledge work, and early scientific research. How does that work internally? Do different people focus on those different parts? How do you get to that result?

A Yeah, we definitely have different teams that are working on specific use cases and are pushing on these use cases. Uh, my team specifically is actually the one that is kind of taking all these vertical improvements and try to put them together in the final model. You could see it as a team that is doing both kind of the smoothing function. So you have all these improvements, but you need to make sure that the model doesn't feel too spiky, doesn't feel differently on different verticals. And also you need to have some teams that are working, and that's basically what my team is doing, on all the horizontal improvements. So there are many things like instruction following, function calling, or like thinking about how much should a model think for on different Uh, problems. Those are very horizontal and that kind of impacts all these use cases. So we have both these more vertical teams and these more horizontal ones. Um, and both are very important, uh, to, to, to improve the, to improve on the model. Um, and the good thing is that these things can kind of be improved orthogonally. So you might have like multiple different teams that are working on certain verticals and maybe for one model, there's only a Half of these teams that made integrations basically in the last run and like improve the model on these capabilities, and maybe for the next model, it'll be the other half. So …

AI assessment note: “we definitely have different teams that are working on specific use cases”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay, great. And, uh, what was your journey to OpenAI?

A Oh, it's a long story, but I'll try to keep it really short. Uh, basically I did my undergrad in biomedical engineering, um, in Switzerland. Um, I'm from Switzerland. And then I went on an exchange in Canada and I learned about what to VEC. So I don't know if you heard about this algorithm, but it basically takes words, which is like a, something discreet, uh, and puts it in a, in a vector space. Um, so puts it basically in a way to think about it as a plane where if words that are more civil to one another will be closer to one another. So it brings these, like, discrete words into, like, some continuous space that is semantically meaningful, and I was absolutely blown away by that algorithm, and that's when I decided that I wanted to work on natural language crossing and just, like, understanding language. Um, at that time, I was very wrong, but I thought that, uh, English, Uh, uh, NLP was basically solved. Well, like, close to being solved. That was in 2017. So that was, uh, uh, right when Transformers started. It was actually right before Transformers. So I was very wrong, but, uh, I decided that I wanted to work on under-researched languages, and basically, um, I wanted to improve, um, NLP on languages where we don't have that much data. Uh, so I went to, uh, work, uh, for Grab in Singapore. And I was basically building the natural language processing, uh, pipeline for the…

AI assessment note: “ended up at Stanford, did my PhD there... and then, uh, went to open it.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And still on the topic of, of reasoning, um, what, what's ultimately the difference between, uh, 5.5, uh, thinking versus 5.5 pro? Is that, is that just more test time compute, more tokens, and more time invested in solving a problem?

A Yes. Basically, it's just a question of, of, uh, how much test time compute we pour into the model, uh, or we pour into this entire, uh, system that we're shipping. Um, so, We, we've seen again and again, the longer the model think for, uh, the better answers we will get. The problem is that this, these curves that we're talking about, um, are not, are definitely not linear, and like they, there's some plateauing effect, and they kind of look, um, look logarithmic, uh, on some, in some sense, um, or depending on which evals. So You can pull, like, two times more compute and actually only get, like, small performance gains. Um, I personally don't use Pro that much because I really don't like waiting. I'm pretty impatient, so I don't like waiting for that long, and, uh, and I know that the probability of being correct definitely improves, but it doesn't improve, like, enough for, for me to use it. Um, but there are some people who use Pro and who really love it, especially actually for academic research. And, uh, I know especially a lot of mathematicians who are using it, and that's because they're kind of just have this in the background that is running for maybe one hour, uh, two hours, and they don't really need to like iterate really quickly with the model. Um, and pro is really good for that.

AI assessment note: “Yes. Basically, it's just a question of, of, uh, how much test time compute”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You had a great tweet the other day where, where you were talking about like the, Actual use cases, uh, that you see as a provider of sandbox of agents. Do you want to go through this? You had a, uh, code command execution, computer use, browser use, an RL environment infra. Do you want to unpack that?

A Sure. We basically, I've made it, I, I think I've structured it better since that tweet, which is right now we have two major use cases or two types of customers consuming Daytona and different use cases. And one is on the researcher side. So it'll be, You know, RL evals benchmarks, and the other will be on what we call background agents or long running agents. And so when you think of background agents or long running agents, the most popular, those are where a human is the end consumer. The human talks to a, let's call it app layer service that has an agent in that. And then the agent will call in the sandbox. So things will be, you know, think of like Harvey or perplexity or whatever as, or lovable. As these types of background long running agents. And so they can both be, uh, sort of headless. So code and command execution and or computer browser use. So depending on what they need to do. And the same thing is on the researcher side, where it's like RL and evals and whatnot, you can do RL and evals for, you know, coding. And for that, it's basically just headless commands and command execution in there, or you can actually teach it To do things in the real world. And then it does have to fire up a, you know, Windows, a Mac, a Linux sort of desktop or a browser to go through the thing. So code and command execution and browser computer use are like two ways an agent can work…

AI assessment note: “we have two major use cases or two types of customers consuming Daytona”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So that's a great, um, sort of sandbox one-on-one introduction to, to, to the concept. Um, help us understand, uh, where sandboxes fit in the overall picture of this Emerging, uh, agent stack, uh, that, uh, I think everybody's trying to figure out at the same time. So there's, there's different components, there's, uh, file systems, there's orchestration, there's like, what, what, what are the different pieces?

A I, I try to think about everything. To me, it's actually quite interesting where, and we'll get this a bit later as well, it's like, a lot of this all exists in real life today. And so people are like overthinking this. I'm not saying that there, there's not going to be new, Products and solutions and technology to solve it. But it's like, it's not a new fundamental way of work. And so when you think about the agent stack itself, it's like, okay, you first have the models and the models are essentially the brain that that is what it is of equivalent to what a human's brain is sort of. So you tell it something, it replies and understands and whatnot. And then under that is like, oh, what are the tools it can do to get Things done. Right. And so that can be anything from like any of the MCP or tool calls can do. It could be the sandbox, the computer, like whatever we as humans, we also have a bunch of tools that we use everything from a hammer to a computer and everything around that. Right. Um, so all of those things exist as well. Then there is memory that exists. Can, can an agent remember these things similar to like, like, do you remember these things that are there? People that talk about orchestration of agents. It is like management. You manage people. I manage people. Managing agents is similar, is not dissimilar to managing humans. Now, what tools will you use to manage…

AI assessment note: “when you think about the agent stack itself, it's like, okay, you first have”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Why does it matter? Uh, how quickly you can initialize a new sandbox and how many it depends on the user and the use case.

A So if you're like a long running background agent, again, Everyone prefers it to be fast. Like no one wants to wait. Like you don't want to wait for a reply. Like everyone wants to be fast to be very clear. So the faster, the better, but generally there's an actual reason why you want it very, very, very fast. And that is, especially if you, so for, I was gonna say for a background agent, a background agent might work for like 10 minutes or an hour or whatever. So the incremental millisecond might not matter, but I still believe that from a user perspective, even a second of waiting is Kind of uncomfortable. You don't want that. And so it's a 60 or 90, maybe less so, but there you want that one, two seconds for sure under that. But the interested, the more interesting part where that is really, really important is for the researchers where when you're doing reinforcement learning, you basically have allotment of GPUs and you want the GPU is more expensive than the CPU. So the vast majority of sandboxes are CPU boxes to be clear. So they're the computers that we all work on. There might be a graphics card in there, but Basically, it's the compute, the RAM, the CPU, um, and the hard disk that's in there, and they are cheaper, or less expensive, and easier to, to get, at least for now, we'll see for how long that lasts, than GPUs, which means you want your GPUs always to be at max…

AI assessment note: “you want your GPUs always to be at maximum utilization”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You had a great tweet the other day where, where you were talking about like the, Actual use cases, uh, that you see as a provider of sandbox of agents. Do you want to go through this? You had a, uh, code command execution, computer use, browser use, an RL environment infra. Do you want to unpack that?

A Sure. We basically, I've made it, I, I think I've structured it better since that tweet, which is right now we have two major use cases or two types of customers consuming Daytona and different use cases. And one is on the researcher side. So it'll be, You know, RL evals benchmarks, and the other will be on what we call background agents or long running agents. And so when you think of background agents or long running agents, the most popular, those are where a human is the end consumer. The human talks to a, let's call it app layer service that has an agent in that. And then the agent will call in the sandbox. So things will be, you know, think of like Harvey or perplexity or whatever as, or lovable. As these types of background long running agents. And so they can both be, uh, sort of headless. So code and command execution and or computer browser use. So depending on what they need to do. And the same thing is on the researcher side, where it's like RL and evals and whatnot, you can do RL and evals for, you know, coding. And for that, it's basically just headless commands and command execution in there, or you can actually teach it To do things in the real world. And then it does have to fire up a, you know, Windows, a Mac, a Linux sort of desktop or a browser to go through the thing. So code and command execution and browser computer use are like two ways an agent can work…

AI assessment note: “I've structured it better since that tweet, which is right now we have two major use cases”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So the sandboxes are fundamentally new. Primitive. Is that, is that correct? Is it like a history of sandbox or did sandbox exist before?

A So I, I would argue, I would say that the, the person that kind of nudged, although he said he did and someone else did, but anyway, there was a company called Code Sandbox way back in the day when we used to compete with our company Code Anywhere, which also was in the, in the realm of like Replet, it's all these cloud-based IDs. And so they called it Code Sandbox because I, I believe it sounded cute. It's like a box where your code, you know, for ID, um, lived. And they were actually one of the firsts. So the team there that actually used micro VMs, did snapshotting, forking, all these things that we do use today. And so that is sort of, and this is maybe a decade ago, but the utilization or the usage from, or the value that it was giving human developers was not there. There's this whole article about like the end of local hosts and people have been talking about this for I've been talking about this for 20 years. Um, and basically developers would say, like, you'll take my local hosts, like, out of my cold dead body. But now that agents are here, local hosts no longer actually, one, you don't want that for a number of reasons. And so we're finally getting to that. So that technology and that thesis around sandboxes originally now seems to be coming to fruition.

AI assessment note: “they were actually one of the firsts... and this is maybe a decade ago”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So I'm curious from your, your vantage points, the, the, the whole acceleration is to versus doomerism, uh, debate that has been raging for the last couple of years that seem to, you know, come and go depending on the moment. Uh, is that, is that at all helpful? Is that how you, you think about it?

A I, I, I, I dislike those labels. A lot on both sides. I think they're oddly enough used as largely, uh, pejoratively by both sides, right? People will dismiss someone as a doomer if they express too much concern about risks of AI systems, or if someone's trying to release models, they'll be called an accelerationist. It's, it's all, I mean, people, some people then, you know, use the terms of pride, I guess, but they're sort of inherently kind of dismissive terms, I think. Um, I believe I am on, I, I, I, I, I have never expressed a P. Doom and things like this. I just think it's a very weird concept as if the world is some stochastic set of dice that you can roll multiple times and that we don't have direct influence over this. Um, so I, I think that, um, I think that the reality is, and the, the, these sort of labels tend to, um, it tends to sort of dismiss a lot of the, the, the reality of the situation right now, which is that AI is not a technology that is, that is wholly bad, in my view, and it's not a technology that has no risks either, that just, we can just, you know, develop however, with no constraints whatsoever. Um, and I would say that, I think. 95% of all researchers, maybe 99% of all researchers feel probably a very similar way that, you know, this technology has great promise. There are massive opportunities, but we have to be mindful of the risks. It's sort of…

AI assessment note: “I dislike those labels. A lot on both sides.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Mm-hmm. And speaking of the Chinese, is, is, is safety a global movement? Like the, the way you have some level of, uh, uh, cooperation in conferences.

A Yeah. There are certainly efforts in many different countries. Um, uh, I'm less familiar with the Chinese efforts, but there are efforts in China certainly, but there's lots of safety in AI safety institutes or AI security institutes in many different countries. So. The UK obviously was the first AI safety, now AI security Institute. Uh, but Singapore has one as well. The US has the Casey, which, which, uh, uh, does similar function. Um, And many other countries have sort of burgeoning institutes as well. There's definitely global understanding of this problem. Now I do think that, um, these things are subject to some degree of political headwind and the fact that the, you know, AI safety, uh, conference was, or AI safety summit was renamed the AI action summit or something is, has some significance actually in terms of the sort of taking temperature of, of where the, Where the world is politically. But, at the same time, I also think a lot of the work being done Is, is a very similar nature. The, the actual researchers and what they're doing, um, they've people, these organizations have continued to do great work, continue to push the frontier and understanding how to assess, how to evaluate systems, how to safeguard them, all these things, they are happening in an ongoing fashion. And I think, um, you know, the good, good work is being done by researchers at companies in acad…

AI assessment note: “Yeah. There are certainly efforts in many different countries.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What happened then? Like how did the labs react?

A When the models were constrained to just be the models themselves, this is not that easy to patch. I mean, you can patch single strings. A lot of labs sort of blocked individual strings, um, that we had published just Which is fine, right? But if you ran the whole process again, you could find another string that would actually actually circumvent it. It wasn't until the development, A, of additional safety classifiers that people started to really kind of be able to detect and stop these things. But then also reasoning models. Reasoning models were much more effective because you can't really do the same trick of optimizing for a probability with a reasoning model that has a whole trace of reasoning that happens in the middle and kind of reflect a bit more. So it's much harder to break reasoning models in the same way. But yeah, the, the, the short is that there, there was certainly some work done to address these things, but it took additional layers of security and, and, and security and the advent of reasoning models before they really became ineffective.

AI assessment note: “A lot of labs sort of blocked individual strings”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q it's been an absolutely epic time at Anthropic and very hard to start this conversation with anything but, uh, the announcement which came out, uh, just yesterday as we were recording this of, uh, Project Glasswing and then, uh, Claude Mythos preview, which you tweeted about and you said it's Pretty hard to overstate what a step function change this model has been inside Anthropic. Can you elaborate on that?

A Yeah, sure. Mythos is a unreleased frontier model. It's, it's a general purpose model that was trained not specifically for cybersecurity or specifically for coding or specifically for software, but, uh, we have discovered what we believe to be outsized capabilities specifically in the aspect of cybersecurity, and we believe that it has Far reaching implications for the safety of software and infrastructure. I think there's two things I'm alluding to, uh, in my tweet. We've obviously used the model internally for a while now. As a software engineer, I think many of us have gone through this exercise of the last couple of years of like our first initial contact with AI was like, you know, probably not that impressive. The first time I touched AI was like sometime in 2013. This was before we had large language models. I was at Microsoft at the time. We had something called Project Oxford where we had an Ngram model. You would give us a token. You would say something like world and the model would return world wide web. And that was sort of the, I want to say the frontier of what language models were capable of doing. And I think a lot of us in the public over the last couple of years had these moments of being like, oh, this model is more capable. I can do more things than I may be expected. Mythos preview is a model that for us as engineers internally, Feels like a dramatic step…

AI assessment note: “Mythos preview is a model that for us as engineers internally, Feels like a dramatic step up”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q famously, uh, was, was, uh, coded in 10 days or so. At least that's, that's, that's, that's, that's the lore of it. Actually, let's, let's spend a minute on this if, uh, if the industry lore is, is not entirely, um, uh, correct. I guess what, what happened? Tell us that story of the, of the 10 days and the core work beyond entirely, uh, built by a cloud code.

A Yeah, I can kind of see, I can kind of see why that caught on. In software, nothing is ever built from scratch, right? And I think the, the exact quote that I gave that people used was that my team sprinted on this for, I think, the last 10 days or so, which is accurate. That is, that is the case. My team got together 10 days before release, and I was like, alright, we should probably release something. What did we release? What does it look like? What is it named? What can it do? However, however, as anyone who's ever built any software can attest to, it's not like you start from scratch with, like, ones and zeros, right? You, like, make use of a lot of libraries. You make use of, like, The research you've done in the past, in particular in Anthropic, the, the core problem that I tried to solve for, which is how do you make it easier to bring the power of cloud code to non-coding work, right? Like general knowledge work. A lot of very smart people have thought about that at length. And, um, it would be inaccurate to say that Anthropic has not thought about this problem. And it would also be inaccurate to say that I feel like slowly came into this cold without benefiting from all that work.

AI assessment note: “my team sprinted on this for, I think, the last 10 days or so, which is accurate”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And where does memory live for co-work to remember you and remember your task? Is that the model? Is that in the harness?

A It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files. It's really just the, the model being instructed, hey, if you feel like anything was pertinent that you might want to remember in the future, just write it down. And then we help the model a little bit with, like, organizing its memory so you can, you can set up projects that have isolated memory versus, like, your overall memory. But the, the underlying technology that sort of is bolted on on top of the model is sometimes surprising to people that it's not, you know, like a, some complex, fancy database technology.

AI assessment note: “It's in the harness, actually”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Would you say that, uh, UX is as important to the success of an AI agent as the technology itself? Like how you take users on a journey, uh, so that they are empowered? And, uh, if so, what, what are some other lessons learned building AI agents from a UX standpoint?

A It's a really good question because I actually think, I actually think that's true. I do think the UX matters quite a bit. Right? Like, even if you go back to our, one of our most popular products, Cloud Code, um, the very genesis of it was, what if Cloud, but instead of, like, in the cloud, it's running on your computer in your terminal. That, that is almost entirely UX. It's the same model. It's the same core capabilities. Um, it's really all around what is the user experience and, like, how do you interact with the model, right? But it's fundamentally the same model, and that's really where a lot of the, a lot of the benefits came from. And I think similarly today, um, the AI products I see resonate with people the most are rarely the ones that deliver the most raw potential, the most raw power. Um, and I, I would actually go one step further and say this is probably true, not just with AI, but maybe with software overall, right? Like, um, I'm going to blindly assume that plenty of startups out there offer email with more features than Gmail. There's plenty of companies that try to, like, sort of jump ahead by offering a larger amount of features or more buttons or, like, more capabilities. Um, I often think a lot about the, um, The silly times of mobile phones right before the smartphone was invented, right? All the things that people bolted onto phones. We had like phones …

AI assessment note: “I actually think that's true. I do think the UX matters quite a bit.”

← previous page 11 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.