Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q How do you, I'm sure this has changed over time, but given platform and broad application surface area, like how do you organize at Rippling given you don't have like contemporary, many contemporary companies with a similar strategy? Um, and you can't be like, I'm going to do what Microsoft does at this scale.
A I mean, I think we try and, um, I mean, we have, we have teams that build capability. You know, that are kind of the, the platform teams. And then we have application teams that are building, you know, sort of hopefully out of the underlying platform capabilities applications. And I think it's like, it's really important to have like the right leader for these application teams. Like former founders are great. Um, but also you need people that can like scale over time, um, as, as the product grows. Um, and so I, you know, ideally, In an ideal world, you want someone with just like real ownership of those areas who can drive it because otherwise you get this, um, you get this sort of bottleneck at the level of executive attention where, you know, if I have to drive it, it's just like, I can't do that. I can only do that for so many things. And then inevitably a lot of balls get dropped. And so you need people like in seat who can like really own it and run like the business holistically, who can think about the product roadmap, the marketing, Um, you know, the sales elements, you know, the competition, you know, like the whole thing and kind of, you know, and synthesize it down to like, okay, here's what we're going to do. Um, and that's, that's always sort of the most challenging thing is to find those people. I mean, we, we try and hire a lot of founders at Ripley to make it w…
AI assessment note: “we have teams that build capability... platform teams. And then we have application teams”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q about the, um, technical approach here, but Jack, you and I met in the context of, you know, you being a beloved engineering and product leader at Stripe, um, coming from the engineering side and looking for like the most interesting problems to work on in AI. Why did you decide to work on this versus like some of the other things we were talking about, like Cogen and such?
A Yes, as you know, Sarah, I spent quite some time thinking about my next steps and what I wanted to do with my, my life after the period I was at, at Stripe. And I give a lot of credit to Josh actually for this, that, you know, we were good friends going back even to the, to college. You know, we were PSET buddies in, in college, uh, at Harvard in, in many of the same classes together. While I was maxing out the CS curriculum, Josh was also doing that somehow for the chemistry and physics and all the other correct scientific curricula as well. Uh, but we, we had landed in a lot of the same classes and, uh, as, uh, as we went our separate ways after college, we really just made a point of keeping in touch every, you know, three, six months. And Josh would always talk to me about, about his research. Once it became clear that, that the research that, that, uh, Josh and us were doing in this space was really no longer just a toy, uh, and, but was, was really going to impact and change the entire industry. That idea became infectious. Right. It sort of become impossible to unsee the future once you have that glimpse. And although you didn't know until very recently that any of this is, it was going to work. And of course there's, there's still a lot left to prove. Once you start to grasp the implications of the fact that over the next few years, we are going to have the ability as a…
AI assessment note: “Once you start to grasp the implications... It's almost hard to work on anything else”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q maybe Jack, I will start with your amazing engineer. And then you guys also have like a, a very software oriented team working on biological problems. Some of those people come from, you know, long-term research in that space in particular, but for, for yourself, Jack, like, as you said, you're, Your software person, how do you get up to speed on the bio area to go do leading work?
A Well, I think it's two things. First of all, uh, ramping up on any new field is always just a total fight. Uh, you have to, to get to the frontier and to be, have read the right papers and to be knowledgeable about the, the areas that you need, you need to learn. You just have to sort of push your head down and push through. And there are, there are waves of, uh, excitement and misery in that experience, but You can get there fast if you really set your mind to it. And I'd say the second part is that surrounding yourself with just the most incredible team is the best thing that you can do far beyond anything that you can learn by yourself. And we have certainly the most special group of people that I've ever worked with within the company. Our co-founders, Matt McPartland and Jack, uh, who, who are, you know, just rare talents, uh, and Then, you know, the entire team beyond that, some of the former heads of AI at other drug discovery companies, some of the top open source contributors, the team is so, uh, multi-talented. It's, it's small. That's, it's around a dozen people, but, but mighty. And I think as we've seen in other areas of AI, small but mighty teams can go a really, really long, long way these days. And so, you know, the, the, uh, I think there are actually surprisingly few people on our, on our team, even with a computer science degree, Josh himself got a chemistry …
AI assessment note: “read the right papers and... surrounding yourself with just the most incredible team”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I want to come back to impact because I think the ramifications here are, are, um, really huge. But, uh, if we just go and, like, think about first problem design, you, I think, looked at 52 problems. Why that many? And, like, how do you specify a target? I'm picturing, like, bind to epitope X, but I'm sure there are other requirements you'd want to have as drug designers.
A It's a great question, Sarah. So in the CHI-II paper, we look at over 50 targets. Most of the existing, uh, papers in this area of doing AI for, for drug discovery are usually looking at like one, two or three targets. But again, it was important for us if we were seeing this as an engineering problem to make sure that this is going to be generalizable. It's like, imagine you had a new LLM paper and you said, oh, I solved like one problem in the USIMO, uh, contest, like really, really cool. It's like, yeah, you need a real benchmark and you need to actually have that benchmark at scale. You need to have enough Problems to convince yourself that the system is working. So that's why whenever we do these experiments, you know, sometimes we'll, we'll try one or two targets just to make sure there's not like a huge bug and, you know, make sure not everything fails. But, you know, even if everything fails in one or two, you know, the hit rate's 50%, you could have just gotten unlucky. So that's one of the reasons why we decided to do a big benchmark here, really convince ourselves things are working. The way we selected the 50 problems, the biology people would laugh at this and engineering people would love it. We actually just went to the vendor catalogs to see what was in stock. Cause we wanted to turn around this experiment quickly. We ordered all of these designs at the same tim…
AI assessment note: “it was important for us if we were seeing this as an engineering problem”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q We have a broad audience for no priors that ranges from like Business people to engineers, machine learning researchers, some scientists in other fields. Like, what intuition can you give listeners for how the model works under the hood? Like, especially for anybody who might start with, um, some familiarity with, like, structure prediction models.
A Yeah, well, structure prediction is really a key part in making these models work, and it's actually the first thing we did when we started the company is we sprinted to build a state-of-the-art structure prediction engine. We actually open-sourced the first version of that. It's called Chi-One, and again, like, Scientists around the world are using that now, but structure prediction basically gives you an atomic level microscope, and it allows you to see where atoms are placed in three D space. So once you can do that and you have this microscope, then the next question is, well, can we start moving those atoms around, right? We can now start to make changes in a sequence, and then we can see the ramifications of those changes in three D space. So the actual design model, you can think of it as a, you prompt it with some information, like here's a target. Uh, that we want to go and, and, and design, uh, an antibody against. And then the model will, will try to place again, these items in three D space in order to satisfy that constraint. Like we tell the model, here's the target and I want you to make a molecule that, you know, binds to that location. And then model will go in and generate, uh, both a, a sequence and a structure that, uh, that kind of fits into that. So that's like the high level intuition for this.
AI assessment note: “the actual design model, you can think of it as a, you prompt it”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What has the reception been like so far? What is the biggest objection? Because this is a, you know, significant challenge to the ideas of high throughput screening or even like the workflow that, um, you know, even innovative pharma and biotechs have today.
A Yeah, it's a great question. You know, usually when these kinds of, uh, papers come out, again, people have, Try to, to do this many times. The critique is, is often, you know, does this really work? You know, you, you show this on maybe COVID, for example, is this going to work for a case where we have less training data? Are the molecules going to be high quality? Do we really, you know, kind of believe the data? So I think the approach we did, like benchmarking this at scale has really helped a lot with that reception. Like I think people really appreciated that approach, which has been great. Some of the questions people have is, okay, like I can already discover drugs. So Uh, you know, so now I have AI that can do it a lot faster, but does that actually change the kinds of molecules I can work on? And it goes back to what we just discussed before. I think there are other folks that are responding to that saying like, no, like the, the transformation here is how about those projects that didn't work for you, uh, or, or where you're really struggling today. Now you've got another tool in the toolkit and you kind of have to use this tool now, or, or you might be left behind. So I think that it's been really interesting to see the, the community kind of digesting this. Of course, a lot of the AI folks are, are really excited, right? Like we're getting artificial antibodies, uh…
AI assessment note: “The critique is, is often, you know, does this really work?”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q economically valuable and interesting, like matrix multiplication, the potential efficiency of it, what is your intuition for, uh, why, you know, those, those solutions have not been found before? Is it simply like the search space is too large or people, In this field were complacent in that, um, they believed a certain solution was like the maximum efficiency or because, uh, clearly there's, there's value to be had here.
A My sort of opinion on this is basically that if you look at the structure of the, of the algorithm, what Strassen produced was quite, uh, uh, sort of, uh, ingenious. It was not a natural thing that you would sort of think of. And that is, uh, that was for only two by two matrices. As you sort of go to larger sizes, the space is so huge. The constructions are not sort of something which is very natural. These are very involved and intricate, intricate, intricate sort of constructions that would be very hard to discover by chance. So it's quite interesting that it is, it has this very special structure, but it's not something that is, that comes naturally. Or, uh, to, to a human computer scientist.
AI assessment note: “As you sort of go to larger sizes, the space is so huge.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Does that imply you have a particular, um, point of view, um, from a, like, a neuroscience perspective of, like, you know, how fundamental visual, you've, I mean, you've always been a leader in, um, uh, computer vision, right? But in how important visual intelligence is versus, let's say, like, Large language models and textual intelligence.
A I actually do. I think from a neural and cognitive science point of view that spatial intelligence is a really hard problem that evolution has to solve for animals. And what's really interesting is I think animals have solved it to an extent, but not fully solved it. It's one of the hardest problem because, um, what is the problem animal has to solve? Animals Have to evolve the capability of collecting lights in something which we call eyes mostly. And then with that collection of eyes, it has to reconstruct a three D world in their mind somehow so that they can navigate and they can do things. And of course they can interact for humans where the most Capable animal in terms of manipulation. We can do a lot of things, and all this is spatial intelligence. To me, that's, um, that's just rooted in, in our intelligence. What is interesting is, it's not a fully solved problem, even in animals. We, uh, for example, uh, for humans, right, um, if I ask you to close your eyes right now and draw out or Or, or build a three D model of the environment around you. It's not that easy. We don't have that much capability to generate extremely complicated three D model till we get trained. You know, there are some of us, whether they're architects or, or designers, or just people with a lot of training and a lot of talent. And that's, that's, uh, That's a hard thing to do. And imagine you do i…
AI assessment note: “I actually do. I think from a neural and cognitive science point of view”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q and managing builds and scale of manufacturing, you're going to have many fewer form factors. And the other argument is the economic value of specialization is very high. And therefore there'll be, you know, thousands and thousands of different form factors as we move to sort of a robot driven future. Do you have a point of view on sort of where we're likely to land between those two viewpoints?
A I think we're going to gradient descending to optimization of productivity and efficiency. My hypothesis is that the requirements of different tasks are so vast that having very few form or, or sticking with one form is. Energy, energy inefficient. And a lot of tasks can be done and should be done by much more energy efficient form factors. Just an extreme and, and trivial example. If we put robots underwater, they should not be in the shape of humans. They better be in the shape of fish, right? Just think about energy efficiency. And the same with flying. I don't think human form is, uh, our airplanes are becoming more and more robots. Um, so I, I, I do think there's gonna be diversity.
AI assessment note: “so I, I, I do think there's gonna be diversity.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What are the biggest challenges in, um, I guess, trying to go down this path of, you know, designing and training world models? I imagine one is, like, you worked on images, you worked on video, but we, we have images, and we have video, and we don't have lots of, you know, three-D worlds, like, in, in a format I assume you're building.
A Yeah, data is absolutely a challenge. You're totally right about that. Um, you know, to create, uh, world models, three-D foundation models, uh, we, we require More and more sophisticated data engineering, data acquisition, data processing, and data, uh, synthesis. So, um, uh, I am envious of my, uh, NLP LLM colleagues that the data is so abundant on the internet, and we don't necessarily have that luxury. So that's definitely one, um, one challenge. Another one challenge is that, um, three D is, this is kind of, um, ironic, right? Every one of us use three D every day, like in so many settings. Basically you open your eye and, and, and the whole life that you experience is three D. Okay. Even when we type on the computer or stare at a screen all the time, yet it's still not as easy a form factor to deliver in the hands of people compared to language. The language is just so easy. And, uh, it's also a very active form of, it's not a passive consumption of viewing. Nobody wakes up and say, I'm just gonna sit here and watch three D, you know. So, um, that, uh, creates challenges for, for, for productization and how to do it in the right way.
AI assessment note: “data is absolutely a challenge. You're totally right about that.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Why do you think it didn't work? Because it, it felt like an awful market.
A It was like a graveyard, like, you know, of all these companies that tried to solve the problem and it didn't. Part of it was just that I think search is a hard problem. In an enterprise, like even getting access to all the data that you want to search, It was such a big problem. In the pre-SaaS world, the, there was no way to sort of go into those data centers, figure out where the servers were, where the storage systems were, try to connect with information in them. It was a big, it was a big challenge. So SaaS actually solved that issue. So like search products, like most of them, most of the companies started in the pre-SaaS world, they failed, uh, because you could just couldn't build a turnkey product. But SaaS actually allowed you to, to actually build something, you know, uh, which is my insight. Was that like, look, you know, the enterprise world has changed. We have these SaaS systems now, and SaaS systems don't have versions. Like everybody, all customers have the same version. You know, they are open, they're interoperable. You can actually hit them with APIs and get all the content. I felt that the biggest problem was actually solved, which was that I could actually easily go and bring all the enterprise information and data in one place, uh, and build this unified search. System on top. So that was actually a big unlock.
AI assessment note: “In the pre-SaaS world, the, there was no way to sort of go into those data centers”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q like, okay, learning is hard and, and like, there is no way around it. Like a smart friend of mine compared learning to like going to the gym, but for your mind, like, It's very simple. You just have to do it. And there's no other way besides like being willing to put the effort and think really hard about it. How do you react to that point of view?
A I would be very, very happy if everybody who is working on an education app thought that way, because then we will have no competitors. It just turns out that the hardest thing about learning is motivation. By the way, I believe that to be true of the gym too. You know, it's like a lot of people talk about, like, you can do, you can spend all your time debating whether an elliptical machine is better than a treadmill. But in the end, what matters is that you do it. Uh, that, that's like, 90% of what matters. Uh, it's the same with learning something. You know, you can spend all your time debating whether you should read that book or that, or this book, or whether you should do it through an app, or you should do it to a tutor. There, there's, of course, differences in how effective each of these are, but what matters the most is that you actually do it, and if you want to get people to actually do it, you have to make it as easy as possible to get started, to get in there, and to be motivated. I believe that is the reason why we have grown so much. I mean, and in fact, it's funny. I mean, we see a lot of, um, language learning apps that pop up that say, oh, we're like Duolingo, but without the gamification. And every time I see that, I'm like, great. You do that. Carry on. 99% of the world's population just is not that motivated for any activity. And, um, You know, there's the …
AI assessment note: “It just turns out that the hardest thing about learning is motivation.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What else are you excited about in terms of, like, AI changing the product itself? Like role play and explain my answer and of these things. Like, what do you think matters or works so far?
A You know, I'm excited about, uh, generally practicing real world language, like conversation. I'm excited about that. Outside of language, I'm really excited about teaching math. We can do a much better job at teaching math with large language models, and we have a math course, um, but it is, it's, you know, we're completely retooling it, um, to be a lot more like a tutor. Now, you know, again, I'll go back to tutors. Tutors can be really good at kind of learning outcomes. They have this one big problem that they're really boring. So we're trying to find ways to make it a quote unquote tutor, but that is also gamified. Just turns out most people would rather play Candy Crush than sit there in front of a tutor. So we're trying to come up with something that is as effective as a tutor, but as fun as Candy Crush. The reality is we'll probably come up with something that is 90% as effective as a tutor and 90% as fun as Candy Crush, but at least the combo of this will be a lot more effective.
AI assessment note: “Outside of language, I'm really excited about teaching math.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What do you know about how people learn that other, other humans don't, that you think is interesting?
A A lot of it is just codified on what the algorithms have adapted to. I mean, we have adaptive algorithms that try to figure out how to teach. And I think it's hard to verbalize a lot of it. I mean, we know a few things that are, that are, that you can verbalize, um, that, you know, may not be that surprising, but the farther your language is from your native language, the harder it is to learn. We have a model that can predict whether you're going to get an exercise right or wrong, and we're actually extremely accurate. So when we give you an exercise, we know if you're going to get it right or wrong. We're very accurate. Um, and the way, the way we do that is we just watch everything you're doing and we see what you're good at and what you're bad at. And for example, we know, we know for each user, this particular user is bad at the past tense. Um, one of the things you may think at first that the right thing to do is because this user is bad at the past tense, we should give them more past tense. And that's roughly true, but it's not exactly true because if all we did was give you lessons that Of things that you're bad at. These would be very, very, very horrible lessons for you. So there's a whole, uh, you know, I was going to say art, but it's actually more of a science. There's a more of a science about when to give you the things that you're bad at. And, and for example, …
AI assessment note: “the right thing to do is to give you an exercise that you're about 83%”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q like, diversity and generalization across the different things you can be doing, but it, it's a, it seems like a very large space and scaling RL, like, new gen RL is not, it's just not obvious, like, how, to me, it's not obvious how you do it or how you choose the path. Is there some sort of organizing framework that, You know, you guys have that you can share?
A I mean, I, I don't know if there's, like, one organizing framework. I think there are a few, like, factors at least that I think about in, like, the very, very grand scheme of things is, like, how much, like, in order to solve this task, like, how much uncertainty with the environment do I have to, like, wrestle with? Like, um, for some things where it's, like, this is a purely fat, like, who was the first president of the United States? Like, there's zero, like, Environment I need to interact with to, like, reach the answer to this question correctly. I just need to remember the answer and say the answer. You know, if I want you to, like, write some code, you know, that, like, solve some problem. Well, now I have to deal with a little bit of, like, not purely internal model stuff, but also, like, okay, I need to execute the code and, like, that code execution environment is maybe more complicated than my model can memorize internally. So I have to do like a little bit of like writing code and then executing it and making sure it does what I thought it did and then testing it and then giving it to the user and things get like the amount of that sort of stuff outside the model that you have to like, you know, you can't just recall the answer and give it to the user. You have to like test something and, you know, run an experiment in the world and then wait for the results of tha…
AI assessment note: “there are a few, like, factors at least that I think about”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q It is clear that reinforcement fine tuning on very powerful base models can do very useful things now. That's super exciting. What advice would you have for, um, startups or other companies who are thinking about doing RFT for a particular task as to, like, When it's worth doing or when they can just try to do, um, sort of just traditional orchestration where agents are a component?
A So I think in general you will always get a model better, better at a specific task if you train on that task, but we also see a lot of generalization from training on one kind of task to, to, you know, other domains. So you can train a reasoning model on mostly math, coding, other reasoning kind of problems, and it will be good at writing, but if you You know, trained on that specific task, it would be better at it. I think if you have a very specific task that you think is so different to anything that the model was likely trained on and you try it a bunch of times yourself and you've tried a lot of different prompts and it's just really not good at it. So maybe it's, um, genetic sequencing task or something that's just so out of distribution for the model that it doesn't know how to figure it out. I think that is a good time to try reinforcement fine tuning or If you have a task that is so critical to your, like, business workflow that getting the extra 10, 15% performance is really make or break, then probably try it. But if it's something that you think, oh, the model's pretty good at, but it gets things wrong, you know, some percentage of the time, and then you see with every next model that's released, it gets a little bit better, it might not be worth the effort if the model naturally is just going to get better at those things. So that would be my recommendation.
AI assessment note: “I think that is a good time to try reinforcement fine tuning”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Maybe we can actually use this as a, uh, like a moment to talk about some of the failure modes. Like, how do you think about some of the, uh, classic issues with agents, like maybe, you know, compounding error or distraction or even safety?
A Yeah, so I think with deep research, since it can't actually take actions that aren't kind of the same class of the typical agent safety problems you would think of, but I think the fact that the, Responses are much more comprehensive and take longer means that people will trust them more. So I think maybe hallucinations is a bigger problem. While this model is hallucinates less than any model that we've ever released, it is still possible for it to hallucinate most times because it will infer something incorrectly from one of its sources. So that's part of the reason we have citations, because it's very important that the user is able to check where the information came from. And if it's not, Correct. They can hopefully figure it out. Um, but yeah, that's definitely one of the biggest model limitations and something that we're actively always working on to improve. In terms of like future agents, I think the ideal agent will be able to do research and take actions for, um, on your behalf. And so I think that's a much harder question that we need to address. And it's kind of at that point when capabilities and safety kind of converge where an agent is not Useful if you can't trust it to do a task in a way that doesn't have unintended side effects that you don't want. Like if you ask it to do a task for you and then in the process it sends an embarrassing email or something like…
AI assessment note: “hallucinations is a bigger problem. While this model is hallucinates less than any model”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you think there are, like, big blockers to progress of, like you said, um, maybe not exactly describing it as the next iteration of deep research, but just confidence that, you know, we're gonna have these unified agent capabilities and it will feel like a co-worker. What stands between us and that?
A There's a lot of really hard safety questions that we need to figure out. You know, we would never ship anything That we don't have, like, very high confidence is safe, and I think it's, the stakes are way higher when it can, when it has access to your GitHub repositories and your passwords and your private data, so I think that's a really big challenge. I guess also, if you want the model to be able to do tasks that take many, many hours, finding efficient ways to manage context, kind of similar to the memory thing, but if you're doing a task for a really long time, you're going to run out of context, so what's an efficient way of, like, dealing with that? Um, allowing the model to continue to do its thing. And then, yeah, the, just the task of making, making the data and making the tools. I mean, I've, I've said this already a few times, but that, that's a lot of work.
AI assessment note: “There's a lot of really hard safety questions that we need to figure out.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Was there any tool or data set that was most crucial for that, uh, murky discovery you made in Zambia? Like, was there a piece of data that others had overlooked? Was it just Looking at that geography, was it a specific tool that?
A There is no one piece of data that enabled that. And that's really a critical theme. This is often new technologies are invented in this industry where people think, ah, this is going to be the silver bullet. It's going to help us find, it's going to help us find all the ore deposits, or this data set alone is going to let us do that. Actually, the data is very high dimensional. And when you can add dimensionality to the data, then you can have improved predictive power. Uh, and so that's the, That's the story there as it is everywhere else. It's a combination of, you know, new analytical methods, the ability to quantify uncertainty and understand the range of possibilities, and critical scientific insights about the way that these OR systems are formed, and all of those things in combination are what make it possible. There is no way to isolate the AI from the HI. There's no way to isolate one piece of data that's uniquely powerful, and that's one of the reasons that I think is limited innovation as well, is that we think, oh, You know, this new airborne gravity gradiometry invented in the 19 nineties was going to find all the ore deposits. It doesn't, but it's really powerful. We're really happy that when we can get that data, we go collect it ourselves, but these are, these are incremental improvements to predictive power, but it's only possible if you can work with all of t…
AI assessment note: “There is no one piece of data that enabled that.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q what people call evaluation crisis around, like, the models are so good, uh, and they're somewhat indistinguishable at the fringe of, of capability today that we, we don't know how to test them, you know, ignoring all the issues with, uh, uh, People, um, gaming, gaming the benchmarks, right? Um, what do you think, how, like, what right ideas are there about, uh, evaluating models, especially as they become superhuman?
A Well, I think one of the most important things is that a lot of the evals historically have been for, like, zero shot of a model, uh, or, like, a test question, right? That might be academic, when the thing that we actually need to eval is, like, what's economically valuable work, right? When a software engineer goes to their job, It's so much more than writing a PR. It's, like, coordinating with all of the relevant parties, uh, to, like, understand what does, like, the product manager want, and how does that fit into, you know, the priorities of each team, and how does that all translate to, like, the end output of work? And so I think we're going to see an immense amount of eval creation for, like, agents, uh, and that is the largest barrier to Automating most knowledge work in the economy.
AI assessment note: “the thing that we actually need to eval is, like, what's economically valuable work”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Given you spend all your time thinking about how to attract high-skilled candidates and determine their effectiveness, like, what advice would you have for people who are hiring in startups and scaling companies?
A Early on, it's hard to stress the importance of talent density, and just, like, there's always a trade-off between hiring speed and hiring quality, and you should just, for those early employees, like, always index on quality. Like, you need to be patient, and you need to make sure that people are extremely high caliber. When you're scaling up an org, uh, you obviously don't want to drop those standards, but people need to be a lot more data-driven around What are the characteristics of people that actually drive the outcomes they care about? And it feels like where a lot of problems happen is when that slips, when it's sort of like this vibes based assessment that doesn't scale very well, where each hiring manager is doing it in a fragmented way. And it's hard to enforce those standards across the board. And so just being very disciplined around like, what are your hiring goals? What are the characteristics of people that you know are actually going to achieve the business outcomes you care about? And how do you measure those things is really important.
AI assessment note: “for those early employees, like, always index on quality.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q That's really cool. And I guess you started physical intelligence with four other co-founders and an incredibly impressive team. Could you tell us a little bit more about what physical intelligence is working on and the approach that you're taking? Because I think it's a pretty unique, um, slant on the whole field and approach.
A Yeah. So we're trying to build a big neural network model that could ultimately control any robot to do anything in any scenario. And like a big part of our vision is that In the past, robotics has focused on, like, trying to go deep on one application and, like, developing a robot to do one thing, and then ultimately gotten kind of stuck in that one application. It's really hard to, like, solve one thing and then try to get out of that and broaden, and instead we're really, um, in it for the, the long term to try to address this broader problem of physical intelligence in the real world. We're thinking a lot about generalization, generalists, uh, and Unlike other robotics companies, we think that being able to leverage all of the possible data is very important. And this, uh, comes down to actually not just leveraging data from one robot, but from any robot platform that might have six joints or seven joints or two arms or one arm. We've seen a lot of evidence that you could actually transfer a lot of rich information across these different embodiments and allows you to use data. And also if you iterate on your robot platform, you don't have to throw all your data away. I have faced a lot of pain In the past where we got a new version of the robot and then your policy doesn't work. Uh, and it's, it's a really painful process to try to get back to where you were, um, on the pre…
AI assessment note: “we're trying to build a big neural network model that could ultimately control any robot”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q To the large language model world where, you know, really a mixture of deep learning, the transformer architecture in scale has really proven out that you can get real generalizability in different forms of transfer between different areas. Could you tell us a little bit more about the architecture you're taking or the approach or, you know, how you're thinking about the basis for the foundation model that you're developing?
A At the beginning, we were just getting off the ground. We were trying to scale data collection, and a big part of that is, unlike in language, we don't have, like, Wikipedia or an internet of robot motions, and we're really excited about scaling data on real robots in the real world. This is, this kind of real data is what has fueled machine learning advances in the past, and a big part of that is we actually need to collect that data, and that looks like teleoperating robots In the physical world. We're also exploring other ways of scaling data as well, but the kind of bread and butter is scaling real robot data. Uh, we released something in late October where we showed, um, some of our initial efforts around scaling data and how we can learn very complex tasks of folding laundry, cleaning tables, built, constructing a cardboard box. Now where we are in that, uh, journey is really thinking a lot about Language interaction, uh, and generalization to different environments. So what we showed in October was the robot in one environment, and it was trained, it had data in that environment. We did, we were able to see some amount of generalization, so it was able to fold shirts that had never seen before, fold shorts that has never seen before, but, um, the degree of generalization was very limited, and you also couldn't interact with it in any way. You couldn't prompt it and tell …
AI assessment note: “in terms of the architecture, we're, we're using transformers, uh, and we are using pre-trained models”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q like Tesla and others under the assumption that the world is designed for people and therefore is the perfect form factor to coexist with people. And then other people have taken very different approaches in terms of saying, well, I need something that's more specialized for the home in certain ways or for factories or manufacturing or you name it. What is your view on kind of humanoid versus not?
A On one hand, I think humanoids are really cool, and I have one in my lab at Stanford. On the other hand, I think that they're a little overrated, and one way it kind of to practically look at it is, I think that we're generally fairly bottlenecked on data right now, and some people argue that with humanoids, you can maybe collect data more easily because it matches the human form factor, and so maybe it'd be easier to mimic humans, and I've actually heard people make those arguments, but if you've ever actually tried to Teleoperate a humanoid. Uh, it's actually a lot harder to teleoperate than that a static manipulator or, or a mobile manipulator with wheels. Optimizing for being able to collect data, I think is very important, uh, because the, if we can get to the point where we have more data than we could ever want, then it just comes down to research and compute and, and evaluations. And so we're optimizing for That's one of the things we're kind of optimizing for, and so we're using cheap robots. We're using robots that we can very easily develop teleoperation interfaces for, in which you can do teleoperation very quickly, uh, and collect diverse data, collect lots of data.
AI assessment note: “On the other hand, I think that they're a little overrated”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q startup, so I think you have a full ability to suggest stuff to people in that area for sure. I've heard a number of different groups doing is really using observational data of people Is part of the training set. So that could be YouTube videos. It could be things that they're recording specifically for the purpose. How do you think about that in the context of training robotic models?
A I think that data can have a lot of value, but I think that by itself, it won't get you very far, uh, and I think that there's actually some really nice analogies you can make where, um, for example, if you watch, like, an Olympic swimmer, swimmer race, uh, even if you had their strength, uh, just their practice at moving their own muscles to do the, to, to accomplish what they're accomplishing is, like, essential for being able to do it, or if you're trying to learn how to hit a tennis ball well, You won't be able to learn it by kind of watching the pros. Now, maybe these examples seem a little bit contrived because they're talking about like experts. The reason why I make those analogies is that we humans are experts at motor control, low level motor control already for a variety of things that our robots are not. And I think the robots actually need experience from their own body, uh, in order to learn. Uh, and so I think that it's really promising to be able to leverage that form of data, especially to expand on, um, the robot's own experience, but it's really going to be essential to like actually Have the data from the robot itself, too.
AI assessment note: “I think that data can have a lot of value, but I think that by itself”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Um, there has been a lot of news recently on a, uh, a different question, which is the rise of, uh, Chinese biotechs. Um, uh, you know, for the core members of the research community here, is, is that a threat? How do you think of it?
A Well, their cost basis is definitely more competitive, right? I think, uh, a lot of the discussion, you know, around the water cooler and the biotech and farm industry is, you know, how are they able to do it at this pace? How are they able to do it at this cost? Why do their data packages look so good, right? They have safety, they have talks, they have all these IND enabling studies, you know, it's really competitive, and, uh, I think folks got really surprised at the Um, efficiency of the pipelining and the ability to manufacture all these different antibodies, uh, primarily. Um, and I think that's great for the industry, right? I think everybody, including patients, investors, you know, the biotech companies themselves want lower cost basis, right? We want, um, the ability to actually make molecules that work faster, and I think all these things will You know, kind of compete right in the system to be able to reduce the right now, like pretty high cost basis of doing these things, uh, you know, stateside, right? I think one of the, one of the core challenges right now is we have a wide array of services and, you know, CROs and contract research collaborators that you can try to chain together. There's, you know, kind of previously the virtual biotech was a concept that was very much in fashion, right? Folks found out just In reality, when you try to do this, even though it …
AI assessment note: “their cost basis is definitely more competitive... and I think that's great for the industry”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q In retrospect, is there, was there like a right year to start a self-driving car company? You were so ahead of the ball on this.
A It takes a long time to sort of spin up automotive pipeline and everything. So probably like, you know, circa of 20, 20 or so would be the right time to get started. And then, you know, around now, I think you'd be having like a combination of hardware and software that's mature enough. If you have a, a nimble enough engineering team that's able to adopt new technologies when they pop up and quickly pull them into the pipeline, you're actually well positioned, even if you started a while ago and your tech stack was based on some older technologies, if you have all the infrastructure in place for validation and testing and training models and deploying on, you know, on public roads with test drivers, Um, I think you can go a lot, uh, a lot faster, uh, even if you have to like sort of rip out and change some of your tech staff to adapt with the times.
AI assessment note: “probably like, you know, circa of 20, 20 or so would be the right time”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q On this topic, how do you imagine deployment To, to work. You're obviously like, hey, Tesla had the right path here. Is there a Tesla analogy where you make billions along the way? Um, it's not obvious. Like, are there constraints you can put around it where you have tele-op, right? Or just a couple tasks or a more constrained environment in the messiness of a home?
A There's a number of ways to, to attack that. Approaches I've seen are like, you sell a really high priced robot today, uh, like a humanoid or, or something resembling a human. That's like fully teleopt and you just tell someone this is going to cost, I don't know, like something crazy, like 50,000 dollars and a thousand dollars a month, but it's a, it's the first robot you can buy that will like do stuff in your house. I think that's one approach to try to, um, you know, make money along the way. I think your market size is pretty small doing that, but that's, that's a viable approach. And the other side would be to sell robots that don't fulfill like the promise of a, of a household robot that does all your chores, but do like little bits of useful things. I just saw at CES this year, They have like little iRobot Roomba type things that have a tiny little hand that could come out and like pick up a sock that's in the way. And so these are like incremental approaches to sell products and get, you know, data and hopefully like learn about what it would take, um, or train, or train models or try things to, to be able to work up that, um, that ladder, I guess, to the Holy grail, which would be like, you know, a robot that, that takes the place of your butler and your housekeeper and just, just about anything else that you would ever want. You know, if you could have an infinite st…
AI assessment note: “these are like incremental approaches to sell products and get, you know, data”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q You've already been through the ringer once on the regulatory front with AV. If you could wave a magic wand, we just call Trump right now. What do you think is the right policy approach to make sure we have a competitive domestic robotics industry, if that's relevant?
A First of all, I mean, I think there should be tons of regulation on AV. It reminds me what's happening in the AV space and sort of You know, companies like Cruise got wiped out. Companies like Waymo are growing or are expanding, I think much more slowly than they should based on the merits and the safety of their technology. Uh, and it reminds me of like when the first airlines were formed in the U S and there was no FAA, there was no regulation. If you had basically any kind of plane crash, you would get sued out of existence and you just wiped out. And many of the first airlines are no longer in existence because of this. And so the FAA was created because You know, the government decided that actually we should have airlines and if they keep going out of business, no one's going to start an airline anymore. And so create the FAA, monitor airplane manufacturers and airlines, make sure they meet safety criteria and in exchange, give them, you know, reasonable protections and limits on liability so they can actually operate, you know, in a, in a society like the U S that hasn't happened for self-driving cars. And so I think the only candidate today, the only companies that stand a chance, um, are the ones where, you know, they can afford to, to take on that liability. Um, cause they're a giant tech company that makes money in other places. You know, otherwise it's very bleak. S…
AI assessment note: “make sure they meet safety criteria and in exchange, give them, you know, reasonable protections”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q You are obviously not focused just on property law now. Like, how do you think about the mission or scope of Harvey today?
A Yeah, um, so mostly we're, we're developing it for legal overall, but I would say that what we're building is the AI platform for legal and professional services, right? And if that sounds vague or it sounds like there aren't incredibly defined use cases for, you know, the small areas that we're building, that's on purpose. Like, the reality is If you are using these tools and you don't think that you can take basically AI and apply to X industry and transform the entire industry, I don't think you're thinking ambitiously enough, right? And I think it's really hard because, you know, these models can't just one shot all of these really complex legal tasks or, or in these other domains like tax and provide other professional services. And so what you have to do is you have to build a platform that is kind of constantly expanding and constantly collapsing. And so what I mean by that is you need to build specific features and maybe agentic workflows, et cetera, that can do parts of a task. And then you need to combine them all together. So the UI is simple and you don't have this like tentacle monster of a platform.
AI assessment note: “what we're building is the AI platform for legal and professional services”