Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q I think on the customer side, they have a lot of data about their users who has bought, they have the action data. Can you kind of walk us through an example of what does someone come to you for? What questions would they want solved in the process of, do you customize a model for them? Do you have something off the shelf? What does that look like?
A Today, uh, when people leverage our models, it's often to better understand the population of their interest. So usually the start of the relationship Uh, we basically come together and hear about what population they want us to model, right? Um, so it might be that if you're a CPG company that's selling to all of the US, that might be fairly straightforward. You want to model the gem pop of the US, but at the same time, there is a vertical, or if there's a market that they're trying to go into, imagine, ah, they want to better understand, let's say people in their twenties and thirties living in California. That's a much more specific population. So we hear about these population and we go recruit these People, uh, with consent, uh, and with incentives, and we basically collect some of their data and create a model of these people. And then what our product allows you to do is basically query them, uh, so it can take us input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. Environment can literally be a survey questions. It can be behavioral experiments. It can be A-B testing. Oftentimes, the core use cases are things like concept testing, uh, to start with, but also, you know, people sometimes want to do focus group, or one of the sort of fun use cases that we also serve is actually even modeli…
AI assessment note: “when people leverage our models, it's often to better understand the population of their interest”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q your own solution to this, but how far off are we, and how do you check if it's grounded? Uh, you have some interesting stuff on your site that actually points to how you run real evals, but If you could take us through that side, you know, I think that's one of the big concerns that people have. They're like LLM's hallucinate. You're just hallucinating layer after layer, right?
A The way we do this, and this is actually the, the paper that we worked on after the Generative Agents paper that really became the, at least for a simile and also the field of simulation and synthetic panels, really became the foundation. Yeah, this is the paper. The paper is called Generative Agents Simulations of Thousand People. Here's what we've done. For this paper, we actually brought thousand people that's representatively sample from the US to a virtual app. And what we basically have done was we spent two hours collecting fairly wide-ranging data. In this particular study, we focused a lot on this interview data, uh, that was, uh, whose script was taken from this project called American Voices Project. And then we would also pair it up with a lot of behavior data and so forth, whatever we can collect within two hours. And then we would actually send these people away for a couple of weeks. And during that time, I would use this data to create their digital twins. And I would bring the humans participants back after two weeks and have them complete a battery of surveys, experiments, behavior studies. So we actually have the list here, which basically included things like the behavior economic games. We would run literally like big five personality tests, general social survey. We would also go ahead and run the randomized control trials that were published on PNAS. And …
AI assessment note: “we basically could replicate people's behaviors and attitudes, 85% as accurately”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Tell us about the company. You guys just raised a lot. You're half a research lab, half a company. I guess you're hiring. Are you based?
A Yeah. So we're based in Mission Rock. Uh, so not too far away from where we are right now. So we're in SF. Uh, but we are also bi-coastal. So we have our, uh, team, uh, I would say our headquarters in NSF, and we have a lot of our technical talent in NSF, and we do have a smaller office that just opened up actually in New York. We are, as a company, an interesting one in that today, obviously, there are AI Neolabs, and then there are AI product companies. Similarly, it truly is both. So this is a company that was founded by four co-founders, myself, Michael Bernstein, Percy Leung, Laney Yellen, Uh, Michael Percy and I are all researchers. So, of course, Michael was one of the co-authors of the ImageNet, kicks her the AI revolution back in 2013, has been instrumental in human-centered AI. Percy coined the term foundation model, and obviously, you know, one of the, the greats of the AI researchers today. And Lainey is my business counterpart, where she led some of the fastest growing AI native companies from their C to A and B. But we have this DNA at the company where The vision of the technology that we're creating is continuously developing that we are getting people who are basically my lab mates. We are right now about 60 or so people. 15%, almost 20% of the company population actually are just my lab mates. And we are, it's actually quite fun because many of them then had g…
AI assessment note: “we're based in Mission Rock... founded by four co-founders”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Do we want to keep going on the paper, uh, routes?
A Yeah, for sure. Uh, so the last one, uh, was sort of an interesting one. So this, uh, paper was the follow-up paper that we had, uh, to the Thousand Agents paper, where basically the idea was now can we augment the models even further and actually post-train the model based on a lot of randomized controlled trials. So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there's this, and there's this platform called Open Science Foundation. So some, uh, the audience might be familiar with this, and there has been, especially in the social sciences over the past five years or so, there has been this concern around replicability of studies, right? So it was a bit of a crisis that scientists acknowledged where we rerun the study and we don't actually see the same finding. It's, it's tough. And the reason why that was often the case was there's basically this survival bias where the papers that get published often need to maintain what we call the p-value of less than 0.05 in the experiments that we ran. That basically suggests that only, there's only five percent chance that the results that we saw is false positive. But the tricky part was all the papers that were not published, and there's still a five percent chance That whatever we publish is actually totally just randomly generated. Like there's a fi…
AI assessment note: “Yeah, for sure. Uh, so the last one, uh, was sort of an interesting one.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q What's the intuition between why you need to do it in the model?
A My intuition behind the actual, when do you train or even post train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it's operating in. So it has to learn new social physics. The places where it doesn't have to train is it already has the physics. We trust the physics. It already has the base statistics, but it's just trying to react to an environment. Then I think you can just prompt your way into getting the, you know, actions out of it. I don't think the model has yet, at least the models that are out in the open, has yet learned the complete mapping of social physics of humanity. Uh, this actually is one of the core thesis of simile, right? And one of the core reason why that is the case is if you look at the data that the model was trained on, these models were trained on the web data and whatever was available on the web. And these are really interesting data sets, but they are fundamentally the self exposed attitudinal data with some behavior data that's sprinkled around here and there. And It has yet to learn really deep behavioral nature of people. Not just what people say they do online, but they, what they actually do in real life. And this is actually one of the sort of, ah, what I would consider to be the dark knowledge of humanity that we haven't quite captured. And it's these kind of data that would also ne…
AI assessment note: “if the model has to learn the underlying physics of the world that it's operating in”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah. Are there other case studies? So you talked about CVS, talked about Gallup, Deloitte, Worldfront?
A Worldfront is an interesting one, um, because one of the things they were trying to do, they were one of the first customers that wanted to actually do product testing. That goes beyond just asking people what they think about, let's say, behavior experiments and so forth. So there, really what we Had to do was reason about multimodal input. So images, but also you can also imagine like these agents traversing through Figma mockups or websites. So some of the things that our agents can also do is you can be given a domain, like a website URL and actually go use it for a while. It's these kind of things. And Wealthfront was one of the first, uh, customers, uh, that was very excited about this possibility.
AI assessment note: “Worldfront is an interesting one, um, because one of the things they were trying”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q So like open claw, all these clients and personal agents, like what, what do you want to see from them that you're, that they don't currently have?
A They do think it's slowly getting there, but I do generally want them to have much deeper understanding of the person. Uh, right now you look at, uh, Models. I mean, open clause, and what it's basically leveraging is basically Markdown file, and I think it's quite clever, right? So if you look at the Generative Agents paper, this actually was the same intuition that we had, where initially when we were creating the memory architecture for the Generative Agents, and this is like back in twenty-twenty-two, so we didn't really quite have the idea of even agent architecture or the term agent, but the intuition that we shared with some of the work that's coming out today was We initially thought, well, do we want to make the memory into, let's say, knowledge graph? Do we want to train a bespoke model? All of these kind of things. And what we decided to do was, no, no, no, just forget about all this. These language models are actually quite good at modeling text and understanding and reasoning about text. So just put everything in a markdown file or a text file. You're done. I thought that was quite interesting that we could do that. And there's a lot of strength in doing that. But also there is limitation. It's the way you retrieve and make sense of data that's extremely large, it takes a lot of work. So I think that technology is getting better. I also do, however, think there are …
AI assessment note: “I do generally want them to have much deeper understanding of the person.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q Yeah. But you put people in the lab, they watch them sleep or what?
A So we do actually care a lot about the consent process. So people know that we are like, we invite them to be a member of this community to both share data and also have their selves represented in different forms. Um, but we bring a lot of people to the lab, uh, or virtual lab where we design experiments that would actually pose them real behavioral decisions. And often in these kinds of experimental setup, What makes the difference between what is attitudinal versus behavioral is if the stake in your decision is real. That's ultimately what makes it behavioral. So in these kind of setups, we are inspired by our colleagues in social sciences, psychology, and so forth. So when they run studies, what the kind of techniques they utilize is imagine there is a online store that you're inviting people to come by, and then whatever they purchase in this experiment, they actually get that item delivered, for instance. Like, these are the kind of things that makes the stakes real. So we run a lot of these experiments, and we also do partner with firms, um, and also right now we also have customers who are quite excited to at least give us a glimpse of the kind of behaviors that their users exhibit so that we can get a little bit deeper understanding of how people behave in these different platforms.
AI assessment note: “we bring a lot of people to the lab, uh, or virtual lab where we design experiments”
Redirected raw tape
D 3 · C 4 · P 4 · Cm 4 3.70
Q away from a car wash, it's a 10 minute drive, should I walk or drive? The model will say, oh, walk to the car wash, and you know, you don't have your car. Is anything like this a problem in simulation? You would assume, like, very simple for human to think about, but if the model is saying you should walk to the car wash, you know. Any, anything here?
A It's less, uh, what can we solve? But I think it's more about what biases or mistakes do people make that models miss? Like for instance, imagine that you are, you know, like the, when I was, instead of Stanford, I lived in Palo Alto, so it's about, I would say, a 40 minute walk from the campus. You ask the model, okay, let's go home. What can I, what can I do? It would likely call an Uber or, you know, give me, you know, the bus time. Um, but for, for the longest time, I actually really liked walking back. And the reason why I wanted to do that was not for efficiency. It actually really helped me think. And I like to walk for, you know, half an hour, 40 minutes or so a day, where I just get to, you know, you know, just think about ideas, research, just get lost in my thoughts. That's very human activity. Unless the model has seen that and actually understands the importance of that activity, It would actually miss these kinds of features. So that actually I think is fundamentally what we're trying to model. Like what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are.
AI assessment note: “It's less, uh, what can we solve? But I think it's more about”
Redirected raw tape
D 2 · C 4 · P 4 · Cm 3 3.25
Q What about a different domain? Say it was, what about all of Amazon data? Shopping data, right?
A Shopping data. So Amazon data is interesting in that it's very much behavioral. Although like what people do on social media, you could sort of squint and say that it's also behavioral, but the transaction data is always interesting. It is also most commonly available, however. If we were to look at purely social media, like if, if you really, you know, if I were, you know, if I had to really pick, Facebook likely is interesting because I actually do think it is most sort of a default version of people because you go to LinkedIn, it's very much professional environment. So people put up their, you know, you know, they have their guards up, right? And that still is interesting because that is true human attitude and behavior. But it is not your base state. Uh, you go to Twitter, and Twitter, people have their own crazy personas, uh, or depending on who you are, like my Twitter profile and, you know, persona is very much, initially was I was very much an academic. Hey, I'm here to share my studies. Now, uh, I share, uh, things that's related similarly. But Facebook is one of those more private space where people just connect with their friends. In that way, I actually do think it shows you a little bit more about who that person is. So if I had to pick, I'd likely pick Facebook.
AI assessment note: “If we were to look at purely social media... if I had to really pick, Facebook”