Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Quick one. What does improve writing quality mean? Uh, what are the evals?
A What are the evals? How do you improve it? The way I'm thinking about it is like, have two various directions. The first direction is like, how do you improve the quality of the writing of the current use cases of Chachapi? And those, most of the use cases are mostly like non-fiction writings. It's like email writing or like some of the maybe blog posts. Cover letters is like one of the main use cases. But then the second one is like, how do we teach the model to literally think more creatively or like write In a more creative manner, such that it will just create novel forms writing. And I think the second one is like much of a longer term, like research question, while the first one is more like, okay, we just need to improve data quality for the writing use cases that between the models are. It is more straightforward question, but the way we evaluated the writing quality, so actually I worked with Jan's team On the model design. So they had a team of, like, model writers, and we would work together, and it's just like a human eval. It's like internal human eval where we would just- Always like that. Yeah, on the prompt distribution that we cared about, like, we want to make sure that the models that we, like, used, that we trained were always, like, better or something. Yeah.
AI assessment note: “It's like internal human eval where we would just... on the prompt distribution”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q What are your tips on O-one versus like cloud prompting, or like what are things that you took away from that experience? And especially now, I know that with Foro for Canvas, you've done RL after on the model. So yeah, just general learning. So now to think about prompting these models differently.
A I actually think like a one, I did not even harness the magic of like a one prompting, but like one thing that I found is that like, if you give a one like hard, like constraints of like what you're looking for, basically the model would be, we'll have a much easier time to like, kind of like select the candidates and, uh, match like the candidate that is most like, fulfill the criteria that you gave. And I think there's a. The class of problems like this that a one excels at. For example, if you have a question, like a bio question on like some, or like in chemistry, right? Like if you have like very specific criteria with the protein or like some of the chemical bindings or something, like then the model would be really, will be really good at like determining the exact candidate that will match the certain criteria.
AI assessment note: “if you give a one like hard, like constraints of like what you're looking for”
Answered raw tape
D 5 · C 4 · P 3 · Cm 3 3.90
Q We'll have, we'll interview Avi at some point. But, like, it sounds like the guiding principle is just what is useful to you. It's a little bit B to C, um, you know. Is there any B to B push at all? Or you don't think about that?
A I personally don't think about that as much, but I definitely feel like B to B is cool. Again, I come back to, like, Cloud and Slack. It's like one of the, like, the first, like, interfaces where, like, the model was operating inside your organization, right? It would be very cool for the model to, like, handle, to, like, become, like, a productive member of your organization. And then either, like, even, like, even process Like I, right now, like I'm thinking like processing, like user feedback. I think it'd be very cool if the model would just like start doing this for us and like, we don't have to hire a new person on this just for this or something. And like, you have like very simple, like data analysis, like data analytics or like how this feature is like.
AI assessment note: “I personally don't think about that as much, but I definitely feel like B to B”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q I think like, Like you weren't going to publish one and then you insisted on it or what?
A I think like, I just like put the section inside it and like, yeah, Jared, My, like, one of my most favorite researchers was like, yeah, that's cool. Let's, let's do that, I guess. Um, yeah, like, nobody had this, like, term of, like, behavioral design necessarily for the models. It's kind of like a new little field of, like, extending, like, product design into, like, the model design, right? Like, so how do you create a behavior for the model in certain contexts? So as, for example, like, in Canvas, right, like, one of the Things that we had to like think about was like, okay, like now the model enters like more collaborative environment, more collaborative context. So like, what's the most appropriate behavior for the model to act like as a collaborator? Should it ask like more follow-up questions? Should it like change? What's the tone should be like? What is the collaborator's tone? It's different from like a chat, like conversationalist versus like collaborator. So how do you shape the persona and the personality around that? It has like some philosophical questions too. Like, yeah, behavioral, I mean, like, I guess like I can talk more about like the methods of like creating the personality.
AI assessment note: “I just like put the section inside it and like, yeah, Jared”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q How do you think about the tasks? So I started creating a bunch of them. Like, do you see this as being, going back to like the composability, like composable together later, like, uh, you're going to be scheduled one task that does multiple tasks chained together. What's the vision?
A I would say task is like a foundational module, obviously to generalize to all sorts of like behaviors that you want. Like sometimes like I see like people have like three tasks in one query. And right now I don't think like the model handles this very well. I think that ideally we learn from like the user behavior and ideally the model will just be more proactive in suggesting of like Oh, I can either do this for you every day because I've observed that you do that every day or something. So it's like more becomes like a proactive behavior. I think right now you have to be more explicit, like, oh yeah, like every day, like remind me of this. But I think like the, the ideally the model will always think about you on the background and like kind of suggest, okay, like I noticed you've been reading Some, uh, this particular, like, Harkin news articles. Maybe I can try to suggest you, like, every day or something. So, like, it's just, like, much more, like, of a natural, like, friend, I think.
AI assessment note: “I think that ideally we learn from like the user behavior and ideally the model”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q Is that your definition of agents? Like, what are you looking for?
A I'm not sure if this is my definition of agents, but I feel like it's more like how I think. It makes sense, right? Like, I feel like for me to, like, trust an agent with my passwords or my credit card, I actually need to build trust with that agent that it will handle my tasks correctly and reliably. And the way I would go about this is how I would naturally like collaborate with other people is that like we first, even with any project, right? Like we first came, when we first come, like we don't even know each other. Like we don't know how each other's like working style, like what I prefer, what do they prefer, how do they prefer to communicate, et cetera, et cetera. So like you spend like the first, like, I don't know, like two weeks to just like learn their style of working. And then like over time you adapt To their working style. And then this is how you create the collaboration. And then like, at the beginning, you don't have much trust. So like, how do you build more trust? Especially like, it's the same thing as like with a manager, right? Like, it's like, how do you build trust with your manager? What does it need to know about you? What do you need to know about them? Over time, as you build trust and trust builds either through collaboration, which is why I feel like building Canvas was kind of like the first steps towards like more collaborative agents. I think w…
AI assessment note: “I feel like for me to, like, trust an agent with my passwords”