Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q this thing is broken. Kind of sucks. How do you think about problems that you want to solve versus research that you do to highlight some of the problems and then hoping that other people will participate? Like does everything that you talk about is on the Chroma roadmap basically, or Are you just advising people, hey, this is bad, work around it, but don't ask us to fix it.
A Going back to what I said a moment ago, like, Chroma's broad mandate is to make the process of building applications more like engineering and less like alchemy. Um, and so, you know, it's a pretty broad tent, but we're a small team and we can only focus on so many things. We've chosen to focus very much on one thing for now. And so I don't think that, I don't have the hubris to think that we can ourselves solve This stuff conclusively for a very dynamic and large and emerging industry. I think it does take a community. It does take, like, a rising tide of people all working together. We intentionally wanted to, like, make very clear that, like, we do not have any, like, commercial motivations in this research. You know, we do not posit any solutions. We don't tell people to use Chroma. It's just, here's the, here's the problem.
AI assessment note: “we do not have any, like, commercial motivations in this research.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q When you say that code embeddings are underrated, what do you think that is?
A Most people just take generic embedding models that are trained on the internet, and they try to use them for code. And like, it works okay for some use cases, but does it work great for all use cases? I don't know. Another way to think about these different primitives and what they're useful for, fundamentally we're trying to find signal. Text search works really well. Lexical search, text search works really well. When the person who's writing the query knows the data. If I want to search my Google Drive, I just, for the spreadsheet that has all my investors, I'm just going to type in CapTable, because I know there's a spreadsheet in my Google Drive called CapTable. Full text search, great, it's perfect. I'm a subject matter expert in my data. Now, if you wanted to find that file, and you didn't know that I had a spreadsheet called CapTable, you're going to type in the spreadsheet that has the list of all the investors, and of course, in embedding space, in semantic space, that's going to match. And so, I think, again, these are just like, Different tools, and it depends on, like, who's writing the queries. It depends on what expertise they have in the data. Like, what blend of those tools is gonna be the right fit. My guess is that, like, for code today, it's something like, 90% of queries or 85% of queries can be satisfactorily run with regex. Regex is obviously, like, the …
AI assessment note: “Most people just take generic embedding models that are trained on the internet”
Answered raw tape
D 4 · C 5 · P 4 · Cm 3 4.15
Q You said that memory is the benefit of context engineering. I think there's, uh, you, you had a rant on Twitter about stop making memory for AI so complicated. How do you think about memory? And what are like maybe the other benefits of context engineering that maybe we were not connecting together?
A I think memory is a good term. It is very legible to a wide population. Again, this is sort of just continuing the anthropomorphization of LLMs. You know, we ourselves understand how we are, we as humans use memory. We're very good at, well, some of us are very good at using memory to learn how to do tasks. And then those learnings being like flexible to new environments. And, you know, the idea of being able to like take an AI, sit down next to an AI, And then instruct it for 10 minutes or a few hours and kind of just, like, tell it what you want it to do, and it does something, and you say, hey, actually do this next time, the same that you would with a human. At the end of that 10 minutes, at the end of those few hours, the AI is able to do it now. And the same level of reliability that a human could do it, like, is an incredibly attractive and exciting vision. I think that that will happen. And I think that memory, again, is, like, the, memory is the term that, like, everybody can understand. Like, We all understand. Our moms all understand. And, and, and the benefits of memory are also very appealing and very attractive. But what is memory under the hood? It's still just context engineering, I think, which is the domain of how do you put the right information into the context window. And so, yeah, I think of memory as the benefit. Context engineering is the tool that gives…
AI assessment note: “what is memory under the hood? It's still just context engineering”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q Don't care. Uh, standard cyborg, mighty hive, and know it. What are those lessons there that you're applying to Chroma?
A Yeah, more, more than I can count. Um, I mean, it's a bit of a cliche, um, and it's very hard to be self-reflective and honest with yourself about a lot of this stuff, but I think Viewing your life as being very short and kind of a, you know, a vapor in the wind and therefore like only doing the work that you absolutely love doing and only doing work that you doing that work with people that you love spending time with and serving customers that you love serving is a very useful, like North star. Um, and you know, I mean, I'd be the North star to like print a ton of money in some sense, There may be faster ways to scam people into making five million dollars or whatever. Um, so, but if I reflect on, and I'm happy to go in more detail, obviously, but if I reflect on like my prior experiences, like I was always making trade-offs. I was making trade-offs with like the people that I was working with, or I was making trade-offs with the customer that I was serving. I was making trade-offs with like the technology and like how proud I was of it. And maybe it's sort of like an age thing. I don't know, but like, you know, the older that I get, I just more and more want to do the best work that I can. And I want that work. To not just be great work, but I also want it to be seen by the most number of people, because ultimately that is what impact looks like. You know, impact is not inve…
AI assessment note: “if I reflect on like my prior experiences, like I was always making trade-offs.”
Partly raw tape
D 3 · C 4 · P 3 · Cm 3 3.30
Q the vision clear to people when on the outside you have, oh, I'll just use PG vector or like, you know, whatever else the thing of the day is. Do you feel like that helps you bring people that are more aligned with the vision versus more of the missionary type on just joining this company before it's hot and maybe any learning that you have from recruiting early on?
A The upstream version of Conway's Law, like you ship your org chart, is you ship your culture, because I think your org chart is downstream of your company's culture. We've always placed an extremely high premium on that, on people that we actually have here on the team. Um, I think that the slope of our future growth is entirely dependent on the people that are here in this office. And, you know, that could mean going back to zero, that could mean, you know, linear growth, that could mean all kinds of versions of, like, hyper-linear growth. Exponential growth, hockey stick growth. And so, yeah, we've just really decided to hire very slowly and be really picky. And I don't know. I mean, you know, the future will determine whether or not that was the right decision, but I think having worked on a few startups before, like that was something that I really cared about was like, I just want to work with people that I love working with and like want to be shoulder to shoulder with in the trenches. And I think in independently execute on the level of like craft and quality that like We owe developers. And so that was how we chose to do it.
AI assessment note: “we've just really decided to hire very slowly and be really picky.”
Not addressed raw tape
D 1 · C 3 · P 3 · Cm 3 2.40
Q like this should be the primary context retrieval paradigm where when you build an agent, you effectively call out to another agent with all these sort of recursive Rerankers and summarizers or another agent with tools. Yep. Um, or do you sort of glom them onto a single agent? I don't know if you have an opinion, obviously, because agent is very ill-defined, but I'll just put it out there.
A You can pull that apart. So, you know, indexing by definition is a trade off. Like when you index data, you're trading write time performance for query time performance. You're making it slower to ingest data, but much faster to query data, which obviously scales as data sets get larger. And so like, if you're only grepping very small, you know, a 15 file code bases, they probably don't have to index it. And that's okay. If you want to search all of the open source dependencies Of that project. You all have done this before in VS code or cursor, right? You've like run a search over like the node modules folder. It takes a really long time to run that search. That's a lot of data. Like to make that indexed and sort of, you can make that trade off of write time performance or create time performance. Like that's what, that's what indexing is. Like just like demystify it. What is this, right? Like that's what it is. You know, embeddings are known for semantic similarity today. Embeddings is just a generic concept of like, Information compression. There's actually like many tools you can use embeddings for. I think embeddings for code are still extremely early and underrated, but regex is obviously an incredibly valuable tool. And, you know, we've actually worked on now inside of Chroma, both single load and distributed, we support regex search natively. So you can do regex search …
AI assessment note: “indexing by definition is a trade off. Like when you index data”