Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q like generically across, um, rag and other things. Like, how does that manifest? If I, um, if the, if the, if the chatbot, for example, mentions the name of, um, uh, you know, ASICs in the shoe example, what happens? Like, do I, do I control this or does it just say, uh, does it kill the answer? Like, how does it fail? What's my sort of end user experience?
A Yeah. Yeah. I think that's a great question. Uh, we thought about this a bunch. So the open source supports a bunch of different policies on failures, uh, that, you know, allow you to configure your level of risk. So in the extreme, uh, what we, what we support is this re-asking, you know, strategy that, you know, guardrails kind of, uh, supports really, really well, where it hooks into the LLM's ability to self heal. Uh, or correct their own outputs if you give them enough context, right? So how that works is that let's say you're, you have an LM output and it fails the validator that says you can't talk about, you know, your competitor. Um, and so what, what guardrails would do or guardrails re-asking would do is it would automatically construct a new prompt. Um, and only give it the information it needs to be able to correct its output. And then, you know, you send your prompt back to your LLM, get a new output. And more often than not, that output ends up being, you know, correct. So once again, it's not a hundred percent, you know, it's not, it doesn't work a hundred percent of the times, but it does end up working like pretty substantially in a lot of outcomes. So that is like one of the strategies of how do you handle failures, but there's, you know, a bunch of other things available, including programmatic fixes when possible, uh, right? So if there's any hallucinated t…
AI assessment note: “the open source supports a bunch of different policies on failures”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q you had a bunch of people that was, that were helping to build, uh, parts of this. So, um, is there like a very large number of validators and part of the idea is that the community is going to build some of those validators? Like, how do you, how do you think about, um, community, not just from a go to market perspective, but from a product development perspective?
A Yeah. Yeah. I think that's a great question. Uh, I think the community is definitely contributing like a number of these validators. I think, um, uh, I'm not sure when the episode will go out, but like this morning, you know, we had like a contribution for like a regex check. Uh, so making sure any output that you have, like matches a specific pattern that you predefined, uh, which was, you know, uh, contributed by the community, et cetera. So people, uh, it's very interesting because like the applications are so wide and varied and then based on that people end up, you know, making like Specific types of contributions. So I think we're definitely focused on that. I also do think that, um, the, um, like what the open source really ends up doing as, uh, you know, a framework is taking down this very abstract problem of what it means to do safe AI development, right? Like it's, it's a very abstract problem. It's almost an academic problem to some degree, and it takes that and it breaks that down into, you know, an engineering problem really, right? Like, okay, if these are the risk areas that you care about, This is exactly how you mitigate and safeguard against those risk areas, right? So it ends up like, um, having like a much more like, uh, uh, like these safety measures that are grounded in code rather than just in policy. Um, and so that I feel like is, is a big value of the…
AI assessment note: “community is definitely contributing like a number of these validators”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Um, just to jump into this, so there's this rag and there's fine tuning, uh, get like to very much to, uh, like you just did to make this, uh, interesting to a broad group of people. What does fine tuning actually mean? Everybody's sort of heard the term, uh, but what does it actually mean?
A Yeah, yeah, yeah, absolutely. So fine tuning, um, used to be, um, for the longest time, like before prompt engineering entered the scene, fine tuning was the primary way to make machine learning models like work for you. Uh, and the idea was that typically in research, uh, you know, uh, research labs or academia or other folks would end up designing like architectures, which are, you know, like, uh, ways of configuring, like, you know, uh, deep Deep learning neural networks together that like typically work really well for certain problems. And then you take that architecture as is, uh, start with like a random initialization of that architecture. And then, you know, do a bunch of passes where you just input your data and then try to make that architecture work really, really well, like for your data. So that was like more traditional. That was, you know, what, what training a deep learning model meant. Um, I think then we ended up like, In the trajectory of building these models ended up getting to a point where the model sizes ended up just being so large and so massive, uh, and the data that you'd need to actually get meaningful signal from, from those large architectures would end up being small, right? So your organization data would be much smaller, uh, than the size of the internet, for example. Um, so then, you know, we ended up in the space where we started doing like …
AI assessment note: “take the small amount of, like, custom organization data that you have, uh, and then”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q uh, like, how do you think about this? Because like the, the goal is to get to a hundred percent, presumably like accuracy, but you have something that's not at a hundred percent. You use something that's typically not a hundred percent either to correct it. Uh, so how do you think about the getting to that a hundred percent or, or can one eventually get to a hundred percent?
A Yeah. Yeah. I think it's a, it's a great question, honestly, and it's also like the key, uh, key, uh, kind of, um, Limitation of working with machine learning, which is like a hundred percent is almost never achievable. Uh, and so in practice, what you end up doing is like a, trying to, with an individual model, trying to improve performance as much as possible by training on better data, by using better techniques and doing a bunch of other things. Uh, and B, and this is, you know, the, this is like one of the things that guardrails really hooks into is using this machine learning kind of like concept of ensembling models together. Um, and what that typically, I think my favorite way of explaining what that does is that you can think of, of ensembling as stacking a bunch of different, like, sieves on top of each other, where each sieve has, you know, holes in, like, different areas, right? So the stack of different sieves ends up being so that, you know, uh, your holes are all, like, non-overlapping, and so you end up getting something that is greater than the sum of its parts and ends up being, you know, like, more watertight and more, like, correct. So ensembling kind of works in a different way, which is you have, like, your large, uh, you know, large language models such as OpenAI or Anthropic or Cohere or, you know, um, Lama too, um, but that has, like, specific failure m…
AI assessment note: “a hundred percent is almost never achievable”
Partly raw tape
D 4 · C 5 · P 5 · Cm 3 4.40
Q spent, um, like a lot of people in the space, a good amount of time with our, uh, vector database friends. And then this, like the whole concept of, uh, RAG, uh, that, um, you know, a lot of people have, I've had a lot of hope, uh, for, do you want to maybe talk about what that is and what's working about it? What's not yet working about it?
A Yeah, yeah, absolutely. Um, so RAG for, I think a lot of folks in your audience would know, but for folks that don't RAG stands for Retrieval Augmented Generation. Uh, and it's essentially a method of building with, uh, large language models so that you're not directly questioning them. You're first sub-selecting context from your organization, from your application that's helpful, uh, in putting that context into the prompt, and then, you know, like asking the LLM to answer only using the context you provided. Uh, so as an example, let's say you have, like, some internal documents, um, Uh, let's say a standard operating procedure at your company, and then you want employees to be able to ask questions from that, uh, standard operating procedure, you would first figure out, like, what the most relevant subsections of that SOP are, and then input that, like, amend, sorry, append that SOP subsections to your prompt, and then ask a question from the LM to, like, answer using the subsections you provide. So that's high level, you know, what retrieval augmented generation does. You retrieve the highlights, and then you augment the prompt and generate, yeah. Just for summarization. Um, so it is, uh, right now the primary way that most people are building with LLMs, uh, especially, I think, like, one of the, one of the most common interfaces that, uh, people end up building are chatbo…
AI assessment note: “right now the primary way that most people are building with LLMs”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q love to spend a little bit of time on the problem itself. So as, as, as users, we all now very familiar with the concept of hallucination, just using JGPD or other systems and, uh, getting results of, of varying levels of quality, let's say. Um, but what, what, what is the spectrum? Like in what spectrum, what, what is the range of ways, uh, generative AI can go bad?
A Yeah. So I think like hallucination is, um, hallucination is the top reason that everyone thinks of, and it definitely is a problem. Like I work with these systems, you know, not just obviously from a, from a, um, you know, from the creator of the open source and from the, you know, a founder of the company, but also like deeply as an engineer. And you do see the extent of the problem, even when the solution is like roughly right, it ends up being, you know, like actually Um, it still is hallucinating, like, facts I never gave it. So that definitely is a problem, but outside of that, there's a lot of, um, there's a lot of, like, implicit assumptions we make out of, you know, our existing workforce, and translating those assumptions into Gen AI-based systems is, is unsolved to, you know, right now. Um, so concretely, like, what some of these risks would end up looking like is, you know, compliance, for example, is a large one, where, um, there's, Two failures with Gen AI. One is that our existing frameworks of making AI compliant within specific settings, uh, suddenly break down. So existing, for example, like model risk management frameworks don't really apply when you haven't built the model yourself. You know, you didn't like curate the data that the model was trained on. And so you can't make any claims to that. Right. Um, I think the second one is that there's a lot of like…
AI assessment note: “compliance, for example, is a large one, where, um, there's Two failures with Gen AI.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q of hugging face, uh, open source and just, uh, doing my own work and, or Lama too, or whichever, um, is there, are all those models Created equal in terms of hallucination, or is it like open source that can heavily customize, fine tune, uh, to my own needs is less likely to hallucinate. Is there, is there a difference or, or it depends on what you do with them?
A Um, so always depends on what you do with them, but right now the models aren't equal. Uh, I think OpenAI is definitely leading the charge in terms of, you know, how effective, I guess just how flexible and generalizable the models are, uh, but like other players, including open source, are quickly catching up. Um, so, um, uh, let's see. Um, I think there's, uh, I'm thinking of like a benchmark called LegalBench. Uh, definitely recommend like everybody goes and check it out. Um, LegalBench was a benchmark that was released on a myriad of legal tasks, like, you know, any, any tasks that, you know, could be augmented or automated by AI in the legal domain. Uh, you know, there's, like, some benchmark for that in this, like, LegalBench, um, you know, set, and in that, they essentially, like, evaluated a bunch of models, including open source, and then some of, you know, you would typically see that, like, OpenAI-based models tend to do very well, On these tasks like they haven't seen before that are, you know, newly created benchmarks. So I think OpenAI, and specifically GPT-IV, is pretty powerful and potent in that, in, in terms of its flexibility. Um, in terms of, um, open source models, there's this one, uh, um, study that came out recently. It was a blog actually published by the AnyScale team. Would also definitely recommend, like, folks go out and check that out. Um, It was e…
AI assessment note: “always depends on what you do with them, but right now the models aren't equal.”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q um, and I have, uh, resources, I have engineers, and, um, I do that kind of like belt and suspender approach of doing both rag and fine tuning. Um, are they, are they, what, what does that mean in terms of hallucination? Are they benchmarks? Does that mean I'm like, I'm at Like, five percent of hallucination, or 15%, or like, how do people measure the severity of the problem?
A Yeah, I think it's a great question. I think you're kind of hitting the nail on, like, what the, what the big, uh, what the big, like, um, challenge in the space is right now, where evaluation is this, uh, you know, big open problem, and, uh, we're, we're at that point where we're already beyond where traditional academic metrics were, uh, and so, you know, it's kind of like the wild west a little bit out here. Um, in terms of, like, organizations that are building, you know, Doing this, like, two-step approach of first fine-tuning a model and then doing RAG on top of that first, um, props to them. Uh, I think, uh, it's, it's definitely, like, especially for fine-tuning, a big challenge ends up being, you know, just getting the data set, right? So for the last decade or so, we've heard that, you know, data is oil and, uh, data-centric machine learning and, you know, just getting that data set ready and together ends up being, you know, um, a substantial investment. Obviously, that investment often results in that lift. Uh, in model performance, and it allows you to work with, you know, much smaller models, uh, that work better for your data, that often have, like, much lower latencies, uh, but, you know, that, that in order to get those results, you typically need to make that initial investment. Um, I think, like, where it's hard to quantify, like, it's hard to give a blanket …
AI assessment note: “hard to give a blanket quantification of how much hallucination decreases by fine-tuning”