Scott Clark, CEO and co-founder of Distributional, discusses failure modes and performance issues enterprise teams encounter when deploying RAG systems.
“And so RAG has obviously become very prevalent in a wide variety of industries and people use it for a lot of different things. We've spoken with different firms that they were like, okay, well, I'm just going to continue to add more and more data to the corpus because more data is better. Like, of course. But this ends up messing up the retrieval mechanism. And so where before it all had very recent data and it was giving very good responses because people were asking about things that had a lot of recency. Now they put their whole history into it. Now it's grabbing old stuff and pretending like it's new.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Scott Clark
Insight
Clark: AI testing should use many weak estimators to detect behavioral differences
“Instead of trying to come up with a small number of strong estimators for performance, where we want to be able to conclusively say A is better than B, Instead, what we want is a large number of potentially weak estimators to be able to determine whether or no…”
Scott ClarkMay 23, 2025▶ 32:03Building AI Systems You Can Trust
Insight
Clark: Untuned deep learning models perform worse than tuned simple algorithms
“An untuned, sophisticated system will underperform a tuned simple system.”
Scott ClarkJan 2, 2019▶ 20:48a16z Podcast | AI, from 'Toy' Problems to Practical Application
Insight
Clark: System trust, not performance, limits enterprise AI value
“The thing that's holding back people getting value from these AI systems is not performance. It's not about squeezing out that last half a percent from some eval function or some performance metric. It's about being able to confidently trust these systems.”
Scott ClarkMay 23, 2025▶ 3:31Building AI Systems You Can Trust
Insight
Clark: High-level LLM evaluations mask undesired AI system behaviors
“We're seeing people do the exact same thing again today with LLMs, where they're focusing on these high-level metrics, these end outputs, these performance evals, and that ends up masking all of these potentially undesired behaviors within the system itself.”
Scott ClarkMay 23, 2025▶ 4:05Building AI Systems You Can Trust
Insight
Clark: Evaluating end-to-end AI performance hides upstream system failures
“And what I think a lot of firms are running into right now is if you're only looking at that last step, if you're only looking at the system's performance as a whole, it can be very difficult to understand when, where, and why behaviors are shifting within thi…”
Scott ClarkMay 23, 2025▶ 12:13Building AI Systems You Can Trust
PredictionNot checkable as stated
Clark: Production generative AI adoption will drive dedicated AI ops teams
“I think as we see the rise of these gen AI platforms, we're going to see the rise of more AI ops, the people who have to make sure the system's working and understand when it isn't and then fix it.”
Scott ClarkMay 23, 2025▶ 42:35Building AI Systems You Can Trust
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 1,000 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.