May 23, 2025 · 45m · a16z

Building AI Systems You Can Trust

Scott Clark · 34m spoken Matt Bornstein · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, Scott Clark, Co-founder and CEO of Distributional, discusses the paradigm shift in artificial intelligence from chasing raw model performance to establishing system trust, reliability, and enterprise governance. He outlines strategies for testing non-deterministic models, managing AI technical debt, and building robust platform engineering to safely scale production Gen AI.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 3.4 Guest teaching 3.8 Guest disagreement 1.1 The host pushing back 1.1
05100:0015:0030:0045:000:54–8:04 · The host as informed peer 3/10 Defining Machine Learning vs. Artificial Intelligence Matt asks foundational questions about the distinction between traditional ML and modern generative AI. Scott reframes the core challenge in AI from pure accuracy optimization to system trust and behavior verification based on his experience building SigOpt and leading AI at Intel.8:04–12:36 · The host as informed peer 2/10 Building Production AI Systems: Then and Now Matt prompts Scott on product owner worries in production AI. Scott educates the host on the three structural complexities of modern GenAI—non-determinism, non-stationarity, and component chaining—showing how upstream shifts propagate chaos downstream.12:36–17:17 · The host as informed peer 3/10 Establishing Enterprise Trust and Defining Model Behavior Matt synthesizes Scott's point using analogies like job interview behavior and AI bedside manner. Scott explains why enterprise trust requires testing non-performance characteristics such as toxicity, reading level, tone, and reasoning steps.17:17–26:46 · The host as informed peer 4/10 Centralization, Platform Engineering, and Mitigating Shadow AI Matt pushes back on whether developers actually want standardized platform gateways. Scott explains how enterprise engineering teams use uniform gateways to manage governance, shadow AI risks, and multi-model access.26:46–30:51 · The host as informed peer 3/10 Scaling AI in Production and the AI Confidence Gap Matt asks Scott for real-world stories of failed deployments. Scott explains the AI confidence gap where prototypes stall because teams fear how non-deterministic models and expanding RAG corpora behave under full production scale.30:51–38:08 · The host as informed peer 4/10 Statistical Testing Approaches and Unsupervised Change Detection Matt frames the technical discussion with a metaphor of putting AI systems in glass tanks with sensors. Scott delivers a detailed mathematical explanation of why distributional testing uses high-entropy weak estimators rather than standard evaluation metrics.38:08–41:43 · The host as informed peer 5/10 Managing AI Tech Debt, Prompt Hygiene, and Cost Trade-Offs Matt demonstrates high technical familiarity by applying Conway's Law to prompt engineering, noting how system prompts mirror corporate bureaucracy. Scott agrees and expands on how behavioral test suites prevent build failures during prompt refactoring.41:43–45:48 · The host as informed peer 3/10 AI Labs, Enterprise Dynamics, and AI Ops Matt highlights operational realities by asking who gets paged late at night when a production model fails. Scott outlines how enterprise platform layers connect frontier research labs with operational needs, predicting the rise of dedicated AI Ops teams.0:54–8:04 · Guest teaching 4/10 Defining Machine Learning vs. Artificial Intelligence Matt asks foundational questions about the distinction between traditional ML and modern generative AI. Scott reframes the core challenge in AI from pure accuracy optimization to system trust and behavior verification based on his experience building SigOpt and leading AI at Intel.8:04–12:36 · Guest teaching 5/10 Building Production AI Systems: Then and Now Matt prompts Scott on product owner worries in production AI. Scott educates the host on the three structural complexities of modern GenAI—non-determinism, non-stationarity, and component chaining—showing how upstream shifts propagate chaos downstream.12:36–17:17 · Guest teaching 4/10 Establishing Enterprise Trust and Defining Model Behavior Matt synthesizes Scott's point using analogies like job interview behavior and AI bedside manner. Scott explains why enterprise trust requires testing non-performance characteristics such as toxicity, reading level, tone, and reasoning steps.17:17–26:46 · Guest teaching 3/10 Centralization, Platform Engineering, and Mitigating Shadow AI Matt pushes back on whether developers actually want standardized platform gateways. Scott explains how enterprise engineering teams use uniform gateways to manage governance, shadow AI risks, and multi-model access.26:46–30:51 · Guest teaching 4/10 Scaling AI in Production and the AI Confidence Gap Matt asks Scott for real-world stories of failed deployments. Scott explains the AI confidence gap where prototypes stall because teams fear how non-deterministic models and expanding RAG corpora behave under full production scale.30:51–38:08 · Guest teaching 5/10 Statistical Testing Approaches and Unsupervised Change Detection Matt frames the technical discussion with a metaphor of putting AI systems in glass tanks with sensors. Scott delivers a detailed mathematical explanation of why distributional testing uses high-entropy weak estimators rather than standard evaluation metrics.38:08–41:43 · Guest teaching 2/10 Managing AI Tech Debt, Prompt Hygiene, and Cost Trade-Offs Matt demonstrates high technical familiarity by applying Conway's Law to prompt engineering, noting how system prompts mirror corporate bureaucracy. Scott agrees and expands on how behavioral test suites prevent build failures during prompt refactoring.41:43–45:48 · Guest teaching 3/10 AI Labs, Enterprise Dynamics, and AI Ops Matt highlights operational realities by asking who gets paged late at night when a production model fails. Scott outlines how enterprise platform layers connect frontier research labs with operational needs, predicting the rise of dedicated AI Ops teams.0:54–8:04 · Guest disagreement 1/10 Defining Machine Learning vs. Artificial Intelligence Matt asks foundational questions about the distinction between traditional ML and modern generative AI. Scott reframes the core challenge in AI from pure accuracy optimization to system trust and behavior verification based on his experience building SigOpt and leading AI at Intel.8:04–12:36 · Guest disagreement 1/10 Building Production AI Systems: Then and Now Matt prompts Scott on product owner worries in production AI. Scott educates the host on the three structural complexities of modern GenAI—non-determinism, non-stationarity, and component chaining—showing how upstream shifts propagate chaos downstream.12:36–17:17 · Guest disagreement 1/10 Establishing Enterprise Trust and Defining Model Behavior Matt synthesizes Scott's point using analogies like job interview behavior and AI bedside manner. Scott explains why enterprise trust requires testing non-performance characteristics such as toxicity, reading level, tone, and reasoning steps.17:17–26:46 · Guest disagreement 2/10 Centralization, Platform Engineering, and Mitigating Shadow AI Matt pushes back on whether developers actually want standardized platform gateways. Scott explains how enterprise engineering teams use uniform gateways to manage governance, shadow AI risks, and multi-model access.26:46–30:51 · Guest disagreement 1/10 Scaling AI in Production and the AI Confidence Gap Matt asks Scott for real-world stories of failed deployments. Scott explains the AI confidence gap where prototypes stall because teams fear how non-deterministic models and expanding RAG corpora behave under full production scale.30:51–38:08 · Guest disagreement 1/10 Statistical Testing Approaches and Unsupervised Change Detection Matt frames the technical discussion with a metaphor of putting AI systems in glass tanks with sensors. Scott delivers a detailed mathematical explanation of why distributional testing uses high-entropy weak estimators rather than standard evaluation metrics.38:08–41:43 · Guest disagreement 1/10 Managing AI Tech Debt, Prompt Hygiene, and Cost Trade-Offs Matt demonstrates high technical familiarity by applying Conway's Law to prompt engineering, noting how system prompts mirror corporate bureaucracy. Scott agrees and expands on how behavioral test suites prevent build failures during prompt refactoring.41:43–45:48 · Guest disagreement 1/10 AI Labs, Enterprise Dynamics, and AI Ops Matt highlights operational realities by asking who gets paged late at night when a production model fails. Scott outlines how enterprise platform layers connect frontier research labs with operational needs, predicting the rise of dedicated AI Ops teams.0:54–8:04 · The host pushing back 1/10 Defining Machine Learning vs. Artificial Intelligence Matt asks foundational questions about the distinction between traditional ML and modern generative AI. Scott reframes the core challenge in AI from pure accuracy optimization to system trust and behavior verification based on his experience building SigOpt and leading AI at Intel.8:04–12:36 · The host pushing back 1/10 Building Production AI Systems: Then and Now Matt prompts Scott on product owner worries in production AI. Scott educates the host on the three structural complexities of modern GenAI—non-determinism, non-stationarity, and component chaining—showing how upstream shifts propagate chaos downstream.12:36–17:17 · The host pushing back 1/10 Establishing Enterprise Trust and Defining Model Behavior Matt synthesizes Scott's point using analogies like job interview behavior and AI bedside manner. Scott explains why enterprise trust requires testing non-performance characteristics such as toxicity, reading level, tone, and reasoning steps.17:17–26:46 · The host pushing back 2/10 Centralization, Platform Engineering, and Mitigating Shadow AI Matt pushes back on whether developers actually want standardized platform gateways. Scott explains how enterprise engineering teams use uniform gateways to manage governance, shadow AI risks, and multi-model access.26:46–30:51 · The host pushing back 1/10 Scaling AI in Production and the AI Confidence Gap Matt asks Scott for real-world stories of failed deployments. Scott explains the AI confidence gap where prototypes stall because teams fear how non-deterministic models and expanding RAG corpora behave under full production scale.30:51–38:08 · The host pushing back 1/10 Statistical Testing Approaches and Unsupervised Change Detection Matt frames the technical discussion with a metaphor of putting AI systems in glass tanks with sensors. Scott delivers a detailed mathematical explanation of why distributional testing uses high-entropy weak estimators rather than standard evaluation metrics.38:08–41:43 · The host pushing back 1/10 Managing AI Tech Debt, Prompt Hygiene, and Cost Trade-Offs Matt demonstrates high technical familiarity by applying Conway's Law to prompt engineering, noting how system prompts mirror corporate bureaucracy. Scott agrees and expands on how behavioral test suites prevent build failures during prompt refactoring.41:43–45:48 · The host pushing back 1/10 AI Labs, Enterprise Dynamics, and AI Ops Matt highlights operational realities by asking who gets paged late at night when a production model fails. Scott outlines how enterprise platform layers connect frontier research labs with operational needs, predicting the rise of dedicated AI Ops teams.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 22:22 Pushback on platform abstraction benefits

Scott gently counters Matt's skepticism about platform gateways, arguing that uniform abstractions solve essential governance and model-switching hurdles for developers.

Hardest push from the host ▶ 22:22 Challenging platform developer appeal

Matt directly challenges Scott's narrative that developers want centralized platform layers, pointing out that abstraction between developers and LLMs can feel like unnecessary friction.

Biggest teaching moment ▶ 32:02 Mathematical rationale for weak estimators

Scott educates Matt on the core mathematical distinction between traditional LLM evals using strong estimators and distributional testing using many weak, high-entropy estimators to detect subtle behavior shifts.

The host holds their own ▶ 39:41 Conway's Law in prompt engineering

Matt demonstrates sharp domain knowledge by connecting Conway's Law to prompt engineering, observing how corporate system prompts reflect internal organizational structure and bureaucracy.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Defining Machine Learning vs. Artificial Intelligence 3411 Matt asks foundational questions about the distinction between traditional ML and modern generative AI. Scott reframes the core challenge in AI from pure accuracy optimization to system trust and behavior verification based on his experience building SigOpt and leading AI at Intel.
Building Production AI Systems: Then and Now 2511 Matt prompts Scott on product owner worries in production AI. Scott educates the host on the three structural complexities of modern GenAI—non-determinism, non-stationarity, and component chaining—showing how upstream shifts propagate chaos downstream.
Establishing Enterprise Trust and Defining Model Behavior 3411 Matt synthesizes Scott's point using analogies like job interview behavior and AI bedside manner. Scott explains why enterprise trust requires testing non-performance characteristics such as toxicity, reading level, tone, and reasoning steps.
Centralization, Platform Engineering, and Mitigating Shadow AI 4322 Matt pushes back on whether developers actually want standardized platform gateways. Scott explains how enterprise engineering teams use uniform gateways to manage governance, shadow AI risks, and multi-model access.
Scaling AI in Production and the AI Confidence Gap 3411 Matt asks Scott for real-world stories of failed deployments. Scott explains the AI confidence gap where prototypes stall because teams fear how non-deterministic models and expanding RAG corpora behave under full production scale.
Statistical Testing Approaches and Unsupervised Change Detection 4511 Matt frames the technical discussion with a metaphor of putting AI systems in glass tanks with sensors. Scott delivers a detailed mathematical explanation of why distributional testing uses high-entropy weak estimators rather than standard evaluation metrics.
Managing AI Tech Debt, Prompt Hygiene, and Cost Trade-Offs 5211 Matt demonstrates high technical familiarity by applying Conway's Law to prompt engineering, noting how system prompts mirror corporate bureaucracy. Scott agrees and expands on how behavioral test suites prevent build failures during prompt refactoring.
AI Labs, Enterprise Dynamics, and AI Ops 3311 Matt highlights operational realities by asking who gets paged late at night when a production model fails. Scott outlines how enterprise platform layers connect frontier research labs with operational needs, predicting the rise of dedicated AI Ops teams.

Statements from this episode (14)

Insight
Clark: Machine learning is normalized tech; AI is cutting-edge novelty
“Like machine learning is the stuff that's now become easy and then AI is all the fun new stuff. And then as soon as it stops becoming the cutting edge, Then it just becomes, oh, that's just machine learning.”
Scott Clark May 23, 2025 ▶ 1:20
Insight
Clark: System trust, not performance, limits enterprise AI value
“The thing that's holding back people getting value from these AI systems is not performance. It's not about squeezing out that last half a percent from some eval function or some performance metric. It's about being able to confidently trust these systems.”
Scott Clark May 23, 2025 ▶ 3:31
Insight
Clark: High-level LLM evaluations mask undesired AI system behaviors
“We're seeing people do the exact same thing again today with LLMs, where they're focusing on these high-level metrics, these end outputs, these performance evals, and that ends up masking all of these potentially undesired behaviors within the system itself.”
Scott Clark May 23, 2025 ▶ 4:05
Assertion Not checkable as stated
Clark: Generative AI platform leaders are traditional machine learning veterans
“A lot of the people who are now in charge of building Gen AI platforms or productionizing these massive use cases are the same people who built those original machine learning systems.”
Scott Clark May 23, 2025 ▶ 7:55
Insight
Clark: Evaluating end-to-end AI performance hides upstream system failures
“And what I think a lot of firms are running into right now is if you're only looking at that last step, if you're only looking at the system's performance as a whole, it can be very difficult to understand when, where, and why behaviors are shifting within thi…”
Scott Clark May 23, 2025 ▶ 12:13
Assertion Not checkable as stated
Clark: Enterprises are moving from generative AI prototypes to centralized platforms
“One thing that we've seen that's really interesting over the last year, year and a half is people have started to shift from kind of science project prototype land where they have a bunch of individual teams trying to roll their own stack and trying to like bu…”
Scott Clark May 23, 2025 ▶ 17:22
Assertion Not checkable as stated
Clark: Generative shadow AI exposes intellectual property to external SaaS vendors
“And it was a somewhat localized problem because like you're doing data science on your laptop versus now I'm just shipping off secret IP to some SaaS company or something like that.”
Scott Clark May 23, 2025 ▶ 19:30
Insight
Clark: AI adoption faces misaligned incentives between providers and enterprise users
“So one big complication is That sometimes the incentives are misaligned. So open AI obviously wants to create the best general purpose foundational models, but an individual business may want a model that solves a very specific problem a very specific way very…”
Scott Clark May 23, 2025 ▶ 24:09
Insight
Clark: An AI confidence gap leaves enterprise generative AI in prototypes
“We talked to a lot of firms that are terrified to cross this AI confidence gap from I've developed something that works good in, in, in theory. How do I actually scale it up in practice? And A lot of times we'll talk to individuals who say, every single time I…”
Scott Clark May 23, 2025 ▶ 27:29
Insight
Clark: Expanding RAG datasets with historical data degrades search quality
“And so RAG has obviously become very prevalent in a wide variety of industries and people use it for a lot of different things. We've spoken with different firms that they were like, okay, well, I'm just going to continue to add more and more data to the corpu…”
Scott Clark May 23, 2025 ▶ 28:46
Insight
Clark: AI testing should use many weak estimators to detect behavioral differences
“Instead of trying to come up with a small number of strong estimators for performance, where we want to be able to conclusively say A is better than B, Instead, what we want is a large number of potentially weak estimators to be able to determine whether or no…”
Scott Clark May 23, 2025 ▶ 32:03
Assertion Not checkable as stated
Clark: Enterprises avoid high-value AI use cases due to unwieldy risk
“We see some firms attacking the low hanging fruit internal chatbots to like ask questions about HR because they're afraid to take that leap to develop the difficult problem because it's so unwieldy and there is so much risk associated with it. A lot of the mos…”
Scott Clark May 23, 2025 ▶ 36:19
Insight
Bornstein: System prompts mirror company organizational structures like Conway's Law
“My biggest takeaway from this is the system prompt often reflects the organization it came from. Like it's some new form of Conway's law. You know, Conway's law is this thing that you kind of ship your org structure as a software company. Like I think as an AI…”
Matt Bornstein May 23, 2025 ▶ 39:56
Prediction Not checkable as stated
Clark: Production generative AI adoption will drive dedicated AI ops teams
“I think as we see the rise of these gen AI platforms, we're going to see the rise of more AI ops, the people who have to make sure the system's working and understand when it isn't and then fix it.”
Scott Clark May 23, 2025 ▶ 42:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.