Nov 22, 2023 · 34m · mad

How to Ship Reliable GenAI Apps: Humanloop CEO on LLM Observability, RAG & Rapid Iteration

Raza Habib · 25m spoken Matt Turck · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, Matt Turck interviews Raza Habib, CEO of Humanloop, to discuss how enterprise teams build, evaluate, and monitor production-ready LLM applications. They explore topics ranging from collaborative prompt engineering and automated evaluators to model adoption trends, agentic workflows, and AI safety regulation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.1% of the talking time here. How this is scored →

Matt as informed peer 3.1 Guest teaching 3.9 Guest disagreement 0.5 Matt pushing back 1.8
05100:0010:0020:0030:000:10–5:33 · Matt as informed peer 2/10 What is Humanloop? Elevator Pitch and LLMOps Category Matt asks foundational landscape questions about category definitions (MLOps vs LLMOps). Raza provides an extensive tutorial on how non-deterministic ML and subjective LLM outputs differ from classical software testing.5:33–8:04 · Matt as informed peer 1/10 Human Feedback and Automated LLM Evaluators Matt acts primarily as a supportive interviewer, asking prompt questions. Raza outlines Humanloop's evaluation methods spanning human feedback triangulation and automated LLM evaluators.8:04–12:17 · Matt as informed peer 5/10 Implementing Guardrails for LLM Evaluators Matt offers a sharp interjection about evaluating evaluators ('turtles all the way down') and challenges Raza on whether detection actually enables fixing. Raza explains how rapid prompt tweaking turns monitoring into active interventions.12:17–15:30 · Matt as informed peer 3/10 Democratizing AI Development Through Prompt Engineering Matt asks for clarification between developer prompt engineering and end-user template prompt libraries. Raza clarifies the product focus while sharing that users currently hack Humanloop for prompt sharing.15:30–18:54 · Matt as informed peer 4/10 Closed vs. Open Source Model Trends and Adoption Patterns Matt shows familiarity with specific open-source models, asking Raza to compare Llama 2, Falcon, and Mistral. Raza breaks down model adoption curves from GPT-4 prototyping to open-source fine-tuning.18:54–21:08 · Matt as informed peer 2/10 Tool Use, Function Calling, and Monitoring AI Agents Matt checks his prep notes regarding tool calling and agent monitoring. Raza explains function calling mechanics and how JSON API outputs expand LLMs beyond text generation.21:08–23:42 · Matt as informed peer 4/10 Navigating the GenAI Ecosystem and Enterprise Needs Matt presses Raza on competitive overlap with ecosystem tools like LangChain, LangSmith, and ChatGPT Enterprise. Raza details Humanloop's focus on enterprise team collaboration and guardrails.23:42–26:44 · Matt as informed peer 3/10 Target Customers, Go-To-Market Strategy, and Tool Centralization Matt probes Humanloop's go-to-market strategy and ideal target profile. Raza explains why a sales-led inbound strategy works best to prevent ad-hoc prompt sprawl across enterprise departments.26:44–28:46 · Matt as informed peer 4/10 Current State of Enterprise GenAI Adoption Matt challenges market sentiment by asking if enterprise GenAI is actually reaching production or if it is mostly hype and consultant fees. Raza offers grounded evidence of internal operational workflows in production.28:46–32:16 · Matt as informed peer 3/10 Perspectives on Open Source AI, Safety, and Governance Matt invites Raza's perspective on the open versus closed AI safety debate. Raza delivers a highly nuanced discourse on capability thresholds and regulating end-use cases.0:10–5:33 · Guest teaching 5/10 What is Humanloop? Elevator Pitch and LLMOps Category Matt asks foundational landscape questions about category definitions (MLOps vs LLMOps). Raza provides an extensive tutorial on how non-deterministic ML and subjective LLM outputs differ from classical software testing.5:33–8:04 · Guest teaching 4/10 Human Feedback and Automated LLM Evaluators Matt acts primarily as a supportive interviewer, asking prompt questions. Raza outlines Humanloop's evaluation methods spanning human feedback triangulation and automated LLM evaluators.8:04–12:17 · Guest teaching 4/10 Implementing Guardrails for LLM Evaluators Matt offers a sharp interjection about evaluating evaluators ('turtles all the way down') and challenges Raza on whether detection actually enables fixing. Raza explains how rapid prompt tweaking turns monitoring into active interventions.12:17–15:30 · Guest teaching 4/10 Democratizing AI Development Through Prompt Engineering Matt asks for clarification between developer prompt engineering and end-user template prompt libraries. Raza clarifies the product focus while sharing that users currently hack Humanloop for prompt sharing.15:30–18:54 · Guest teaching 3/10 Closed vs. Open Source Model Trends and Adoption Patterns Matt shows familiarity with specific open-source models, asking Raza to compare Llama 2, Falcon, and Mistral. Raza breaks down model adoption curves from GPT-4 prototyping to open-source fine-tuning.18:54–21:08 · Guest teaching 4/10 Tool Use, Function Calling, and Monitoring AI Agents Matt checks his prep notes regarding tool calling and agent monitoring. Raza explains function calling mechanics and how JSON API outputs expand LLMs beyond text generation.21:08–23:42 · Guest teaching 4/10 Navigating the GenAI Ecosystem and Enterprise Needs Matt presses Raza on competitive overlap with ecosystem tools like LangChain, LangSmith, and ChatGPT Enterprise. Raza details Humanloop's focus on enterprise team collaboration and guardrails.23:42–26:44 · Guest teaching 3/10 Target Customers, Go-To-Market Strategy, and Tool Centralization Matt probes Humanloop's go-to-market strategy and ideal target profile. Raza explains why a sales-led inbound strategy works best to prevent ad-hoc prompt sprawl across enterprise departments.26:44–28:46 · Guest teaching 3/10 Current State of Enterprise GenAI Adoption Matt challenges market sentiment by asking if enterprise GenAI is actually reaching production or if it is mostly hype and consultant fees. Raza offers grounded evidence of internal operational workflows in production.28:46–32:16 · Guest teaching 5/10 Perspectives on Open Source AI, Safety, and Governance Matt invites Raza's perspective on the open versus closed AI safety debate. Raza delivers a highly nuanced discourse on capability thresholds and regulating end-use cases.0:10–5:33 · Guest disagreement 1/10 What is Humanloop? Elevator Pitch and LLMOps Category Matt asks foundational landscape questions about category definitions (MLOps vs LLMOps). Raza provides an extensive tutorial on how non-deterministic ML and subjective LLM outputs differ from classical software testing.5:33–8:04 · Guest disagreement 0/10 Human Feedback and Automated LLM Evaluators Matt acts primarily as a supportive interviewer, asking prompt questions. Raza outlines Humanloop's evaluation methods spanning human feedback triangulation and automated LLM evaluators.8:04–12:17 · Guest disagreement 1/10 Implementing Guardrails for LLM Evaluators Matt offers a sharp interjection about evaluating evaluators ('turtles all the way down') and challenges Raza on whether detection actually enables fixing. Raza explains how rapid prompt tweaking turns monitoring into active interventions.12:17–15:30 · Guest disagreement 0/10 Democratizing AI Development Through Prompt Engineering Matt asks for clarification between developer prompt engineering and end-user template prompt libraries. Raza clarifies the product focus while sharing that users currently hack Humanloop for prompt sharing.15:30–18:54 · Guest disagreement 0/10 Closed vs. Open Source Model Trends and Adoption Patterns Matt shows familiarity with specific open-source models, asking Raza to compare Llama 2, Falcon, and Mistral. Raza breaks down model adoption curves from GPT-4 prototyping to open-source fine-tuning.18:54–21:08 · Guest disagreement 1/10 Tool Use, Function Calling, and Monitoring AI Agents Matt checks his prep notes regarding tool calling and agent monitoring. Raza explains function calling mechanics and how JSON API outputs expand LLMs beyond text generation.21:08–23:42 · Guest disagreement 1/10 Navigating the GenAI Ecosystem and Enterprise Needs Matt presses Raza on competitive overlap with ecosystem tools like LangChain, LangSmith, and ChatGPT Enterprise. Raza details Humanloop's focus on enterprise team collaboration and guardrails.23:42–26:44 · Guest disagreement 0/10 Target Customers, Go-To-Market Strategy, and Tool Centralization Matt probes Humanloop's go-to-market strategy and ideal target profile. Raza explains why a sales-led inbound strategy works best to prevent ad-hoc prompt sprawl across enterprise departments.26:44–28:46 · Guest disagreement 1/10 Current State of Enterprise GenAI Adoption Matt challenges market sentiment by asking if enterprise GenAI is actually reaching production or if it is mostly hype and consultant fees. Raza offers grounded evidence of internal operational workflows in production.28:46–32:16 · Guest disagreement 0/10 Perspectives on Open Source AI, Safety, and Governance Matt invites Raza's perspective on the open versus closed AI safety debate. Raza delivers a highly nuanced discourse on capability thresholds and regulating end-use cases.0:10–5:33 · Matt pushing back 1/10 What is Humanloop? Elevator Pitch and LLMOps Category Matt asks foundational landscape questions about category definitions (MLOps vs LLMOps). Raza provides an extensive tutorial on how non-deterministic ML and subjective LLM outputs differ from classical software testing.5:33–8:04 · Matt pushing back 0/10 Human Feedback and Automated LLM Evaluators Matt acts primarily as a supportive interviewer, asking prompt questions. Raza outlines Humanloop's evaluation methods spanning human feedback triangulation and automated LLM evaluators.8:04–12:17 · Matt pushing back 4/10 Implementing Guardrails for LLM Evaluators Matt offers a sharp interjection about evaluating evaluators ('turtles all the way down') and challenges Raza on whether detection actually enables fixing. Raza explains how rapid prompt tweaking turns monitoring into active interventions.12:17–15:30 · Matt pushing back 2/10 Democratizing AI Development Through Prompt Engineering Matt asks for clarification between developer prompt engineering and end-user template prompt libraries. Raza clarifies the product focus while sharing that users currently hack Humanloop for prompt sharing.15:30–18:54 · Matt pushing back 1/10 Closed vs. Open Source Model Trends and Adoption Patterns Matt shows familiarity with specific open-source models, asking Raza to compare Llama 2, Falcon, and Mistral. Raza breaks down model adoption curves from GPT-4 prototyping to open-source fine-tuning.18:54–21:08 · Matt pushing back 1/10 Tool Use, Function Calling, and Monitoring AI Agents Matt checks his prep notes regarding tool calling and agent monitoring. Raza explains function calling mechanics and how JSON API outputs expand LLMs beyond text generation.21:08–23:42 · Matt pushing back 3/10 Navigating the GenAI Ecosystem and Enterprise Needs Matt presses Raza on competitive overlap with ecosystem tools like LangChain, LangSmith, and ChatGPT Enterprise. Raza details Humanloop's focus on enterprise team collaboration and guardrails.23:42–26:44 · Matt pushing back 1/10 Target Customers, Go-To-Market Strategy, and Tool Centralization Matt probes Humanloop's go-to-market strategy and ideal target profile. Raza explains why a sales-led inbound strategy works best to prevent ad-hoc prompt sprawl across enterprise departments.26:44–28:46 · Matt pushing back 4/10 Current State of Enterprise GenAI Adoption Matt challenges market sentiment by asking if enterprise GenAI is actually reaching production or if it is mostly hype and consultant fees. Raza offers grounded evidence of internal operational workflows in production.28:46–32:16 · Matt pushing back 1/10 Perspectives on Open Source AI, Safety, and Governance Matt invites Raza's perspective on the open versus closed AI safety debate. Raza delivers a highly nuanced discourse on capability thresholds and regulating end-use cases.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 28.5% · guest 71.5%0:00 · Matt 28.5% · guest 71.5%3:00 · Matt 3.6% · guest 96.4%3:00 · Matt 3.6% · guest 96.4%6:00 · Matt 2.6% · guest 97.4%6:00 · Matt 2.6% · guest 97.4%9:00 · Matt 12% · guest 88%9:00 · Matt 12% · guest 88%12:00 · Matt 28.6% · guest 71.4%12:00 · Matt 28.6% · guest 71.4%15:00 · Matt 14.8% · guest 85.2%15:00 · Matt 14.8% · guest 85.2%18:00 · Matt 12.7% · guest 87.3%18:00 · Matt 12.7% · guest 87.3%21:00 · Matt 34.2% · guest 65.8%21:00 · Matt 34.2% · guest 65.8%24:00 · Matt 13.6% · guest 86.4%24:00 · Matt 13.6% · guest 86.4%27:00 · Matt 34.3% · guest 65.7%27:00 · Matt 34.3% · guest 65.7%30:00 · Matt 7.5% · guest 92.5%30:00 · Matt 7.5% · guest 92.5%33:00 · Matt 9.7% · guest 90.3%33:00 · Matt 9.7% · guest 90.3%
Sharpest disagreement ▶ 10:09 Rejecting Passive Monitoring Framing

Raza rejects the assumption that traditional monitoring models like Datadog suffice for LLMs, arguing that passive logging is far less productive than integrated prompt iteration environments.

Hardest push from Matt ▶ 27:39 Challenging Enterprise AI Hype

Matt directly challenges the prevailing industry narrative, asking whether real production deployments exist or if consultants are simply making money selling potential.

Biggest teaching moment ▶ 8:26 Evaluating the Evaluator Guardrails

Raza educates Matt on the recursive risk of using LLMs as evaluators, explaining how bias and ordering effects necessitate strict binary and numerical guardrails.

Matt holds his own ▶ 21:08 Grilling Guest on Competitive Ecosystem

Matt displays clear domain expertise by citing specific framework players like LangChain, LangSmith, and ChatGPT Enterprise to press Raza on category boundaries.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
What is Humanloop? Elevator Pitch and LLMOps Category 2511 Matt asks foundational landscape questions about category definitions (MLOps vs LLMOps). Raza provides an extensive tutorial on how non-deterministic ML and subjective LLM outputs differ from classical software testing.
Human Feedback and Automated LLM Evaluators 1400 Matt acts primarily as a supportive interviewer, asking prompt questions. Raza outlines Humanloop's evaluation methods spanning human feedback triangulation and automated LLM evaluators.
Implementing Guardrails for LLM Evaluators 5414 Matt offers a sharp interjection about evaluating evaluators ('turtles all the way down') and challenges Raza on whether detection actually enables fixing. Raza explains how rapid prompt tweaking turns monitoring into active interventions.
Democratizing AI Development Through Prompt Engineering 3402 Matt asks for clarification between developer prompt engineering and end-user template prompt libraries. Raza clarifies the product focus while sharing that users currently hack Humanloop for prompt sharing.
Closed vs. Open Source Model Trends and Adoption Patterns 4301 Matt shows familiarity with specific open-source models, asking Raza to compare Llama 2, Falcon, and Mistral. Raza breaks down model adoption curves from GPT-4 prototyping to open-source fine-tuning.
Tool Use, Function Calling, and Monitoring AI Agents 2411 Matt checks his prep notes regarding tool calling and agent monitoring. Raza explains function calling mechanics and how JSON API outputs expand LLMs beyond text generation.
Navigating the GenAI Ecosystem and Enterprise Needs 4413 Matt presses Raza on competitive overlap with ecosystem tools like LangChain, LangSmith, and ChatGPT Enterprise. Raza details Humanloop's focus on enterprise team collaboration and guardrails.
Target Customers, Go-To-Market Strategy, and Tool Centralization 3301 Matt probes Humanloop's go-to-market strategy and ideal target profile. Raza explains why a sales-led inbound strategy works best to prevent ad-hoc prompt sprawl across enterprise departments.
Current State of Enterprise GenAI Adoption 4314 Matt challenges market sentiment by asking if enterprise GenAI is actually reaching production or if it is mostly hype and consultant fees. Raza offers grounded evidence of internal operational workflows in production.
Perspectives on Open Source AI, Safety, and Governance 3501 Matt invites Raza's perspective on the open versus closed AI safety debate. Raza delivers a highly nuanced discourse on capability thresholds and regulating end-use cases.

Statements from this episode (5)

Insight
LLM observability must be tightly connected to prompt engineering tools
“Because prompt engineering allows you to intervene really quickly, I think it's very important to have evaluation and observability very closely connected to the prompt engineering tools.”
Raza Habib Nov 22, 2023 ▶ 12:08
Assertion Supported
LangChain's LangSmith is fundamentally focused on monitoring chains and agents
“Their LangSmith product probably has a little bit of overlap with us, but it is not, you know, it has some overlap that is fundamentally, I think focused more around monitoring chains and agents.”
Raza Habib Nov 22, 2023 ▶ 22:37
What-if
Restricting GPT-2 access would have cost years of AI progress
“It would have been really sad for the world, I think, if at the point of GPT-II, we had decided, hey, you know what? This is too dangerous. No one can have access. Cause we would have missed out on all of the past two or three years of incredible progress.”
Raza Habib Nov 22, 2023 ▶ 30:18
Opinion
Habib advocates regulating AI end-use cases rather than banning base models
“And so I'm generally in favor of finding ways to regulate end use cases and make the malicious use of the models be regulated and then punish that very strongly and make people responsible for the end outcomes that I am for banning the underlying technology.”
Raza Habib Nov 22, 2023 ▶ 30:39
Opinion
AI models will cross a capabilities threshold requiring restricted access
“Like there is a capabilities threshold above which it is correct that we would want to restrict access.”
Raza Habib Nov 22, 2023 ▶ 31:52
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.