Nov 22, 2023 · 34m · mad
How to Ship Reliable GenAI Apps: Humanloop CEO on LLM Observability, RAG & Rapid Iteration
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, Matt Turck interviews Raza Habib, CEO of Humanloop, to discuss how enterprise teams build, evaluate, and monitor production-ready LLM applications. They explore topics ranging from collaborative prompt engineering and automated evaluators to model adoption trends, agentic workflows, and AI safety regulation.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.1% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Raza rejects the assumption that traditional monitoring models like Datadog suffice for LLMs, arguing that passive logging is far less productive than integrated prompt iteration environments.
Hardest push from Matt ▶ 27:39 Challenging Enterprise AI HypeMatt directly challenges the prevailing industry narrative, asking whether real production deployments exist or if consultants are simply making money selling potential.
Biggest teaching moment ▶ 8:26 Evaluating the Evaluator GuardrailsRaza educates Matt on the recursive risk of using LLMs as evaluators, explaining how bias and ordering effects necessitate strict binary and numerical guardrails.
Matt holds his own ▶ 21:08 Grilling Guest on Competitive EcosystemMatt displays clear domain expertise by citing specific framework players like LangChain, LangSmith, and ChatGPT Enterprise to press Raza on category boundaries.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| What is Humanloop? Elevator Pitch and LLMOps Category | 2 | 5 | 1 | 1 | Matt asks foundational landscape questions about category definitions (MLOps vs LLMOps). Raza provides an extensive tutorial on how non-deterministic ML and subjective LLM outputs differ from classical software testing. | |
| Human Feedback and Automated LLM Evaluators | 1 | 4 | 0 | 0 | Matt acts primarily as a supportive interviewer, asking prompt questions. Raza outlines Humanloop's evaluation methods spanning human feedback triangulation and automated LLM evaluators. | |
| Implementing Guardrails for LLM Evaluators | 5 | 4 | 1 | 4 | Matt offers a sharp interjection about evaluating evaluators ('turtles all the way down') and challenges Raza on whether detection actually enables fixing. Raza explains how rapid prompt tweaking turns monitoring into active interventions. | |
| Democratizing AI Development Through Prompt Engineering | 3 | 4 | 0 | 2 | Matt asks for clarification between developer prompt engineering and end-user template prompt libraries. Raza clarifies the product focus while sharing that users currently hack Humanloop for prompt sharing. | |
| Closed vs. Open Source Model Trends and Adoption Patterns | 4 | 3 | 0 | 1 | Matt shows familiarity with specific open-source models, asking Raza to compare Llama 2, Falcon, and Mistral. Raza breaks down model adoption curves from GPT-4 prototyping to open-source fine-tuning. | |
| Tool Use, Function Calling, and Monitoring AI Agents | 2 | 4 | 1 | 1 | Matt checks his prep notes regarding tool calling and agent monitoring. Raza explains function calling mechanics and how JSON API outputs expand LLMs beyond text generation. | |
| Navigating the GenAI Ecosystem and Enterprise Needs | 4 | 4 | 1 | 3 | Matt presses Raza on competitive overlap with ecosystem tools like LangChain, LangSmith, and ChatGPT Enterprise. Raza details Humanloop's focus on enterprise team collaboration and guardrails. | |
| Target Customers, Go-To-Market Strategy, and Tool Centralization | 3 | 3 | 0 | 1 | Matt probes Humanloop's go-to-market strategy and ideal target profile. Raza explains why a sales-led inbound strategy works best to prevent ad-hoc prompt sprawl across enterprise departments. | |
| Current State of Enterprise GenAI Adoption | 4 | 3 | 1 | 4 | Matt challenges market sentiment by asking if enterprise GenAI is actually reaching production or if it is mostly hype and consultant fees. Raza offers grounded evidence of internal operational workflows in production. | |
| Perspectives on Open Source AI, Safety, and Governance | 3 | 5 | 0 | 1 | Matt invites Raza's perspective on the open versus closed AI safety debate. Raza delivers a highly nuanced discourse on capability thresholds and regulating end-use cases. |