AI Agent Evaluation
topic on 2 shows · 3 statements across 3 episodes
3 statements about AI Agent Evaluation, every show
Massa: AI agent evals should measure conversion, not call volume or duration
“Like I see companies like measuring number of calls or minutes during the call or some like superficial KPIs that give you some information, but that doesn't really work. Like the important thing is Did this customer convert? Is it bringing value to the custom…”
Petersson: Reducing AI agent evaluations to scalar metrics discards critical trace data
“When you run it for that long, you create so much data and to just say like, oh, the number is X. And then you throw away everything else. That's just very wasteful. There's so much insight from the things leading up to that number and reading the traces is li…”