LLM as judges
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement May 15, 2026 by Ameya Bhatawdekar · across every show →
Everything said about LLM as judges, oldest first
May 15, 2026 positive
Bhatawdekar: Span-level scorers pinpoint errors in AI agent execution
“You can define those as deterministic functions, you know, implemented in code, or you can use LLM as judges, but then you can evaluate like, how did each span perform? And that can give you a fairly good way to zero in on problematic areas of your agents.”