LLM as judges

1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement May 15, 2026 by Ameya Bhatawdekar · across every show →

Everything said about LLM as judges, oldest first

May 15, 2026 positive
Insight
Bhatawdekar: Span-level scorers pinpoint errors in AI agent execution
“You can define those as deterministic functions, you know, implemented in code, or you can use LLM as judges, but then you can evaluate like, how did each span perform? And that can give you a fairly good way to zero in on problematic areas of your agents.”
Ameya Bhatawdekar May 15, 2026 ▶ 22:24 The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.