Jul 14, 2025 · 34m · latent-space
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Latent Space Lightning Pod, hosts Swyx and Alessio Fanelli interview Pratik Bhavsar of Galileo Labs to explore the methodologies, surprising model rankings, and future multi-turn simulation architectures shaping AI agent evaluation.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.6% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
When Swyx suggests Pratik misspoke and meant Meta rather than Mistral, Pratik immediately corrects the premise, reaffirming that Mistral 3.1 performed well while Llama models struggled.
Hardest push from the hosts ▶ 9:59 Swyx questions guest's model attributionSwyx interrupts the flow to directly challenge Pratik's statement, asking if he actually meant Meta instead of Mistral.
Biggest teaching moment ▶ 7:59 Explaining reasoning models' failure in multi-tool callingPratik breaks down counterintuitive benchmark data explaining how OpenAI o-series models underperformed because they collapsed multiple required tool executions into single calls.
The host holds their own ▶ 13:14 Swyx introduces the freshly released Tau-squared benchmarkSwyx displays superior immediate awareness of the eval landscape by informing the guest about the previous day's Tau-squared release from Sierra, including its newly added telecom domain.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| The Shift from General LLMs to Agent Evaluations | 5 | 1 | 0 | 0 | Swyx frames the shift in LLM evaluation towards agentic tool calling and emphasizes the importance of cost-awareness metrics. Pratik completely agrees and corroborates with customer trends seen at Galileo. | |
| Leaderboard Architecture and Tool Selection Quality Metric | 3 | 3 | 0 | 0 | Alessio asks for an introductory breakdown of the benchmark architecture. Pratik gives a comprehensive overview of how Galileo aggregated 14 filtered datasets and built the Tool Selection Quality metric. | |
| Surprising Leaderboard Results and Model Rankings | 4 | 4 | 1 | 2 | Alessio highlights surprising findings like Mistral Small's performance, prompting Pratik to share unexpected takeaways like reasoning models failing multi-tool requests. Swyx briefly steps in to clarify a model name mix-up. | |
| Comparative Evaluation Datasets and the Role of Tau-bench | 7 | 4 | 0 | 1 | Swyx demonstrates up-to-the-minute domain knowledge by informing Pratik of Sierra's brand-new Tau-squared release and its telecom domain addition. Pratik explains the value of Tau-bench's deterministic database verification. | |
| Validating LLM-as-a-Judge and the Tool Selection Quality Metric | 5 | 5 | 1 | 2 | Swyx presses on calibration errors and industry skepticism around LLM-as-a-judge approaches. Pratik responds with a detailed explanation of their prompt engineering rigor, AUROC correlation validation, and majority voting. | |
| Prompt Engineering and Model Selection for Evaluator Judges | 4 | 4 | 0 | 0 | Alessio asks about prompt sensitivity and ranking shifts during judge model evaluation. Pratik explains the team's prompt refinement workflow and emphasizes the necessity of using frontier models like Claude 3.7 to judge difficult outputs. | |
| Designing Agent Leaderboard V2 with Multi-Turn Simulation | 2 | 6 | 0 | 0 | Swyx invites Pratik to outline Leaderboard V2. Pratik delivers a detailed breakdown of user simulation, synthetic tool response generation via GPT-4.1 mini, and the introduction of the Action Completion metric to avoid score saturation. |