Jul 14, 2025 · 34m · latent-space

⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo

Pratik Bhavsar · 25m spoken Shawn Wang · 4m spoken Alessio Fanelli · 51s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Lightning Pod, hosts Swyx and Alessio Fanelli interview Pratik Bhavsar of Galileo Labs to explore the methodologies, surprising model rankings, and future multi-turn simulation architectures shaping AI agent evaluation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.6% of the talking time here. How this is scored →

The hosts as informed peer 4.3 Guest teaching 3.9 Guest disagreement 0.3 The hosts pushing back 0.7
05100:0010:0020:0030:000:39–2:56 · The hosts as informed peer 5/10 The Shift from General LLMs to Agent Evaluations Swyx frames the shift in LLM evaluation towards agentic tool calling and emphasizes the importance of cost-awareness metrics. Pratik completely agrees and corroborates with customer trends seen at Galileo.3:04–6:39 · The hosts as informed peer 3/10 Leaderboard Architecture and Tool Selection Quality Metric Alessio asks for an introductory breakdown of the benchmark architecture. Pratik gives a comprehensive overview of how Galileo aggregated 14 filtered datasets and built the Tool Selection Quality metric.6:42–11:26 · The hosts as informed peer 4/10 Surprising Leaderboard Results and Model Rankings Alessio highlights surprising findings like Mistral Small's performance, prompting Pratik to share unexpected takeaways like reasoning models failing multi-tool requests. Swyx briefly steps in to clarify a model name mix-up.11:29–16:33 · The hosts as informed peer 7/10 Comparative Evaluation Datasets and the Role of Tau-bench Swyx demonstrates up-to-the-minute domain knowledge by informing Pratik of Sierra's brand-new Tau-squared release and its telecom domain addition. Pratik explains the value of Tau-bench's deterministic database verification.16:33–21:19 · The hosts as informed peer 5/10 Validating LLM-as-a-Judge and the Tool Selection Quality Metric Swyx presses on calibration errors and industry skepticism around LLM-as-a-judge approaches. Pratik responds with a detailed explanation of their prompt engineering rigor, AUROC correlation validation, and majority voting.21:21–24:06 · The hosts as informed peer 4/10 Prompt Engineering and Model Selection for Evaluator Judges Alessio asks about prompt sensitivity and ranking shifts during judge model evaluation. Pratik explains the team's prompt refinement workflow and emphasizes the necessity of using frontier models like Claude 3.7 to judge difficult outputs.24:07–33:42 · The hosts as informed peer 2/10 Designing Agent Leaderboard V2 with Multi-Turn Simulation Swyx invites Pratik to outline Leaderboard V2. Pratik delivers a detailed breakdown of user simulation, synthetic tool response generation via GPT-4.1 mini, and the introduction of the Action Completion metric to avoid score saturation.0:39–2:56 · Guest teaching 1/10 The Shift from General LLMs to Agent Evaluations Swyx frames the shift in LLM evaluation towards agentic tool calling and emphasizes the importance of cost-awareness metrics. Pratik completely agrees and corroborates with customer trends seen at Galileo.3:04–6:39 · Guest teaching 3/10 Leaderboard Architecture and Tool Selection Quality Metric Alessio asks for an introductory breakdown of the benchmark architecture. Pratik gives a comprehensive overview of how Galileo aggregated 14 filtered datasets and built the Tool Selection Quality metric.6:42–11:26 · Guest teaching 4/10 Surprising Leaderboard Results and Model Rankings Alessio highlights surprising findings like Mistral Small's performance, prompting Pratik to share unexpected takeaways like reasoning models failing multi-tool requests. Swyx briefly steps in to clarify a model name mix-up.11:29–16:33 · Guest teaching 4/10 Comparative Evaluation Datasets and the Role of Tau-bench Swyx demonstrates up-to-the-minute domain knowledge by informing Pratik of Sierra's brand-new Tau-squared release and its telecom domain addition. Pratik explains the value of Tau-bench's deterministic database verification.16:33–21:19 · Guest teaching 5/10 Validating LLM-as-a-Judge and the Tool Selection Quality Metric Swyx presses on calibration errors and industry skepticism around LLM-as-a-judge approaches. Pratik responds with a detailed explanation of their prompt engineering rigor, AUROC correlation validation, and majority voting.21:21–24:06 · Guest teaching 4/10 Prompt Engineering and Model Selection for Evaluator Judges Alessio asks about prompt sensitivity and ranking shifts during judge model evaluation. Pratik explains the team's prompt refinement workflow and emphasizes the necessity of using frontier models like Claude 3.7 to judge difficult outputs.24:07–33:42 · Guest teaching 6/10 Designing Agent Leaderboard V2 with Multi-Turn Simulation Swyx invites Pratik to outline Leaderboard V2. Pratik delivers a detailed breakdown of user simulation, synthetic tool response generation via GPT-4.1 mini, and the introduction of the Action Completion metric to avoid score saturation.0:39–2:56 · Guest disagreement 0/10 The Shift from General LLMs to Agent Evaluations Swyx frames the shift in LLM evaluation towards agentic tool calling and emphasizes the importance of cost-awareness metrics. Pratik completely agrees and corroborates with customer trends seen at Galileo.3:04–6:39 · Guest disagreement 0/10 Leaderboard Architecture and Tool Selection Quality Metric Alessio asks for an introductory breakdown of the benchmark architecture. Pratik gives a comprehensive overview of how Galileo aggregated 14 filtered datasets and built the Tool Selection Quality metric.6:42–11:26 · Guest disagreement 1/10 Surprising Leaderboard Results and Model Rankings Alessio highlights surprising findings like Mistral Small's performance, prompting Pratik to share unexpected takeaways like reasoning models failing multi-tool requests. Swyx briefly steps in to clarify a model name mix-up.11:29–16:33 · Guest disagreement 0/10 Comparative Evaluation Datasets and the Role of Tau-bench Swyx demonstrates up-to-the-minute domain knowledge by informing Pratik of Sierra's brand-new Tau-squared release and its telecom domain addition. Pratik explains the value of Tau-bench's deterministic database verification.16:33–21:19 · Guest disagreement 1/10 Validating LLM-as-a-Judge and the Tool Selection Quality Metric Swyx presses on calibration errors and industry skepticism around LLM-as-a-judge approaches. Pratik responds with a detailed explanation of their prompt engineering rigor, AUROC correlation validation, and majority voting.21:21–24:06 · Guest disagreement 0/10 Prompt Engineering and Model Selection for Evaluator Judges Alessio asks about prompt sensitivity and ranking shifts during judge model evaluation. Pratik explains the team's prompt refinement workflow and emphasizes the necessity of using frontier models like Claude 3.7 to judge difficult outputs.24:07–33:42 · Guest disagreement 0/10 Designing Agent Leaderboard V2 with Multi-Turn Simulation Swyx invites Pratik to outline Leaderboard V2. Pratik delivers a detailed breakdown of user simulation, synthetic tool response generation via GPT-4.1 mini, and the introduction of the Action Completion metric to avoid score saturation.0:39–2:56 · The hosts pushing back 0/10 The Shift from General LLMs to Agent Evaluations Swyx frames the shift in LLM evaluation towards agentic tool calling and emphasizes the importance of cost-awareness metrics. Pratik completely agrees and corroborates with customer trends seen at Galileo.3:04–6:39 · The hosts pushing back 0/10 Leaderboard Architecture and Tool Selection Quality Metric Alessio asks for an introductory breakdown of the benchmark architecture. Pratik gives a comprehensive overview of how Galileo aggregated 14 filtered datasets and built the Tool Selection Quality metric.6:42–11:26 · The hosts pushing back 2/10 Surprising Leaderboard Results and Model Rankings Alessio highlights surprising findings like Mistral Small's performance, prompting Pratik to share unexpected takeaways like reasoning models failing multi-tool requests. Swyx briefly steps in to clarify a model name mix-up.11:29–16:33 · The hosts pushing back 1/10 Comparative Evaluation Datasets and the Role of Tau-bench Swyx demonstrates up-to-the-minute domain knowledge by informing Pratik of Sierra's brand-new Tau-squared release and its telecom domain addition. Pratik explains the value of Tau-bench's deterministic database verification.16:33–21:19 · The hosts pushing back 2/10 Validating LLM-as-a-Judge and the Tool Selection Quality Metric Swyx presses on calibration errors and industry skepticism around LLM-as-a-judge approaches. Pratik responds with a detailed explanation of their prompt engineering rigor, AUROC correlation validation, and majority voting.21:21–24:06 · The hosts pushing back 0/10 Prompt Engineering and Model Selection for Evaluator Judges Alessio asks about prompt sensitivity and ranking shifts during judge model evaluation. Pratik explains the team's prompt refinement workflow and emphasizes the necessity of using frontier models like Claude 3.7 to judge difficult outputs.24:07–33:42 · The hosts pushing back 0/10 Designing Agent Leaderboard V2 with Multi-Turn Simulation Swyx invites Pratik to outline Leaderboard V2. Pratik delivers a detailed breakdown of user simulation, synthetic tool response generation via GPT-4.1 mini, and the introduction of the Action Completion metric to avoid score saturation.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 64.4% · guest 35.6%0:00 · the hosts 64.4% · guest 35.6%3:00 · the hosts 9.3% · guest 90.7%3:00 · the hosts 9.3% · guest 90.7%6:00 · the hosts 6.8% · guest 93.2%6:00 · the hosts 6.8% · guest 93.2%9:00 · the hosts 19.2% · guest 80.8%9:00 · the hosts 19.2% · guest 80.8%12:00 · the hosts 33% · guest 67%12:00 · the hosts 33% · guest 67%15:00 · the hosts 34.3% · guest 65.7%15:00 · the hosts 34.3% · guest 65.7%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 9.1% · guest 90.9%21:00 · the hosts 9.1% · guest 90.9%24:00 · the hosts 6.2% · guest 93.8%24:00 · the hosts 6.2% · guest 93.8%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 5.3% · guest 94.7%30:00 · the hosts 5.3% · guest 94.7%33:00 · the hosts 13.8% · guest 86.2%33:00 · the hosts 13.8% · guest 86.2%
Sharpest disagreement ▶ 10:00 Firm correction on Mistral vs Meta models

When Swyx suggests Pratik misspoke and meant Meta rather than Mistral, Pratik immediately corrects the premise, reaffirming that Mistral 3.1 performed well while Llama models struggled.

Hardest push from the hosts ▶ 9:59 Swyx questions guest's model attribution

Swyx interrupts the flow to directly challenge Pratik's statement, asking if he actually meant Meta instead of Mistral.

Biggest teaching moment ▶ 7:59 Explaining reasoning models' failure in multi-tool calling

Pratik breaks down counterintuitive benchmark data explaining how OpenAI o-series models underperformed because they collapsed multiple required tool executions into single calls.

The host holds their own ▶ 13:14 Swyx introduces the freshly released Tau-squared benchmark

Swyx displays superior immediate awareness of the eval landscape by informing the guest about the previous day's Tau-squared release from Sierra, including its newly added telecom domain.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Shift from General LLMs to Agent Evaluations 5100 Swyx frames the shift in LLM evaluation towards agentic tool calling and emphasizes the importance of cost-awareness metrics. Pratik completely agrees and corroborates with customer trends seen at Galileo.
Leaderboard Architecture and Tool Selection Quality Metric 3300 Alessio asks for an introductory breakdown of the benchmark architecture. Pratik gives a comprehensive overview of how Galileo aggregated 14 filtered datasets and built the Tool Selection Quality metric.
Surprising Leaderboard Results and Model Rankings 4412 Alessio highlights surprising findings like Mistral Small's performance, prompting Pratik to share unexpected takeaways like reasoning models failing multi-tool requests. Swyx briefly steps in to clarify a model name mix-up.
Comparative Evaluation Datasets and the Role of Tau-bench 7401 Swyx demonstrates up-to-the-minute domain knowledge by informing Pratik of Sierra's brand-new Tau-squared release and its telecom domain addition. Pratik explains the value of Tau-bench's deterministic database verification.
Validating LLM-as-a-Judge and the Tool Selection Quality Metric 5512 Swyx presses on calibration errors and industry skepticism around LLM-as-a-judge approaches. Pratik responds with a detailed explanation of their prompt engineering rigor, AUROC correlation validation, and majority voting.
Prompt Engineering and Model Selection for Evaluator Judges 4400 Alessio asks about prompt sensitivity and ranking shifts during judge model evaluation. Pratik explains the team's prompt refinement workflow and emphasizes the necessity of using frontier models like Claude 3.7 to judge difficult outputs.
Designing Agent Leaderboard V2 with Multi-Turn Simulation 2600 Swyx invites Pratik to outline Leaderboard V2. Pratik delivers a detailed breakdown of user simulation, synthetic tool response generation via GPT-4.1 mini, and the introduction of the Action Completion metric to avoid score saturation.

Statements from this episode (12)

Insight
General LLM benchmark rankings do not translate to agent performance
“What we have understood from our previous experiments is that it's not necessary that what you see as top models, In let's say LM arena or other specific evaluation or general evaluations, they might not be also the same ranking for other tasks like agentic ta…”
Pratik Bhavsar Jul 14, 2025 ▶ 2:22
Insight
Top LLMs hold marginal performance edges over cheaper tiers
“Historically, what we have seen from our previous initiatives is that, okay, maybe the best GPT or best cloud is the top model, but there might be very small gap with the model just below it. And then it becomes a cost performance trade off so that users can k…”
Pratik Bhavsar Jul 14, 2025 ▶ 5:26
Assertion Contradicted
Google Gemini Flash-Lite ranks in top five on Galileo Agent Leaderboard
“In the, on the original leaderboard, we also have flashlight, which is, nobody talks about, I think, even now, but I think it's got a very decent score, which is, like, in the top five models. It's extremely cheap model, like, it's so dirt cheap that it makes …”
Pratik Bhavsar Jul 14, 2025 ▶ 7:50
Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Pratik Bhavsar Jul 14, 2025 ▶ 8:22
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:08
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:58
Opinion
LLM-as-a-judge methodology is now reliable enough for production evaluation
“Two years back where we were using our own similar LLMS judge methodology, and we had certain issues and there were some correlation issues at the time, but last year it got pretty strong and this year I feel it's so strong. LLMS judges is so strong that you c…”
Pratik Bhavsar Jul 14, 2025 ▶ 20:02
Insight
Changing LLM judge models can heavily alter benchmark leaderboards
“The leaderboard, if you are using different judges with different models, it can, there can be heavy shakeup of the leaderboard also.”
Pratik Bhavsar Jul 14, 2025 ▶ 22:45
Opinion
GPT-4o Mini cannot accurately evaluate complex Claude 3.7 outputs
“GPD for a mini can't find, ah, cannot evaluate the hard outputs that's 3.7 might be doing correctly, right?”
Pratik Bhavsar Jul 14, 2025 ▶ 23:12
Assertion Supported
Galileo Agent Leaderboard top scores have saturated at 0.95
“Then we saw with V-one that the scores are saturating. We see that the best score has already reached .95.”
Pratik Bhavsar Jul 14, 2025 ▶ 25:17
Opinion
Berkeley Function Calling Leaderboard is unsuitable for realistic agent evaluation
“This BFCL if you just open and check it you would be like, okay, this is okay, but you don't want to use it for realistic evaluation because it's like some math. Problem statement is given. Some simple tool is given just like, let's say some exponential tool, …”
Pratik Bhavsar Jul 14, 2025 ▶ 25:47
Opinion
Frontier LLMs remain unreliable at realistic multi-turn tool calling
“Our last leaderboard is saying that models are really great at tool calling. So it's like safe, but they actually not, right? They're making mistakes and this is going to recalibrate the expectation of the users that Be careful because they're still not perfec…”
Pratik Bhavsar Jul 14, 2025 ▶ 33:22
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.