Pratik Bhavsar

AI Engineer, Cisco / Splunk · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

engineerauthor@ptkbhv ↗LinkedIn ↗pratik.ai ↗

Pratik Bhavsar specializes in AI evaluation, reliability, and agent observability, having built tools including Galileo's Agent Leaderboard and the Hallucination Index. He previously worked at Enterpret and Morningstar, founded the AI builder community Maxpool, and authored the Mastering GenAI book series covering agentic systems and LLM evaluations.

12statements → 5claims → 5claims resolved → 60%fully supported → 3.83/5average certainty → 2.33/5average debate potential →

3 supported 0 partly supported 2 contradicted how the 5 claims stand · each chip opens the sources

5 assertions · 4 opinions · 3 insights · every statement was checked. The predictions and assertions are the 5 claims: statements the public record can support or contradict. 5 are resolved. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Pratik argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Pratik Bhavsar Jul 14, 2025 ▶ 8:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo

Their most notable contradicted claim

Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:58 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Pratik Bhavsar on measured tape to publish a rate. This says nothing about how they speak.

Everything Pratik Bhavsar said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Pratik Bhavsar Jul 14, 2025 ▶ 8:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:08 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
LLM-as-a-judge methodology is now reliable enough for production evaluation
“Two years back where we were using our own similar LLMS judge methodology, and we had certain issues and there were some correlation issues at the time, but last year it got pretty strong and this year I feel it's so strong. LLMS judges is so strong that you c…”
Pratik Bhavsar Jul 14, 2025 ▶ 20:02 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
GPT-4o Mini cannot accurately evaluate complex Claude 3.7 outputs
“GPD for a mini can't find, ah, cannot evaluate the hard outputs that's 3.7 might be doing correctly, right?”
Pratik Bhavsar Jul 14, 2025 ▶ 23:12 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
Frontier LLMs remain unreliable at realistic multi-turn tool calling
“Our last leaderboard is saying that models are really great at tool calling. So it's like safe, but they actually not, right? They're making mistakes and this is going to recalibrate the expectation of the users that Be careful because they're still not perfec…”
Pratik Bhavsar Jul 14, 2025 ▶ 33:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:58 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
Berkeley Function Calling Leaderboard is unsuitable for realistic agent evaluation
“This BFCL if you just open and check it you would be like, okay, this is okay, but you don't want to use it for realistic evaluation because it's like some math. Problem statement is given. Some simple tool is given just like, let's say some exponential tool, …”
Pratik Bhavsar Jul 14, 2025 ▶ 25:47 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Insight
General LLM benchmark rankings do not translate to agent performance
“What we have understood from our previous experiments is that it's not necessary that what you see as top models, In let's say LM arena or other specific evaluation or general evaluations, they might not be also the same ranking for other tasks like agentic ta…”
Pratik Bhavsar Jul 14, 2025 ▶ 2:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Insight
Top LLMs hold marginal performance edges over cheaper tiers
“Historically, what we have seen from our previous initiatives is that, okay, maybe the best GPT or best cloud is the top model, but there might be very small gap with the model just below it. And then it becomes a cost performance trade off so that users can k…”
Pratik Bhavsar Jul 14, 2025 ▶ 5:26 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Insight
Changing LLM judge models can heavily alter benchmark leaderboards
“The leaderboard, if you are using different judges with different models, it can, there can be heavy shakeup of the leaderboard also.”
Pratik Bhavsar Jul 14, 2025 ▶ 22:45 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Contradicted
Google Gemini Flash-Lite ranks in top five on Galileo Agent Leaderboard
“In the, on the original leaderboard, we also have flashlight, which is, nobody talks about, I think, even now, but I think it's got a very decent score, which is, like, in the top five models. It's extremely cheap model, like, it's so dirt cheap that it makes …”
Pratik Bhavsar Jul 14, 2025 ▶ 7:50 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Supported
Galileo Agent Leaderboard top scores have saturated at 0.95
“Then we saw with V-one that the scores are saturating. We see that the best score has already reached .95.”
Pratik Bhavsar Jul 14, 2025 ▶ 25:17 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo

Appearances (1)

EpisodeDateSpeaking time
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo Jul 14, 2025 25m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.