Assertion Supported AI assessment confidence: 85% certainty 4/5 debate potential 3/5

Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks

Pratik Bhavsar · ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo · Jul 14, 2025 · at 10:08

Pratik Bhavsar discusses open-source model rankings on Galileo's AI Agent Leaderboard.

0:00 / 0:12exact quote · 13.0s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Pratik Bhavsar

Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Pratik Bhavsar Jul 14, 2025 ▶ 8:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
LLM-as-a-judge methodology is now reliable enough for production evaluation
“Two years back where we were using our own similar LLMS judge methodology, and we had certain issues and there were some correlation issues at the time, but last year it got pretty strong and this year I feel it's so strong. LLMS judges is so strong that you c…”
Pratik Bhavsar Jul 14, 2025 ▶ 20:02 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
GPT-4o Mini cannot accurately evaluate complex Claude 3.7 outputs
“GPD for a mini can't find, ah, cannot evaluate the hard outputs that's 3.7 might be doing correctly, right?”
Pratik Bhavsar Jul 14, 2025 ▶ 23:12 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
Frontier LLMs remain unreliable at realistic multi-turn tool calling
“Our last leaderboard is saying that models are really great at tool calling. So it's like safe, but they actually not, right? They're making mistakes and this is going to recalibrate the expectation of the users that Be careful because they're still not perfec…”
Pratik Bhavsar Jul 14, 2025 ▶ 33:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:58 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Opinion
Berkeley Function Calling Leaderboard is unsuitable for realistic agent evaluation
“This BFCL if you just open and check it you would be like, okay, this is okay, but you don't want to use it for realistic evaluation because it's like some math. Problem statement is given. Some simple tool is given just like, let's say some exponential tool, …”
Pratik Bhavsar Jul 14, 2025 ▶ 25:47 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.