Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Opinion
LLM-as-a-judge methodology is now reliable enough for production evaluation
“Two years back where we were using our own similar LLMS judge methodology, and we had certain issues and there were some correlation issues at the time, but last year it got pretty strong and this year I feel it's so strong. LLMS judges is so strong that you c…”
Opinion
GPT-4o Mini cannot accurately evaluate complex Claude 3.7 outputs
“GPD for a mini can't find, ah, cannot evaluate the hard outputs that's 3.7 might be doing correctly, right?”
Opinion
Frontier LLMs remain unreliable at realistic multi-turn tool calling
“Our last leaderboard is saying that models are really great at tool calling. So it's like safe, but they actually not, right? They're making mistakes and this is going to recalibrate the expectation of the users that Be careful because they're still not perfec…”
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Opinion
Berkeley Function Calling Leaderboard is unsuitable for realistic agent evaluation
“This BFCL if you just open and check it you would be like, okay, this is okay, but you don't want to use it for realistic evaluation because it's like some math. Problem statement is given. Some simple tool is given just like, let's say some exponential tool, …”
Insight
General LLM benchmark rankings do not translate to agent performance
“What we have understood from our previous experiments is that it's not necessary that what you see as top models, In let's say LM arena or other specific evaluation or general evaluations, they might not be also the same ranking for other tasks like agentic ta…”
Insight
Top LLMs hold marginal performance edges over cheaper tiers
“Historically, what we have seen from our previous initiatives is that, okay, maybe the best GPT or best cloud is the top model, but there might be very small gap with the model just below it. And then it becomes a cost performance trade off so that users can k…”
Insight
Changing LLM judge models can heavily alter benchmark leaderboards
“The leaderboard, if you are using different judges with different models, it can, there can be heavy shakeup of the leaderboard also.”
Assertion Contradicted
Google Gemini Flash-Lite ranks in top five on Galileo Agent Leaderboard
“In the, on the original leaderboard, we also have flashlight, which is, nobody talks about, I think, even now, but I think it's got a very decent score, which is, like, in the top five models. It's extremely cheap model, like, it's so dirt cheap that it makes …”
Assertion Supported
Galileo Agent Leaderboard top scores have saturated at 0.95
“Then we saw with V-one that the scores are saturating. We see that the best score has already reached .95.”