Galileo

includes Galileo Agent Leaderboard

0 statements across 0 episodes · 3 bullish · 2 bearish · 1 people on the record · first statement Jul 14, 2025 by Pratik Bhavsar · said 13 times in 1 episodes since 2025 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (5), Pratik Bhavsar (5), Alessio Fanelli (3)

tap a year for its mentions
00811512025episodesmentions
0112025episodes it came up in
007.50.51512025episodesmentions per episode
2025 13 mentions in 1 episode

every mention, scene by scene, with the transcript →

Everything said about Galileo, oldest first

Jul 14, 2025 bullish
Assertion Contradicted
Google Gemini Flash-Lite ranks in top five on Galileo Agent Leaderboard
“In the, on the original leaderboard, we also have flashlight, which is, nobody talks about, I think, even now, but I think it's got a very decent score, which is, like, in the top five models. It's extremely cheap model, like, it's so dirt cheap that it makes …”
Pratik Bhavsar Jul 14, 2025 ▶ 7:50 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Jul 14, 2025 bullish
Opinion
LLM-as-a-judge methodology is now reliable enough for production evaluation
“Two years back where we were using our own similar LLMS judge methodology, and we had certain issues and there were some correlation issues at the time, but last year it got pretty strong and this year I feel it's so strong. LLMS judges is so strong that you c…”
Pratik Bhavsar Jul 14, 2025 ▶ 20:02 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Jul 14, 2025 negative
Assertion Supported
OpenAI o-series reasoning models fail at multi-tool calling benchmarks
“Then another surprise for me was that the reasoning models were not performing well enough. They had certain kind of limitation when we probed into it, like, why are they scoring less overall? They were like the O-one, the O-four, O-three, they, When not perfo…”
Pratik Bhavsar Jul 14, 2025 ▶ 8:22 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Jul 14, 2025 bullish
Assertion Contradicted
Claude 3.7 remains unbeaten on Galileo Agent Leaderboard
“When we released the leaderboard and just in a week that launched 3.7, And that went straight up, and nobody has beaten it so far.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:58 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Jul 14, 2025 bearish
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Pratik Bhavsar Jul 14, 2025 ▶ 10:08 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Jul 14, 2025 neutral
Assertion Supported
Galileo Agent Leaderboard top scores have saturated at 0.95
“Then we saw with V-one that the scores are saturating. We see that the best score has already reached .95.”
Pratik Bhavsar Jul 14, 2025 ▶ 25:17 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Jul 14, 2025 neutral
Insight
Changing LLM judge models can heavily alter benchmark leaderboards
“The leaderboard, if you are using different judges with different models, it can, there can be heavy shakeup of the leaderboard also.”
Pratik Bhavsar Jul 14, 2025 ▶ 22:45 ⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.