Oct 5, 2024 · 41m · latent-space
[Paper Club] Berkeley Function Calling Paper Club! — Sam Julien, Writer
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Sam Julien presents the architectural evolution of the Berkeley Function Calling Leaderboard across three versions, leading into an interactive technical discussion on multi-turn evaluation, Gorilla open-source tooling, and future benchmarking projects.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
A participant pushes back on the premise that Gorilla acts as an orchestration API layer, insisting it is strictly a model trained to produce API calls.
Hardest push from the hosts ▶ 38:40 Moderator Intervenes to Redirect ConversationThe Paper Club host firmly cuts off the tangential LLM prediction market pitch to keep the meeting focused on the paper review schedule.
Biggest teaching moment ▶ 23:10 Participant Clarifies API Zoo Hallucination BenchmarkWhen Sam notes the lack of detail on hallucination metrics, a participant steps in with knowledge about Gorilla's API Zoo repository serving as ground truth.
The host holds their own ▶ 18:40 Sam Explains Directed Graph Edge ConstructionSam demonstrates domain mastery by explaining precisely how downstream input/output mappings form executable traversal paths in the benchmark.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| BFCL Version 1: Benchmark Foundations and AST Evaluation | 0 | 0 | 0 | 0 | Sam Julien delivers an uninterrupted monologue introducing the Berkeley Function Calling Leaderboard version 1, focusing on code dataset composition and abstract syntax tree evaluation. | |
| BFCL Version 2: Live Dataset and Data Quality Improvements | 0 | 0 | 0 | 0 | Sam Julien continues his solo slide presentation detailing BFCL version 2 improvements, including deduplication with ROUGE scores, filtering, and live user datasets. | |
| BFCL Version 3: Multi-Turn Function Calling and Graph Generation | 0 | 0 | 0 | 0 | Sam Julien presents BFCL version 3, outlining the transition to multi-turn benchmarks, mock API codebase creation, and graph traversal for generating execution paths. | |
| State-Based Evaluation and Common LLM Error Modes | 0 | 0 | 0 | 0 | Sam Julien concludes the monologue portion by detailing state-based evaluation and specific LLM failure modes like redundant authentication and missing context awareness. | |
| Discussion on Graph Generation, Personas, and API Evaluation | 4 | 3 | 1 | 1 | A collaborative open discussion where a participant asks for clarification on graph generation and persona augmentation, and shares insights regarding Gorilla's API Zoo. | |
| Exploring Model State Management and Gorilla Architecture | 4 | 3 | 2 | 2 | Participants dig into how state tracking is implemented across diverse APIs and clarify whether Gorilla is an orchestrating API or a specialized model. | |
| Gorilla CLI, Open Functions, and Hands-on Session Proposals | 3 | 2 | 1 | 1 | The group explores practical applications, brainstorming an AI in Action session and discussing practical utilities of Gorilla CLI and Open Functions. | |
| Participant Pitch: LLM Event Prediction Quality Benchmark | 2 | 1 | 2 | 3 | A participant pitches a prediction-market benchmark for LLMs, met with mild confusion and a gentle redirect from the club host to move the discussion to Discord. |