Oct 5, 2024 · 41m · latent-space

[Paper Club] Berkeley Function Calling Paper Club! — Sam Julien, Writer

Sam Julien · 20m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Sam Julien presents the architectural evolution of the Berkeley Function Calling Leaderboard across three versions, leading into an interactive technical discussion on multi-turn evaluation, Gorilla open-source tooling, and future benchmarking projects.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 1.6 Guest teaching 1.1 Guest disagreement 0.8 The hosts pushing back 0.9
05100:0015:0030:002:01–4:57 · The hosts as informed peer 0/10 BFCL Version 1: Benchmark Foundations and AST Evaluation Sam Julien delivers an uninterrupted monologue introducing the Berkeley Function Calling Leaderboard version 1, focusing on code dataset composition and abstract syntax tree evaluation.4:58–7:32 · The hosts as informed peer 0/10 BFCL Version 2: Live Dataset and Data Quality Improvements Sam Julien continues his solo slide presentation detailing BFCL version 2 improvements, including deduplication with ROUGE scores, filtering, and live user datasets.7:33–13:22 · The hosts as informed peer 0/10 BFCL Version 3: Multi-Turn Function Calling and Graph Generation Sam Julien presents BFCL version 3, outlining the transition to multi-turn benchmarks, mock API codebase creation, and graph traversal for generating execution paths.13:22–17:10 · The hosts as informed peer 0/10 State-Based Evaluation and Common LLM Error Modes Sam Julien concludes the monologue portion by detailing state-based evaluation and specific LLM failure modes like redundant authentication and missing context awareness.17:19–24:30 · The hosts as informed peer 4/10 Discussion on Graph Generation, Personas, and API Evaluation A collaborative open discussion where a participant asks for clarification on graph generation and persona augmentation, and shares insights regarding Gorilla's API Zoo.24:31–28:32 · The hosts as informed peer 4/10 Exploring Model State Management and Gorilla Architecture Participants dig into how state tracking is implemented across diverse APIs and clarify whether Gorilla is an orchestrating API or a specialized model.28:32–33:04 · The hosts as informed peer 3/10 Gorilla CLI, Open Functions, and Hands-on Session Proposals The group explores practical applications, brainstorming an AI in Action session and discussing practical utilities of Gorilla CLI and Open Functions.33:14–38:53 · The hosts as informed peer 2/10 Participant Pitch: LLM Event Prediction Quality Benchmark A participant pitches a prediction-market benchmark for LLMs, met with mild confusion and a gentle redirect from the club host to move the discussion to Discord.2:01–4:57 · Guest teaching 0/10 BFCL Version 1: Benchmark Foundations and AST Evaluation Sam Julien delivers an uninterrupted monologue introducing the Berkeley Function Calling Leaderboard version 1, focusing on code dataset composition and abstract syntax tree evaluation.4:58–7:32 · Guest teaching 0/10 BFCL Version 2: Live Dataset and Data Quality Improvements Sam Julien continues his solo slide presentation detailing BFCL version 2 improvements, including deduplication with ROUGE scores, filtering, and live user datasets.7:33–13:22 · Guest teaching 0/10 BFCL Version 3: Multi-Turn Function Calling and Graph Generation Sam Julien presents BFCL version 3, outlining the transition to multi-turn benchmarks, mock API codebase creation, and graph traversal for generating execution paths.13:22–17:10 · Guest teaching 0/10 State-Based Evaluation and Common LLM Error Modes Sam Julien concludes the monologue portion by detailing state-based evaluation and specific LLM failure modes like redundant authentication and missing context awareness.17:19–24:30 · Guest teaching 3/10 Discussion on Graph Generation, Personas, and API Evaluation A collaborative open discussion where a participant asks for clarification on graph generation and persona augmentation, and shares insights regarding Gorilla's API Zoo.24:31–28:32 · Guest teaching 3/10 Exploring Model State Management and Gorilla Architecture Participants dig into how state tracking is implemented across diverse APIs and clarify whether Gorilla is an orchestrating API or a specialized model.28:32–33:04 · Guest teaching 2/10 Gorilla CLI, Open Functions, and Hands-on Session Proposals The group explores practical applications, brainstorming an AI in Action session and discussing practical utilities of Gorilla CLI and Open Functions.33:14–38:53 · Guest teaching 1/10 Participant Pitch: LLM Event Prediction Quality Benchmark A participant pitches a prediction-market benchmark for LLMs, met with mild confusion and a gentle redirect from the club host to move the discussion to Discord.2:01–4:57 · Guest disagreement 0/10 BFCL Version 1: Benchmark Foundations and AST Evaluation Sam Julien delivers an uninterrupted monologue introducing the Berkeley Function Calling Leaderboard version 1, focusing on code dataset composition and abstract syntax tree evaluation.4:58–7:32 · Guest disagreement 0/10 BFCL Version 2: Live Dataset and Data Quality Improvements Sam Julien continues his solo slide presentation detailing BFCL version 2 improvements, including deduplication with ROUGE scores, filtering, and live user datasets.7:33–13:22 · Guest disagreement 0/10 BFCL Version 3: Multi-Turn Function Calling and Graph Generation Sam Julien presents BFCL version 3, outlining the transition to multi-turn benchmarks, mock API codebase creation, and graph traversal for generating execution paths.13:22–17:10 · Guest disagreement 0/10 State-Based Evaluation and Common LLM Error Modes Sam Julien concludes the monologue portion by detailing state-based evaluation and specific LLM failure modes like redundant authentication and missing context awareness.17:19–24:30 · Guest disagreement 1/10 Discussion on Graph Generation, Personas, and API Evaluation A collaborative open discussion where a participant asks for clarification on graph generation and persona augmentation, and shares insights regarding Gorilla's API Zoo.24:31–28:32 · Guest disagreement 2/10 Exploring Model State Management and Gorilla Architecture Participants dig into how state tracking is implemented across diverse APIs and clarify whether Gorilla is an orchestrating API or a specialized model.28:32–33:04 · Guest disagreement 1/10 Gorilla CLI, Open Functions, and Hands-on Session Proposals The group explores practical applications, brainstorming an AI in Action session and discussing practical utilities of Gorilla CLI and Open Functions.33:14–38:53 · Guest disagreement 2/10 Participant Pitch: LLM Event Prediction Quality Benchmark A participant pitches a prediction-market benchmark for LLMs, met with mild confusion and a gentle redirect from the club host to move the discussion to Discord.2:01–4:57 · The hosts pushing back 0/10 BFCL Version 1: Benchmark Foundations and AST Evaluation Sam Julien delivers an uninterrupted monologue introducing the Berkeley Function Calling Leaderboard version 1, focusing on code dataset composition and abstract syntax tree evaluation.4:58–7:32 · The hosts pushing back 0/10 BFCL Version 2: Live Dataset and Data Quality Improvements Sam Julien continues his solo slide presentation detailing BFCL version 2 improvements, including deduplication with ROUGE scores, filtering, and live user datasets.7:33–13:22 · The hosts pushing back 0/10 BFCL Version 3: Multi-Turn Function Calling and Graph Generation Sam Julien presents BFCL version 3, outlining the transition to multi-turn benchmarks, mock API codebase creation, and graph traversal for generating execution paths.13:22–17:10 · The hosts pushing back 0/10 State-Based Evaluation and Common LLM Error Modes Sam Julien concludes the monologue portion by detailing state-based evaluation and specific LLM failure modes like redundant authentication and missing context awareness.17:19–24:30 · The hosts pushing back 1/10 Discussion on Graph Generation, Personas, and API Evaluation A collaborative open discussion where a participant asks for clarification on graph generation and persona augmentation, and shares insights regarding Gorilla's API Zoo.24:31–28:32 · The hosts pushing back 2/10 Exploring Model State Management and Gorilla Architecture Participants dig into how state tracking is implemented across diverse APIs and clarify whether Gorilla is an orchestrating API or a specialized model.28:32–33:04 · The hosts pushing back 1/10 Gorilla CLI, Open Functions, and Hands-on Session Proposals The group explores practical applications, brainstorming an AI in Action session and discussing practical utilities of Gorilla CLI and Open Functions.33:14–38:53 · The hosts pushing back 3/10 Participant Pitch: LLM Event Prediction Quality Benchmark A participant pitches a prediction-market benchmark for LLMs, met with mild confusion and a gentle redirect from the club host to move the discussion to Discord.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 26:40 Disputing Gorilla Model Architecture Framing

A participant pushes back on the premise that Gorilla acts as an orchestration API layer, insisting it is strictly a model trained to produce API calls.

Hardest push from the hosts ▶ 38:40 Moderator Intervenes to Redirect Conversation

The Paper Club host firmly cuts off the tangential LLM prediction market pitch to keep the meeting focused on the paper review schedule.

Biggest teaching moment ▶ 23:10 Participant Clarifies API Zoo Hallucination Benchmark

When Sam notes the lack of detail on hallucination metrics, a participant steps in with knowledge about Gorilla's API Zoo repository serving as ground truth.

The host holds their own ▶ 18:40 Sam Explains Directed Graph Edge Construction

Sam demonstrates domain mastery by explaining precisely how downstream input/output mappings form executable traversal paths in the benchmark.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
BFCL Version 1: Benchmark Foundations and AST Evaluation 0000 Sam Julien delivers an uninterrupted monologue introducing the Berkeley Function Calling Leaderboard version 1, focusing on code dataset composition and abstract syntax tree evaluation.
BFCL Version 2: Live Dataset and Data Quality Improvements 0000 Sam Julien continues his solo slide presentation detailing BFCL version 2 improvements, including deduplication with ROUGE scores, filtering, and live user datasets.
BFCL Version 3: Multi-Turn Function Calling and Graph Generation 0000 Sam Julien presents BFCL version 3, outlining the transition to multi-turn benchmarks, mock API codebase creation, and graph traversal for generating execution paths.
State-Based Evaluation and Common LLM Error Modes 0000 Sam Julien concludes the monologue portion by detailing state-based evaluation and specific LLM failure modes like redundant authentication and missing context awareness.
Discussion on Graph Generation, Personas, and API Evaluation 4311 A collaborative open discussion where a participant asks for clarification on graph generation and persona augmentation, and shares insights regarding Gorilla's API Zoo.
Exploring Model State Management and Gorilla Architecture 4322 Participants dig into how state tracking is implemented across diverse APIs and clarify whether Gorilla is an orchestrating API or a specialized model.
Gorilla CLI, Open Functions, and Hands-on Session Proposals 3211 The group explores practical applications, brainstorming an AI in Action session and discussing practical utilities of Gorilla CLI and Open Functions.
Participant Pitch: LLM Event Prediction Quality Benchmark 2123 A participant pitches a prediction-market benchmark for LLMs, met with mild confusion and a gentle redirect from the club host to move the discussion to Discord.

Statements from this episode (4)

Assertion Contradicted
Julien: BFCL v1 was the first dedicated LLM function calling benchmark
“So the first one that came out, like I said, in March was really the first of its kind to Make a function calling leaderboard.”
Sam Julien Oct 5, 2024 ▶ 2:10
Opinion
BFCL v1 introduced AST-based evaluation for executable functions
“And then the main, like, I think sort of the main innovation of this first version of the leaderboard was that it used this, it introduced this like abstract syntax tree. Evaluation where it would look at the executable functions and evaluate them”
Sam Julien Oct 5, 2024 ▶ 3:52
Assertion Supported
Julien: BFCL v3 shifted from AST to state-based evaluation
“So they really updated their evaluation from the abstract syntax tree to something called state-based evaluation, which we'll, we'll talk about in just a second.”
Sam Julien Oct 5, 2024 ▶ 8:24
Assertion Not yet assessed · timeframe Oct 2024
BFCL v3 generates evaluation tasks using graph edge construction
“They basically created their own API for the sake of this testing, and then did this like mapping to create a graph edge construction and like generate tasks through that.”
Sam Julien Oct 5, 2024 ▶ 10:46
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.