Insight certainty 3/5 debate potential 3/5

Function calling benchmarks like BFCL are largely saturated

Michelle Pokrass · Building AGI with OpenAI's Structured Outputs API · Sep 17, 2024 · at 23:54

Michelle Pokrass, Tech Lead at OpenAI, discusses the limitations of current AI function calling evaluations with Shawn Wang.

0:00 / 0:20exact quote · 20.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I find that a lot of these evals are mostly saturated, like for BFCL. All the models are near, near the top. Already. And kind of the errors are more, I would say like just differences in default behaviors. I think most of the models on the leaderboard can kind of get a hundred percent with different prompting, but it's more kind of, you're just pulling apart different default defaults at this point.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Michelle Pokrass

Prediction Not checkable as stated
Pokrass predicts developers will abandon RAG vector stores for direct long-context
“So we do expect a lot of developers to start, you know, uploading their full context more directly to the model. So for smaller tasks, you maybe don't need The whole vector store.”
Michelle Pokrass Apr 15, 2025 ▶ 15:21 GPT 4.1: The New OpenAI Workhorse
Insight
Pokrass: AI model gains now driven by post-training, not larger pre-trains
“We find that actually a significant amount of the gains come from new post-training techniques. So I think in the past the narrative is that you need to pre-train these larger and larger models to get better performance, and we're finding that we're able to sq…”
Michelle Pokrass Apr 15, 2025 ▶ 7:57 GPT 4.1: The New OpenAI Workhorse
Insight
Pokrass: Prototype with GPT-4.1, then downscale for latency or upscale for reasoning
“I think the answer is always going to be the fastest model that accomplishes your task, right? So maybe you start prompting 4.1 as a starting point if it does your task super well, Then maybe you could drop down a 4.1 mini and save latency, or even nano. Where…”
Michelle Pokrass Apr 15, 2025 ▶ 27:58 GPT 4.1: The New OpenAI Workhorse
Opinion
Pokras: Vision fine-tuning is the most underrated release for bespoke OCR
“Vision fine-tuning is so underrated. For the past, like, two months, whenever I talk to founders, they tell me this is the thing they need most. A lot of people are doing, like, OCR on, on very bespoke formats, like government documents, and vision fine-tuning…”
Michelle Pokrass Oct 4, 2024 ▶ 56:21 Building AGI in Real Time (OpenAI Dev Day 2024)
Insight
Pokrass: Every successful company eventually outgrows Postgres for NoSQL
“At some point, every company gets the scale, every successful company gets the scale where Postgres is not cutting it. And then you migrate to some sort of NoSQL database.”
Michelle Pokrass Sep 17, 2024 ▶ 5:49 Building AGI with OpenAI's Structured Outputs API
Insight
Multi-step agentic apps fail at 95% reliability due to compounded errors
“Like if something is 95% reliable, but you're chaining together a bunch of calls, if you magnify that error rate, it makes your like application not work. So that's a really exciting thing here from going from like 95% to a hundred percent. I'm very biased wor…”
Michelle Pokrass Sep 17, 2024 ▶ 28:10 Building AGI with OpenAI's Structured Outputs API
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.