Michelle Pokrass

7 statements across 2 episodes · 0 bullish · 4 bearish · 1 people on the record · first statement Sep 17, 2024 by Michelle Pokrass · said 3 times in 3 episodes since 2024 · across every show →

On the record as a speaker too: Michelle Pokrass's record, appearances and statements → this page counts the times other people say the name.

Mentions by year

brought up most by Alessio Fanelli (1)

tap a year for its mentions
00112220242025episodesmentions
01220242025episodes it came up in
000.511220242025episodesmentions per episode
2025 1 mention in 1 episode
2024 2 mentions in 2 episodes 1 per episode

every mention, scene by scene, with the transcript →

Everything said about Michelle Pokrass, oldest first

Sep 17, 2024
Assertion Not checkable as stated
OpenAI's API team had just five engineers before ChatGPT launched
“I would say the applied team was maybe like 30 or 40 people, and yeah, probably closer to 30, and there was maybe like five-ish total working on the API at most.”
Michelle Pokrass Sep 17, 2024 ▶ 8:37 Building AGI with OpenAI's Structured Outputs API
Sep 17, 2024 negative
Insight
Function calling benchmarks like BFCL are largely saturated
“I find that a lot of these evals are mostly saturated, like for BFCL. All the models are near, near the top. Already. And kind of the errors are more, I would say like just differences in default behaviors. I think most of the models on the leaderboard can kin…”
Michelle Pokrass Sep 17, 2024 ▶ 23:54 Building AGI with OpenAI's Structured Outputs API
Apr 15, 2025 negative
Insight
Open-source AI benchmarks omit critical tasks because they are hard to grade
“And these are useful instructions, but we find that many of the really interesting instructions are actually challenging to grade. And so the open source evals often don't have them.”
Michelle Pokrass Apr 15, 2025 ▶ 18:56 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 neutral
Disclosure
OpenAI uses its own models to categorize anonymized developer prompts
“Well, I will say we do use our own products internally where we can, and so we're not manually by hand reading every prompt. After they're like anonymized, we scrub them with any identifying data, then we use our models to take passes to categorize them.”
Michelle Pokrass Apr 15, 2025 ▶ 19:55 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 negative
Insight
Pokrass warns against close collaboration between AI evaluators and model developers
“Honestly, I think it's best when eval authors and model developers don't collab too much because you want things you know, as objective as possible, not trying to game any evals.”
Michelle Pokrass Apr 15, 2025 ▶ 17:38 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025
Assertion Not checkable as stated
GPT-4.1's multimodal vision improvements stem from pre-training, not post-training
“We talked about like coding instruction following long context, a lot of gains coming from post training, but in particular multimodal, like basically everything you're seeing, the gains are there from pre-training.”
Michelle Pokrass Apr 15, 2025 ▶ 35:09 GPT 4.1: The New OpenAI Workhorse
Apr 15, 2025 bearish
Prediction Not checkable as stated
Pokrass predicts developers will abandon RAG vector stores for direct long-context
“So we do expect a lot of developers to start, you know, uploading their full context more directly to the model. So for smaller tasks, you maybe don't need The whole vector store.”
Michelle Pokrass Apr 15, 2025 ▶ 15:21 GPT 4.1: The New OpenAI Workhorse
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.