Oct 19, 2024 · 1h 1m · latent-space

[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu

Jesse Hu · 34m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Paper Club presentation, Jesse Hu and host Eugene explore the architecture, evolution, and practical realities of SWE-bench, SWE-bench Verified, SWE-bench Multimodal, and MLE-bench. They analyze how multi-step agent scaffolding, human verification pipelines, and structured retrieval strategies bridge the gap between frontier language models and complex software engineering tasks.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 2.7 Guest teaching 3.9 Guest disagreement 0.3 The hosts pushing back 1.2
05100:0015:0030:0045:001:00:000:04–2:06 · The hosts as informed peer 1/10 Technical Setup and Screen Sharing Confirmation Jesse starts off the session introducing the purpose of SWE-bench compared to HumanEval. Eugene facilitates the tech setup without confrontation.2:07–4:17 · The hosts as informed peer 0/10 Mining Pull Requests and Constructing SWE-bench Jesse gives an uninterrupted overview of how SWE-bench mines GitHub pull requests and test patches across popular Python repositories.4:20–8:22 · The hosts as informed peer 3/10 Explaining Fail-to-Pass and Pass-to-Pass Unit Tests Eugene asks for clarification regarding how unit tests in pull requests are isolated and evaluated. Jesse breaks down the distinction between fail-to-pass and pass-to-pass tests.8:24–13:55 · The hosts as informed peer 4/10 Deconstructing Benchmarks and Comparing OpenAI o1 Eugene relays audience questions comparing SWE-bench to OpenAI o1 and chain-of-thought methods. Jesse clarifies the distinction between agent scaffolds and raw model capabilities.13:58–17:05 · The hosts as informed peer 0/10 Environment Setup Challenges and Cheating Vulnerabilities Jesse details the practical engineering pain points of reproducing environments via Docker and the vulnerabilities that allow agents to cheat by leaking future information.17:05–21:09 · The hosts as informed peer 0/10 Introducing SWE-bench Verified and Solvability Criteria Jesse covers the motivation behind OpenAI SWE-bench Verified, explaining how human annotators evaluated issue solvability and filtered ambiguous tasks.21:10–23:55 · The hosts as informed peer 7/10 Human Annotation Pipeline and Quality Filtering A participant chimes in with deep practical experience on human annotation calibration, non-neutral 4-point scales, and conservative ensembling methods.24:02–28:36 · The hosts as informed peer 6/10 Human Demonstrations vs LLM Verification and CriticGPT The participant references OpenAI CriticGPT and context-length degradation curves, sparking a technical discussion on retrieval vs long-context LLMs.28:39–33:13 · The hosts as informed peer 2/10 Analyzing Top Agent Strategies from Execution Trajectories Jesse shares practical insights derived from inspecting agent trajectories in the benchmark experiments repo, contrasting Agentless with Honeycomb and Guru.33:14–38:56 · The hosts as informed peer 0/10 SWE-bench Multimodal for JavaScript and UI Tasks Jesse reviews the SWE-bench Multimodal paper, highlighting the shift to JavaScript/TypeScript and the strict requirement for visual assets in problem statements.38:57–42:53 · The hosts as informed peer 7/10 Challenges in Objective UI and Web Agent Evaluation Eugene steps in with domain authority from running a UI testing startup, highlighting human unreliability and multi-screen PSD translation subjectivity.42:55–47:57 · The hosts as informed peer 0/10 Introducing MLE-bench and Autonomous Kaggle Agents Jesse outlines MLE-bench by OpenAI, detailing how agents are evaluated on end-to-end Kaggle competitions with standard compute budgets and submission APIs.47:58–52:12 · The hosts as informed peer 1/10 Practical Limitations, Costs, and Real-World ML Disconnect Jesse performs cost math on MLE-bench runs and critiques the gap between clean Kaggle competition definitions and messy real-world data science workflows.52:12–54:51 · The hosts as informed peer 6/10 Evaluating Kaggle Data Contamination and Overfitting Eugene pushes back regarding Kaggle memorization, arguing that masking task descriptions fails if the model has memorized task-to-code implementations from GitHub.54:51–59:25 · The hosts as informed peer 4/10 Code Smells and Solution Quality in SWE-bench Audience members ask about generated code maintainability and AI code smells before Eugene concludes the session with announcements for upcoming paper reviews.0:04–2:06 · Guest teaching 2/10 Technical Setup and Screen Sharing Confirmation Jesse starts off the session introducing the purpose of SWE-bench compared to HumanEval. Eugene facilitates the tech setup without confrontation.2:07–4:17 · Guest teaching 4/10 Mining Pull Requests and Constructing SWE-bench Jesse gives an uninterrupted overview of how SWE-bench mines GitHub pull requests and test patches across popular Python repositories.4:20–8:22 · Guest teaching 4/10 Explaining Fail-to-Pass and Pass-to-Pass Unit Tests Eugene asks for clarification regarding how unit tests in pull requests are isolated and evaluated. Jesse breaks down the distinction between fail-to-pass and pass-to-pass tests.8:24–13:55 · Guest teaching 4/10 Deconstructing Benchmarks and Comparing OpenAI o1 Eugene relays audience questions comparing SWE-bench to OpenAI o1 and chain-of-thought methods. Jesse clarifies the distinction between agent scaffolds and raw model capabilities.13:58–17:05 · Guest teaching 5/10 Environment Setup Challenges and Cheating Vulnerabilities Jesse details the practical engineering pain points of reproducing environments via Docker and the vulnerabilities that allow agents to cheat by leaking future information.17:05–21:09 · Guest teaching 5/10 Introducing SWE-bench Verified and Solvability Criteria Jesse covers the motivation behind OpenAI SWE-bench Verified, explaining how human annotators evaluated issue solvability and filtered ambiguous tasks.21:10–23:55 · Guest teaching 3/10 Human Annotation Pipeline and Quality Filtering A participant chimes in with deep practical experience on human annotation calibration, non-neutral 4-point scales, and conservative ensembling methods.24:02–28:36 · Guest teaching 4/10 Human Demonstrations vs LLM Verification and CriticGPT The participant references OpenAI CriticGPT and context-length degradation curves, sparking a technical discussion on retrieval vs long-context LLMs.28:39–33:13 · Guest teaching 5/10 Analyzing Top Agent Strategies from Execution Trajectories Jesse shares practical insights derived from inspecting agent trajectories in the benchmark experiments repo, contrasting Agentless with Honeycomb and Guru.33:14–38:56 · Guest teaching 5/10 SWE-bench Multimodal for JavaScript and UI Tasks Jesse reviews the SWE-bench Multimodal paper, highlighting the shift to JavaScript/TypeScript and the strict requirement for visual assets in problem statements.38:57–42:53 · Guest teaching 2/10 Challenges in Objective UI and Web Agent Evaluation Eugene steps in with domain authority from running a UI testing startup, highlighting human unreliability and multi-screen PSD translation subjectivity.42:55–47:57 · Guest teaching 5/10 Introducing MLE-bench and Autonomous Kaggle Agents Jesse outlines MLE-bench by OpenAI, detailing how agents are evaluated on end-to-end Kaggle competitions with standard compute budgets and submission APIs.47:58–52:12 · Guest teaching 5/10 Practical Limitations, Costs, and Real-World ML Disconnect Jesse performs cost math on MLE-bench runs and critiques the gap between clean Kaggle competition definitions and messy real-world data science workflows.52:12–54:51 · Guest teaching 3/10 Evaluating Kaggle Data Contamination and Overfitting Eugene pushes back regarding Kaggle memorization, arguing that masking task descriptions fails if the model has memorized task-to-code implementations from GitHub.54:51–59:25 · Guest teaching 3/10 Code Smells and Solution Quality in SWE-bench Audience members ask about generated code maintainability and AI code smells before Eugene concludes the session with announcements for upcoming paper reviews.0:04–2:06 · Guest disagreement 0/10 Technical Setup and Screen Sharing Confirmation Jesse starts off the session introducing the purpose of SWE-bench compared to HumanEval. Eugene facilitates the tech setup without confrontation.2:07–4:17 · Guest disagreement 0/10 Mining Pull Requests and Constructing SWE-bench Jesse gives an uninterrupted overview of how SWE-bench mines GitHub pull requests and test patches across popular Python repositories.4:20–8:22 · Guest disagreement 0/10 Explaining Fail-to-Pass and Pass-to-Pass Unit Tests Eugene asks for clarification regarding how unit tests in pull requests are isolated and evaluated. Jesse breaks down the distinction between fail-to-pass and pass-to-pass tests.8:24–13:55 · Guest disagreement 1/10 Deconstructing Benchmarks and Comparing OpenAI o1 Eugene relays audience questions comparing SWE-bench to OpenAI o1 and chain-of-thought methods. Jesse clarifies the distinction between agent scaffolds and raw model capabilities.13:58–17:05 · Guest disagreement 0/10 Environment Setup Challenges and Cheating Vulnerabilities Jesse details the practical engineering pain points of reproducing environments via Docker and the vulnerabilities that allow agents to cheat by leaking future information.17:05–21:09 · Guest disagreement 0/10 Introducing SWE-bench Verified and Solvability Criteria Jesse covers the motivation behind OpenAI SWE-bench Verified, explaining how human annotators evaluated issue solvability and filtered ambiguous tasks.21:10–23:55 · Guest disagreement 1/10 Human Annotation Pipeline and Quality Filtering A participant chimes in with deep practical experience on human annotation calibration, non-neutral 4-point scales, and conservative ensembling methods.24:02–28:36 · Guest disagreement 0/10 Human Demonstrations vs LLM Verification and CriticGPT The participant references OpenAI CriticGPT and context-length degradation curves, sparking a technical discussion on retrieval vs long-context LLMs.28:39–33:13 · Guest disagreement 0/10 Analyzing Top Agent Strategies from Execution Trajectories Jesse shares practical insights derived from inspecting agent trajectories in the benchmark experiments repo, contrasting Agentless with Honeycomb and Guru.33:14–38:56 · Guest disagreement 0/10 SWE-bench Multimodal for JavaScript and UI Tasks Jesse reviews the SWE-bench Multimodal paper, highlighting the shift to JavaScript/TypeScript and the strict requirement for visual assets in problem statements.38:57–42:53 · Guest disagreement 1/10 Challenges in Objective UI and Web Agent Evaluation Eugene steps in with domain authority from running a UI testing startup, highlighting human unreliability and multi-screen PSD translation subjectivity.42:55–47:57 · Guest disagreement 0/10 Introducing MLE-bench and Autonomous Kaggle Agents Jesse outlines MLE-bench by OpenAI, detailing how agents are evaluated on end-to-end Kaggle competitions with standard compute budgets and submission APIs.47:58–52:12 · Guest disagreement 0/10 Practical Limitations, Costs, and Real-World ML Disconnect Jesse performs cost math on MLE-bench runs and critiques the gap between clean Kaggle competition definitions and messy real-world data science workflows.52:12–54:51 · Guest disagreement 2/10 Evaluating Kaggle Data Contamination and Overfitting Eugene pushes back regarding Kaggle memorization, arguing that masking task descriptions fails if the model has memorized task-to-code implementations from GitHub.54:51–59:25 · Guest disagreement 0/10 Code Smells and Solution Quality in SWE-bench Audience members ask about generated code maintainability and AI code smells before Eugene concludes the session with announcements for upcoming paper reviews.0:04–2:06 · The hosts pushing back 0/10 Technical Setup and Screen Sharing Confirmation Jesse starts off the session introducing the purpose of SWE-bench compared to HumanEval. Eugene facilitates the tech setup without confrontation.2:07–4:17 · The hosts pushing back 0/10 Mining Pull Requests and Constructing SWE-bench Jesse gives an uninterrupted overview of how SWE-bench mines GitHub pull requests and test patches across popular Python repositories.4:20–8:22 · The hosts pushing back 1/10 Explaining Fail-to-Pass and Pass-to-Pass Unit Tests Eugene asks for clarification regarding how unit tests in pull requests are isolated and evaluated. Jesse breaks down the distinction between fail-to-pass and pass-to-pass tests.8:24–13:55 · The hosts pushing back 3/10 Deconstructing Benchmarks and Comparing OpenAI o1 Eugene relays audience questions comparing SWE-bench to OpenAI o1 and chain-of-thought methods. Jesse clarifies the distinction between agent scaffolds and raw model capabilities.13:58–17:05 · The hosts pushing back 0/10 Environment Setup Challenges and Cheating Vulnerabilities Jesse details the practical engineering pain points of reproducing environments via Docker and the vulnerabilities that allow agents to cheat by leaking future information.17:05–21:09 · The hosts pushing back 0/10 Introducing SWE-bench Verified and Solvability Criteria Jesse covers the motivation behind OpenAI SWE-bench Verified, explaining how human annotators evaluated issue solvability and filtered ambiguous tasks.21:10–23:55 · The hosts pushing back 2/10 Human Annotation Pipeline and Quality Filtering A participant chimes in with deep practical experience on human annotation calibration, non-neutral 4-point scales, and conservative ensembling methods.24:02–28:36 · The hosts pushing back 2/10 Human Demonstrations vs LLM Verification and CriticGPT The participant references OpenAI CriticGPT and context-length degradation curves, sparking a technical discussion on retrieval vs long-context LLMs.28:39–33:13 · The hosts pushing back 0/10 Analyzing Top Agent Strategies from Execution Trajectories Jesse shares practical insights derived from inspecting agent trajectories in the benchmark experiments repo, contrasting Agentless with Honeycomb and Guru.33:14–38:56 · The hosts pushing back 0/10 SWE-bench Multimodal for JavaScript and UI Tasks Jesse reviews the SWE-bench Multimodal paper, highlighting the shift to JavaScript/TypeScript and the strict requirement for visual assets in problem statements.38:57–42:53 · The hosts pushing back 4/10 Challenges in Objective UI and Web Agent Evaluation Eugene steps in with domain authority from running a UI testing startup, highlighting human unreliability and multi-screen PSD translation subjectivity.42:55–47:57 · The hosts pushing back 0/10 Introducing MLE-bench and Autonomous Kaggle Agents Jesse outlines MLE-bench by OpenAI, detailing how agents are evaluated on end-to-end Kaggle competitions with standard compute budgets and submission APIs.47:58–52:12 · The hosts pushing back 0/10 Practical Limitations, Costs, and Real-World ML Disconnect Jesse performs cost math on MLE-bench runs and critiques the gap between clean Kaggle competition definitions and messy real-world data science workflows.52:12–54:51 · The hosts pushing back 5/10 Evaluating Kaggle Data Contamination and Overfitting Eugene pushes back regarding Kaggle memorization, arguing that masking task descriptions fails if the model has memorized task-to-code implementations from GitHub.54:51–59:25 · The hosts pushing back 1/10 Code Smells and Solution Quality in SWE-bench Audience members ask about generated code maintainability and AI code smells before Eugene concludes the session with announcements for upcoming paper reviews.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 52:58 Disputing Kaggle obfuscation robustness

Eugene challenges Jesse's point about task obfuscation, emphasizing that memorized code on GitHub nullifies text-masking defenses.

Hardest push from the hosts ▶ 52:58 Pressing on training data contamination

Eugene directly reframes the overfitting problem from task comprehension to underlying code memorization across open-source repositories.

Biggest teaching moment ▶ 21:10 Masterclass on human annotator calibration

A participant provides deep domain knowledge on inter-rater reliability, forced-choice rating scales, and high-severity ensembling.

The host holds their own ▶ 39:40 Firsthand UI testing founder experience

Eugene draws on his background founding a UI testing company to explain why subjective visual testing fails across varied screen resolutions.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Technical Setup and Screen Sharing Confirmation 1200 Jesse starts off the session introducing the purpose of SWE-bench compared to HumanEval. Eugene facilitates the tech setup without confrontation.
Mining Pull Requests and Constructing SWE-bench 0400 Jesse gives an uninterrupted overview of how SWE-bench mines GitHub pull requests and test patches across popular Python repositories.
Explaining Fail-to-Pass and Pass-to-Pass Unit Tests 3401 Eugene asks for clarification regarding how unit tests in pull requests are isolated and evaluated. Jesse breaks down the distinction between fail-to-pass and pass-to-pass tests.
Deconstructing Benchmarks and Comparing OpenAI o1 4413 Eugene relays audience questions comparing SWE-bench to OpenAI o1 and chain-of-thought methods. Jesse clarifies the distinction between agent scaffolds and raw model capabilities.
Environment Setup Challenges and Cheating Vulnerabilities 0500 Jesse details the practical engineering pain points of reproducing environments via Docker and the vulnerabilities that allow agents to cheat by leaking future information.
Introducing SWE-bench Verified and Solvability Criteria 0500 Jesse covers the motivation behind OpenAI SWE-bench Verified, explaining how human annotators evaluated issue solvability and filtered ambiguous tasks.
Human Annotation Pipeline and Quality Filtering 7312 A participant chimes in with deep practical experience on human annotation calibration, non-neutral 4-point scales, and conservative ensembling methods.
Human Demonstrations vs LLM Verification and CriticGPT 6402 The participant references OpenAI CriticGPT and context-length degradation curves, sparking a technical discussion on retrieval vs long-context LLMs.
Analyzing Top Agent Strategies from Execution Trajectories 2500 Jesse shares practical insights derived from inspecting agent trajectories in the benchmark experiments repo, contrasting Agentless with Honeycomb and Guru.
SWE-bench Multimodal for JavaScript and UI Tasks 0500 Jesse reviews the SWE-bench Multimodal paper, highlighting the shift to JavaScript/TypeScript and the strict requirement for visual assets in problem statements.
Challenges in Objective UI and Web Agent Evaluation 7214 Eugene steps in with domain authority from running a UI testing startup, highlighting human unreliability and multi-screen PSD translation subjectivity.
Introducing MLE-bench and Autonomous Kaggle Agents 0500 Jesse outlines MLE-bench by OpenAI, detailing how agents are evaluated on end-to-end Kaggle competitions with standard compute budgets and submission APIs.
Practical Limitations, Costs, and Real-World ML Disconnect 1500 Jesse performs cost math on MLE-bench runs and critiques the gap between clean Kaggle competition definitions and messy real-world data science workflows.
Evaluating Kaggle Data Contamination and Overfitting 6325 Eugene pushes back regarding Kaggle memorization, arguing that masking task descriptions fails if the model has memorized task-to-code implementations from GitHub.
Code Smells and Solution Quality in SWE-bench 4301 Audience members ask about generated code maintainability and AI code smells before Eugene concludes the session with announcements for upcoming paper reviews.

Statements from this episode (21)

Assertion Supported
Jesse Hu: Frontier models have effectively capped out HumanEval performance
“I think through multiple techniques and through the most recent models, you can actually basically cap out on the performance on human eval.”
Jesse Hu Oct 19, 2024 ▶ 1:21
Opinion
Hu: SWE-bench repos do not represent typical app development work
“And I'll say the critique here is that these are, like, really robust you know, Repos that are, have a lot of history, have a lot of people, but they don't really represent like what it is to do day-to-day work for us that are building like apps and client thi…”
Jesse Hu Oct 19, 2024 ▶ 6:36
Assertion Supported
Hu: SWE-bench scores jumped from 3% on GPT-4 RAG to 43%
“If we just use GPT-IV plus RAG, what do we get? It's, like, a measly three percent. And then up to the most recent submissions where they get up to 43%.”
Jesse Hu Oct 19, 2024 ▶ 7:40
Insight
Jesse Hu: SWE-bench gains over baseline GPT-4 come entirely from agent scaffolding
“The diff between that and something like Devin is all in like, sort of like the agent scaffold or the agent code, right? So that's, what's really exciting about this stuff. It shows off what you can do just from prompting and just from adding tools.”
Jesse Hu Oct 19, 2024 ▶ 12:09
Assertion Not checkable as stated
Jesse Hu: SWE-bench competitors use identical base model APIs, differentiating via prompting
“Everyone is using the same base models here, or the same model APIs period. I don't know if anyone except for cosine was even training, but like the diff here is all in prompting.”
Jesse Hu Oct 19, 2024 ▶ 13:39
Assertion Supported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Jesse Hu Oct 19, 2024 ▶ 15:00
Insight
Jesse Hu: SWE-Bench trajectory requirements deter commercial agents from submitting
“They actually require you to submit these things called trajectories that are proving what you did and that you're not cheating. And a lot of this ends in controversy. A lot of this ends in people say, well, we don't want to submit at all because if we review …”
Jesse Hu Oct 19, 2024 ▶ 15:45
Opinion
Hu: Some SWE-bench Tasks Are Unsolvable Due to Missing Context
“Sometimes it feels like you try to read the issue. And you're just like, okay, even if I was like an Oracle or like some sort of God coder, I couldn't, I would not be able to solve this because there's some context that was left out here. There's something tha…”
Jesse Hu Oct 19, 2024 ▶ 18:47
Prediction Open · timeframe Oct 2029
Hu: AI Models Should Eventually Reach 100% on SWE-bench Verified
“And in that way, I think we should be able to hit up a hundred percent eventually.”
Jesse Hu Oct 19, 2024 ▶ 20:49
Opinion
Hu: Long-context accuracy degrades; RAG remains necessary for entire large codebases
“My guess would be that, like, long context works, but it's sort of a lie as far as your accuracy, and that rag matters no matter what, because even in the longest context windows, you can't fit the whole code base.”
Jesse Hu Oct 19, 2024 ▶ 27:29
Assertion Supported
Hu: Honeycomb agent framework ranks number one on full SWE-bench dataset
“And then there's another one called Honeycomb, which also scored really well, and I think is number one on the full set.”
Jesse Hu Oct 19, 2024 ▶ 31:30
Assertion Supported
SWE-bench Multimodal paper baseline scores 12 percent
“They want to show off that this is, you know, guys, this is really hard. It's a really hard benchmark, so we can only get 12% on it today. According to, you know, their implementation.”
Jesse Hu Oct 19, 2024 ▶ 34:48
Opinion
Hu: Evaluating UI correctness in SWE-bench Multimodal is highly subjective
“Zooming out, if you're true to try to judge whether a UI is correct, it's like extremely subjective. It might be iterative. It might be, you know, I have to interact with it first to get it right. So I'm super curious to see how they actually do the judging cr…”
Jesse Hu Oct 19, 2024 ▶ 36:56
Insight
Jesse Hu: Realistic AI evaluation requires measuring multi-turn clarification, not single-turn fixes
“I think the broader thing is that the more and more realistic you get, the more you run into sort of like multi-turn or iterative things where now the task of the AI isn't just to just, you know, extract from your brain what the problem is and kind of directly…”
Jesse Hu Oct 19, 2024 ▶ 40:53
Assertion Supported
Hu: OpenAI o1-preview achieves bronze medals in 17% of MLE-bench competitions
“Their final results with a one preview and this a scaffolding from a different company was that they got a bronze medal. I don't think I've ever achieved once but I haven't competed that much in. 17% of competitions.”
Jesse Hu Oct 19, 2024 ▶ 45:06
Assertion Supported
Hu: OpenAI o1-preview surpasses human Kaggle Grandmasters with seven gold medals
“Since a grandmaster requires five gold medals and oh, and preview gets an average of eight or sorry, seven gold medals. They're out competing even capital grandmasters.”
Jesse Hu Oct 19, 2024 ▶ 47:29
Insight
Hu: AI agents lack intrinsic time awareness, failing to budget execution limits
“What's interesting is like, and I, I've seen this in practice, it's like, it's hard to get the agent to say, to think in numbers of steps, and especially in time, because it doesn't know time. So if you tell it like, please complete under 50 steps, it won't do…”
Jesse Hu Oct 19, 2024 ▶ 48:45
Assertion Supported
Hu: GPU setups showed virtually no agent performance gain over CPU-only
“They compared a CPU only setup to a GPU setup to a multi GPU setup, and it kind of made no difference really.”
Jesse Hu Oct 19, 2024 ▶ 49:40
Assertion Supported
Hu: Single MLE-bench evaluation run with OpenAI o1-preview costs $4,000
“Just for one seed, For one run of these things cost 4000 dollars all in with the GPU plus the tokens. And a bulk of the cost was actually the token, so even if you cut the GPU out, it'll still cost you three grand to run on one preview.”
Jesse Hu Oct 19, 2024 ▶ 50:38
Assertion Supported
Hu: MLE-bench authors found obfuscating competition details did not show overfitting
“They do a lot of checks against overfitting on the Kaggle tasks themselves, and so they do something where they obfuscate some of the details of the Of the competitions, and then they rerun it. And I guess if they were overfitting on the competitions themselve…”
Jesse Hu Oct 19, 2024 ▶ 52:19
Insight
Hu: Coding agents produce bloated edits unless constrained by brevity priors
“There's something nuanced about this data set in particular where all the edits are super short and it's like a prior that you can put into your code. But if you don't have that, then yeah, like agents tend to just keep making obnoxiously long edits”
Jesse Hu Oct 19, 2024 ▶ 57:49
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.