“I think the broader thing is that the more and more realistic you get, the more you run into sort of like multi-turn or iterative things where now the task of the AI isn't just to just, you know, extract from your brain what the problem is and kind of directly solve it. The task is like, how can it ask you clarifying questions and extract more knowledge from your brain and iterate just really, really, really hard to evaluate.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Jesse Hu
Insight
Jesse Hu: SWE-bench gains over baseline GPT-4 come entirely from agent scaffolding
“The diff between that and something like Devin is all in like, sort of like the agent scaffold or the agent code, right? So that's, what's really exciting about this stuff. It shows off what you can do just from prompting and just from adding tools.”
Jesse HuOct 19, 2024▶ 12:09[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
AssertionSupported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Jesse HuOct 19, 2024▶ 15:00[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Insight
Jesse Hu: SWE-Bench trajectory requirements deter commercial agents from submitting
“They actually require you to submit these things called trajectories that are proving what you did and that you're not cheating. And a lot of this ends in controversy. A lot of this ends in people say, well, we don't want to submit at all because if we review …”
Jesse HuOct 19, 2024▶ 15:45[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Opinion
Hu: SWE-bench repos do not represent typical app development work
“And I'll say the critique here is that these are, like, really robust you know, Repos that are, have a lot of history, have a lot of people, but they don't really represent like what it is to do day-to-day work for us that are building like apps and client thi…”
Jesse HuOct 19, 2024▶ 6:36[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Opinion
Hu: Long-context accuracy degrades; RAG remains necessary for entire large codebases
“My guess would be that, like, long context works, but it's sort of a lie as far as your accuracy, and that rag matters no matter what, because even in the longest context windows, you can't fit the whole code base.”
Jesse HuOct 19, 2024▶ 27:29[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
AssertionSupported
Hu: OpenAI o1-preview surpasses human Kaggle Grandmasters with seven gold medals
“Since a grandmaster requires five gold medals and oh, and preview gets an average of eight or sorry, seven gold medals. They're out competing even capital grandmasters.”
Jesse HuOct 19, 2024▶ 47:29[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.