People, every show

Jesse Hu

Co-Founder, Abundant. On 1 show, 1 appearance. The Shows tab opens the full record on each.

founderengineerscientist@huyouare ↗LinkedIn ↗abundant.ai ↗

Jesse Hu is the co-founder of Abundant, an AI infrastructure startup building reinforcement learning environments and evaluation platforms for autonomous coding agents. He previously developed machine learning planning systems for autonomous vehicles at Waymo and worked on large-scale recommendation systems at YouTube and Google.

1shows
1appearances
21statements
10resolved
10supported
0contradicted
100%fully supported

Everything Jesse Hu said on any show that made the record, most notable first. Each card names its show and opens the statement there.

Jesse Hu: SWE-bench gains over baseline GPT-4 come entirely from agent scaffolding
“The diff between that and something like Devin is all in like, sort of like the agent scaffold or the agent code, right? So that's, what's really exciting about this stuff. It shows off what you can do just from prompting and just from adding tools.”
Jesse Hu Oct 19, 2024 ▶ 12:09 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
SWE-Bench public test splits enable trivial cheating via runtime PR retrieval
“The entire test split here is public. So you can do things like just overfit to the patches in the test set. You can do things like, let me add at runtime, pull the PR and just get the answer and just use it.”
Jesse Hu Oct 19, 2024 ▶ 15:00 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Jesse Hu: SWE-Bench trajectory requirements deter commercial agents from submitting
“They actually require you to submit these things called trajectories that are proving what you did and that you're not cheating. And a lot of this ends in controversy. A lot of this ends in people say, well, we don't want to submit at all because if we review …”
Jesse Hu Oct 19, 2024 ▶ 15:45 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Hu: SWE-bench repos do not represent typical app development work
“And I'll say the critique here is that these are, like, really robust you know, Repos that are, have a lot of history, have a lot of people, but they don't really represent like what it is to do day-to-day work for us that are building like apps and client thi…”
Jesse Hu Oct 19, 2024 ▶ 6:36 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Hu: Long-context accuracy degrades; RAG remains necessary for entire large codebases
“My guess would be that, like, long context works, but it's sort of a lie as far as your accuracy, and that rag matters no matter what, because even in the longest context windows, you can't fit the whole code base.”
Jesse Hu Oct 19, 2024 ▶ 27:29 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
Hu: OpenAI o1-preview surpasses human Kaggle Grandmasters with seven gold medals
“Since a grandmaster requires five gold medals and oh, and preview gets an average of eight or sorry, seven gold medals. They're out competing even capital grandmasters.”
Jesse Hu Oct 19, 2024 ▶ 47:29 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
Jesse Hu: Frontier models have effectively capped out HumanEval performance
“I think through multiple techniques and through the most recent models, you can actually basically cap out on the performance on human eval.”
Jesse Hu Oct 19, 2024 ▶ 1:21 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Not checkable as stated
Jesse Hu: SWE-bench competitors use identical base model APIs, differentiating via prompting
“Everyone is using the same base models here, or the same model APIs period. I don't know if anyone except for cosine was even training, but like the diff here is all in prompting.”
Jesse Hu Oct 19, 2024 ▶ 13:39 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Hu: Some SWE-bench Tasks Are Unsolvable Due to Missing Context
“Sometimes it feels like you try to read the issue. And you're just like, okay, even if I was like an Oracle or like some sort of God coder, I couldn't, I would not be able to solve this because there's some context that was left out here. There's something tha…”
Jesse Hu Oct 19, 2024 ▶ 18:47 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Prediction Open · timeframe Oct 2029
Hu: AI Models Should Eventually Reach 100% on SWE-bench Verified
“And in that way, I think we should be able to hit up a hundred percent eventually.”
Jesse Hu Oct 19, 2024 ▶ 20:49 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Hu: Evaluating UI correctness in SWE-bench Multimodal is highly subjective
“Zooming out, if you're true to try to judge whether a UI is correct, it's like extremely subjective. It might be iterative. It might be, you know, I have to interact with it first to get it right. So I'm super curious to see how they actually do the judging cr…”
Jesse Hu Oct 19, 2024 ▶ 36:56 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Jesse Hu: Realistic AI evaluation requires measuring multi-turn clarification, not single-turn fixes
“I think the broader thing is that the more and more realistic you get, the more you run into sort of like multi-turn or iterative things where now the task of the AI isn't just to just, you know, extract from your brain what the problem is and kind of directly…”
Jesse Hu Oct 19, 2024 ▶ 40:53 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Hu: AI agents lack intrinsic time awareness, failing to budget execution limits
“What's interesting is like, and I, I've seen this in practice, it's like, it's hard to get the agent to say, to think in numbers of steps, and especially in time, because it doesn't know time. So if you tell it like, please complete under 50 steps, it won't do…”
Jesse Hu Oct 19, 2024 ▶ 48:45 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
Hu: GPU setups showed virtually no agent performance gain over CPU-only
“They compared a CPU only setup to a GPU setup to a multi GPU setup, and it kind of made no difference really.”
Jesse Hu Oct 19, 2024 ▶ 49:40 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
Hu: Single MLE-bench evaluation run with OpenAI o1-preview costs $4,000
“Just for one seed, For one run of these things cost 4000 dollars all in with the GPU plus the tokens. And a bulk of the cost was actually the token, so even if you cut the GPU out, it'll still cost you three grand to run on one preview.”
Jesse Hu Oct 19, 2024 ▶ 50:38 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
Hu: MLE-bench authors found obfuscating competition details did not show overfitting
“They do a lot of checks against overfitting on the Kaggle tasks themselves, and so they do something where they obfuscate some of the details of the Of the competitions, and then they rerun it. And I guess if they were overfitting on the competitions themselve…”
Jesse Hu Oct 19, 2024 ▶ 52:19 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Hu: Coding agents produce bloated edits unless constrained by brevity priors
“There's something nuanced about this data set in particular where all the edits are super short and it's like a prior that you can put into your code. But if you don't have that, then yeah, like agents tend to just keep making obnoxiously long edits”
Jesse Hu Oct 19, 2024 ▶ 57:49 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
Hu: SWE-bench scores jumped from 3% on GPT-4 RAG to 43%
“If we just use GPT-IV plus RAG, what do we get? It's, like, a measly three percent. And then up to the most recent submissions where they get up to 43%.”
Jesse Hu Oct 19, 2024 ▶ 7:40 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
Hu: Honeycomb agent framework ranks number one on full SWE-bench dataset
“And then there's another one called Honeycomb, which also scored really well, and I think is number one on the full set.”
Jesse Hu Oct 19, 2024 ▶ 31:30 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
SWE-bench Multimodal paper baseline scores 12 percent
“They want to show off that this is, you know, guys, this is really hard. It's a really hard benchmark, so we can only get 12% on it today. According to, you know, their implementation.”
Jesse Hu Oct 19, 2024 ▶ 34:48 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
LATENT SPACE Assertion Supported
Hu: OpenAI o1-preview achieves bronze medals in 17% of MLE-bench competitions
“Their final results with a one preview and this a scaffolding from a different company was that they got a bronze medal. I don't think I've ever achieved once but I haven't competed that much in. 17% of competitions.”
Jesse Hu Oct 19, 2024 ▶ 45:06 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu

One line per show, most statements first. The link opens Jesse's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
LATENT SPACELEDGER Co-Founder, Abundant 1 21 100% 10/10 full record on Latent Space →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.