Ali Pullen is the CEO of Cosine, creator of the Genie AI software engineering agent. He discusses why training models purely on raw git diffs is insufficient for real-world software engineering benchmarks.
“What the, very crudely, what the pre-trained models are reading is they're reading those final diffs and they're Emulating that and then being able to output it. Right. But of course it's a super lossy thing, a PR. You have no idea why or how, for the most part, unless there are some comments, which, you know, anyone who's worked in a company realizes PR reviews can be a bit dodgy at times, but you see that you lose so much information at the end. And that's perfectly fine because PRs aren't designed to be something that perfectly preserves everything that happened. But what we realized was if you want something that's A software engineer, and very crudely, we started with like something that can do PRs for you, essentially. You need to be able to figure out why those things happened. Otherwise, you're just going to rely, essentially, you just have a code writing model. You have something that's good at human eval, but not very good at SweetBenge, essentially.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Alistair Pullen
AssertionSupported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Alistair PullenOct 4, 2024▶ 1:13:04Building AGI in Real Time (OpenAI Dev Day 2024)
Opinion
Pullen: SWE-bench is a poor proxy for real-world AI coding competence
“I know Sweebench is, like, the most commonly talked about thing, and honestly, it's a very, it's an amazing project, but one of the things we've learned the most from actually shipping this product to users is, it's a pretty bad proxy at telling us how compete…”
Alistair PullenOct 4, 2024▶ 1:16:10Building AGI in Real Time (OpenAI Dev Day 2024)
Insight
Extracting problem-solving history most determines Cosine Genie's SWE-bench performance
“The way that we Decided we want to try to extract what happened in the past, like as forensically as possible, has been and is currently like one of the main things that we focus all our time on. Because doing that as we're getting as much signal out as possib…”
Cosine refuses to publish SWE-bench trajectories to prevent competitor model distillation
“For the moment, as a closed source company, like fighting for an edge, we've decided not to publish that information for that exact reason. I don't want someone basically taking my tragedies and then taking a model that's suing them in GA and just distilling i…”
Pullen: You cannot build a successful startup solely on the YC advantage
“You can't build a startup on the YC advantage. It's obviously nice and it makes you feel warm and fuzzy inside, but like at the end of the day, it's not that that's going to make you win.”
Pullen: SWE-bench is currently the best metric for software engineering agents
“It was actually a very useful tool in building Genie because beforehand it was like, yes, vibe check this thing and see if it's useful. And then all of a sudden you have an actual measure to see like, couldn't it do software engineering? Not the best measure, …”
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.