Disclosure certainty 4/5 debate potential 1/5

Cosine injects synthetic AST errors into training data to teach error recovery

Alistair Pullen · Is finetuning GPT4o worth it? · Aug 22, 2024 · at 46:15

Ali Pullen is the CEO of Cosine, creators of Genie, an autonomous AI software engineering agent. He discusses creating training data for debugging non-working code.

0:00 / 0:24exact quote · 24.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“And that was in sort of two parts. We synthetically generated runtime errors where we would Intentionally mess with the AST to make stuff not work or index out of bounds or refer to a variable that doesn't exist or errors that the foundational models just make sometimes that you can't really avoid. You can't expect it to be perfect. So we threw some of those in with a probability of happening.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Alistair Pullen

Assertion Supported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Alistair Pullen Oct 4, 2024 ▶ 1:13:04 Building AGI in Real Time (OpenAI Dev Day 2024)
Opinion
Pullen: SWE-bench is a poor proxy for real-world AI coding competence
“I know Sweebench is, like, the most commonly talked about thing, and honestly, it's a very, it's an amazing project, but one of the things we've learned the most from actually shipping this product to users is, it's a pretty bad proxy at telling us how compete…”
Alistair Pullen Oct 4, 2024 ▶ 1:16:10 Building AGI in Real Time (OpenAI Dev Day 2024)
Insight
Extracting problem-solving history most determines Cosine Genie's SWE-bench performance
“The way that we Decided we want to try to extract what happened in the past, like as forensically as possible, has been and is currently like one of the main things that we focus all our time on. Because doing that as we're getting as much signal out as possib…”
Alistair Pullen Aug 22, 2024 ▶ 18:10 Is finetuning GPT4o worth it?
Disclosure
Cosine refuses to publish SWE-bench trajectories to prevent competitor model distillation
“For the moment, as a closed source company, like fighting for an edge, we've decided not to publish that information for that exact reason. I don't want someone basically taking my tragedies and then taking a model that's suing them in GA and just distilling i…”
Alistair Pullen Aug 22, 2024 ▶ 50:35 Is finetuning GPT4o worth it?
Insight
Pullen: You cannot build a successful startup solely on the YC advantage
“You can't build a startup on the YC advantage. It's obviously nice and it makes you feel warm and fuzzy inside, but like at the end of the day, it's not that that's going to make you win.”
Alistair Pullen Aug 22, 2024 ▶ 13:23 Is finetuning GPT4o worth it?
Opinion
Pullen: SWE-bench is currently the best metric for software engineering agents
“It was actually a very useful tool in building Genie because beforehand it was like, yes, vibe check this thing and see if it's useful. And then all of a sudden you have an actual measure to see like, couldn't it do software engineering? Not the best measure, …”
Alistair Pullen Aug 22, 2024 ▶ 15:39 Is finetuning GPT4o worth it?
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.