Aug 22, 2024 · 1h 1m · latent-space
Is finetuning GPT4o worth it?
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of Latent Space, Cosine CEO Ali Pullen joins Alessio Fanelli and Swix to break down the technical architecture, synthetic training pipelines, and large-scale OpenAI fine-tuning behind Genie, their state-of-the-art autonomous AI software engineer. The discussion explores Cosine's founding journey, empirical discoveries in context window degradation, and benchmark leadership on SWE-bench.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.7% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Ali firmly rejects Swyx's assertion that he deemed browser tools unimportant, clarifying that he was attacking superficial wrapper architectures rather than the utility of browsing.
Hardest push from the hosts ▶ 22:07 Challenging Genie's dismissal of agent toolingSwyx directly challenges Ali by quoting his claim that competitors are mere wrappers and questioning why Genie downplays tools like browsers.
Biggest teaching moment ▶ 36:15 Context window degradation thresholdAli delivers proprietary empirical data showing that SWE-bench problem resolution degrades linearly, falling to a 50% failure rate past 60k tokens.
The host holds their own ▶ 47:51 Synthesizing Llama 3 back-translation methodsSwyx demonstrates deep domain awareness by connecting Ali's synthetic data methodology to the Llama 3 paper's back-translation findings.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Ali Pullen's Early Background and the Fancy Acquisition | 4 | 3 | 1 | 0 | Swyx introduces Ali and summarizes his resume from Exeter to Fancy/GoPuff. Ali recounts his transition into startups and the acquisition in a narrative, collaborative tone. | |
| Discovering GPT-3 and Early App Generation Prototypes | 5 | 3 | 1 | 0 | Ali and the hosts bond over early GPT-3 playground completions and early Codex prototypes. The dynamic is conversational and nostalgic with shared domain context. | |
| The Y Combinator Experience and the Genesis of Build | 5 | 3 | 1 | 1 | Ali details the brutal YC interview and the early pivot toward codebase retrieval. Swyx teases him about the poor performance and unpronounceable name of 'Build'. | |
| Expanding Context Windows and the Thesis for Fine-Tuning | 5 | 5 | 2 | 1 | Ali explains why 128k context windows unlocked viable software engineering agents via fine-tuning. Swyx observes how OpenAI's fast-tracking of fine-tuning diminishes the YC Slack channel moat. | |
| Partnering with OpenAI and Adopting SWE-bench | 5 | 6 | 1 | 0 | Ali details getting early experimental access to GPT-4 Turbo fine-tuning and diving headfirst into SWE-bench after Devin launched. The hosts listen attentively as Ali breaks down early training setbacks. | |
| Distinguishing Code Generation from True Software Engineering | 6 | 7 | 2 | 0 | Alessio asks how Genie differentiates code generation from software engineering. Ali explains that PR diffs are lossy artifacts, educating the hosts on reconstructing the developer's decision trajectory. | |
| Data Cleansing and Customer Appetite for Enterprise Code Sharing | 6 | 5 | 1 | 0 | Alessio queries customer willingness to share proprietary codebases for model training. Ali explains enterprise appetite for productivity outweighs security resistance in practice. | |
| Genie's Core Architecture vs. Generic Agent Tooling | 6 | 5 | 4 | 3 | Swyx challenges Ali on claiming browser and code interpreter tools are unimportant compared to Genie's core loop. Ali corrects the premise, asserting he derided generic wrapper architecture, not the tools themselves. | |
| Advancing Codebase Retrieval through Self-Play and Language Servers | 6 | 7 | 2 | 1 | Ali explains why standard semantic embeddings fail for code functionality and how training Genie on LSP traversal and self-play boosted retrieval accuracy to 66%. Swyx follows with technical suggestions. | |
| Foundation Model Agnosticism vs. Custom Architectures | 6 | 5 | 2 | 1 | Swyx brings up Magic.dev's LTM approach with long contexts. Ali defends remaining model-agnostic, arguing that fine-tuning data can easily transfer across foundation models like Gemini or Anthropic. | |
| Context Window Degradation and Token Log Probabilities | 6 | 8 | 1 | 0 | Ali shares proprietary insights shared with OpenAI showing that task success drops to 50% past 60k tokens and breaks down token log probabilities. Alessio asks about context scaling factors. | |
| Multi-Language Data Distribution and Model Generalization | 6 | 5 | 1 | 0 | Alessio examines multi-language distribution and enterprise test execution. Ali clarifies that Genie offloads test runs to existing GitHub Actions and CI pipelines rather than managing local repo environments. | |
| Deep Collaboration with OpenAI on LoRA and Adapter Scaling | 5 | 7 | 1 | 0 | Ali describes deep technical collaboration with OpenAI, discussing LoRA adapter capacity, learning rate schedules, and scaling laws when pushing billions of tokens through fine-tuning APIs. | |
| Generating Synthetic Data and Modeling Iterative Error Correction | 7 | 6 | 1 | 0 | Ali explains synthetically corrupting ASTs to teach models iterative error correction. Swyx connects this methodology to Llama 3's back-translation findings in code synthesis. | |
| SWE-bench Trajectory Controversies and IP Protection | 6 | 6 | 3 | 1 | Swyx asks why Genie's score isn't on the official SWE-bench leaderboard. Ali explains his refusal to submit full model trajectories, prioritizing IP protection against distillation over academic listing. | |
| SWE-bench Verified Results and Granular Evaluation Metrics | 5 | 6 | 1 | 0 | Ali discusses achieving 43.8% on SWE-bench Verified and outlines plans to move beyond binary pass/fail evals toward step-level divergence metrics. | |
| Future Horizons: Scaling Data and Repo-Specific Fine-Tuning | 5 | 4 | 1 | 0 | Swyx and Ali conclude by discussing fine-tuning models on specific customer repositories, Cursor's UX excellence, hiring needs, and Ali's viral launch video bloopers. |