Genie

product on 10 shows · 24 statements across 6 episodes · said 84 times in 20 episodes since 2024

Latent Space 61 the a16z Podcast 9 All-In 5 Invest Like the Best 2 TBPN 2 20VC 2 the Y Combinator Startup Podcast 1 Lenny's Podcast 1 Big Technology 1 Acquired

Mentions by year, every show

tap a year for its mentions
00304608202420252026episodesmentions
048202420252026episodes it came up in
007.54158202420252026episodesmentions per episode

Latent Space 61the a16z Podcast 9All-In 520VC 2Invest Like the Best 2TBPN 2Big Technology 1Lenny's Podcast 11 more shows

2026 9 mentions in 7 episodes 1 per episode
2025 23 mentions in 8 episodes 3 per episode
2024 52 mentions in 5 episodes 10 per episode

every mention on every show, scene by scene, with the transcript →

24 statements about Genie, every show

BIG TECHNOLOGY Disclosure
Manyika: Google DeepMind develops specialized AI models for Waymo's autonomous driver
“And then you've got all these other, you know, special kind of ambitious projects working on things like Genie build world models. You've got work going on to build special things for Waymo and enhance the models that Waymo's that lead to Waymo, the Waymo driv…”
James Manyika Feb 18, 2026 ▶ 9:46 How Google DeepMind Operates & Experiments — With Lila Ibrahim and James Manyika
ALL-IN Assertion Supported
Hassabis: DeepMind's Genie was trained on video and game engine synthetic data
“It was trained off a video and some synthetic data from game engines.”
Demis Hassabis Sep 12, 2025 ▶ 7:14 Inside Google DeepMind: AGI, Robotics, & World Models Explained - Demis Hassabis
a16z Assertion Supported
Fruchter: Veo cannot currently navigate environments or take actions like Genie
“Genie allows you to navigate environment and then maybe take actions. And that's not something that Veo at this point can do.”
Shlomi Fruchter Aug 16, 2025 ▶ 21:05 Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building
LATENT SPACE Disclosure
Pullen: Replacing Genie's reasoning traces with o1 traces improves performance
“Even now we've started Replacing some of the reasoning traces in our Genie model with reasoning traces generated by O-one, or at least in tandem with O-one, and we've already started seeing improvements in performance from that point.”
Alistair Pullen Oct 4, 2024 ▶ 1:09:59 Building AGI in Real Time (OpenAI Dev Day 2024)
LATENT SPACE Assertion Supported
Pullen: CoScene's Genie outperforms OpenAI o1 out of the box on SWE-bench
“So it was obviously great to see, like, we still are better than O-one out of the box. You know, even with an older model, and I'm sure that that, that Delta will continue to grow once we're able to train O-one and once we've done more work on our dataset usin…”
Alistair Pullen Oct 4, 2024 ▶ 1:13:04 Building AGI in Real Time (OpenAI Dev Day 2024)
Autonomous coding agents require grounding all generated code in context windows
“Fundamentally to build a product like this, you need to get as much information in front of the model as possible and make sure that everything ever writes in output can be Traced back to something in the context window, so it's not hallucinating it.”
Alistair Pullen Aug 22, 2024 ▶ 11:32 Is finetuning GPT4o worth it?
LATENT SPACE Disclosure
Access to GPT-4 Turbo experimental fine-tuning enabled Cosine Genie's creation
“Eventually we were able to get on the experimental access program and we got access to four turbo fine tuning. As soon as we did that, because in the entire run up to that, we'd built the data pipeline. We already had all that set up. So we're like, right, we …”
Alistair Pullen Aug 22, 2024 ▶ 14:14 Is finetuning GPT4o worth it?
Pullen: SWE-bench is currently the best metric for software engineering agents
“It was actually a very useful tool in building Genie because beforehand it was like, yes, vibe check this thing and see if it's useful. And then all of a sudden you have an actual measure to see like, couldn't it do software engineering? Not the best measure, …”
Alistair Pullen Aug 22, 2024 ▶ 15:39 Is finetuning GPT4o worth it?
Extracting problem-solving history most determines Cosine Genie's SWE-bench performance
“The way that we Decided we want to try to extract what happened in the past, like as forensically as possible, has been and is currently like one of the main things that we focus all our time on. Because doing that as we're getting as much signal out as possib…”
Alistair Pullen Aug 22, 2024 ▶ 18:10 Is finetuning GPT4o worth it?
File modification is more fundamental for coding agents than browser access
“At least with what we've seen, the browser is helpful, but it's not as helpful as like writing the correct files. If that makes sense. Like, it is still helpful, but obviously there are more fundamental things you have to get right before you get to like, oh y…”
Alistair Pullen Aug 22, 2024 ▶ 22:07 Is finetuning GPT4o worth it?
LATENT SPACE Disclosure
Cosine's Genie uses the Perplexity API and URL reading tools
“The genie has both of those tools available to it as well. So yeah, yeah. So we have a tool where you can like put in URLs and it will just read the URLs and you can, it also uses perplexities API under the hood as well to be able to actually ask questions if …”
Alistair Pullen Aug 22, 2024 ▶ 23:02 Is finetuning GPT4o worth it?
LATENT SPACE Assertion Supported
Cosine's Genie achieved roughly 66% codebase retrieval accuracy across benchmark tasks
“And I think in our technical report, I can't remember the exact number, but I think it was around 65 or 66% retrieval accuracy overall measured on. We know what lines we need for these tasks to find for the task to actually be able to be completed. And we foun…”
Alistair Pullen Aug 22, 2024 ▶ 27:31 Is finetuning GPT4o worth it?
LATENT SPACE Prediction Open · timeframe Aug 2027
Cosine will fine-tune and run Genie on Gemini 1.5 once supported
“As soon as you can fine-tune Gemini 1.5, then you best be sure that Genie will have, will work, will run on Gemini 1.5, and, like, we'll probably get very good performance out of that.”
Alistair Pullen Aug 22, 2024 ▶ 30:41 Is finetuning GPT4o worth it?
LATENT SPACE Prediction Not checkable as stated
Upgraded frontier models will automatically improve Cosine's synthetic data flywheel
“When models like that come out, obviously the signal in my data, when I regenerate it goes up. And then I can then train that model that's already better at reasoning with improved reasoning data. And just like, I can keep bootstrapping and keep leapfrogging e…”
Alistair Pullen Aug 22, 2024 ▶ 32:20 Is finetuning GPT4o worth it?
LATENT SPACE Disclosure
Cosine will make all Genie plans and generated code patches editable
“So we're going to make everything editable, including the code it writes. Like you can, if it makes a small error in a patch, you can just change it yourself and let it continue and it will be fine. So yeah, like those things are super important. We'll be doin…”
Alistair Pullen Aug 22, 2024 ▶ 34:18 Is finetuning GPT4o worth it?
LATENT SPACE Assertion Not publicly verifiable
Token log probabilities show models are most certain when writing code
“The certainty of code writing is so much more certain than every other aspect of Genie's loop. So whatever's going on under the hood, the model is really comfortable with writing code. There is no doubt, and it's like in, in the token probabilities.”
Alistair Pullen Aug 22, 2024 ▶ 35:23 Is finetuning GPT4o worth it?
LATENT SPACE Disclosure
Cosine trains Genie to output diffs rather than full file rewrites
“We train Genie to write diffs and, you know, essentially patches, right? Because it's more token efficient”
Alistair Pullen Aug 22, 2024 ▶ 35:49 Is finetuning GPT4o worth it?
LATENT SPACE Assertion Not publicly verifiable
Genie's SWE-bench success rate drops to roughly 50% past 60k tokens
“Performance of Jeannie over the length of the context window degrades fairly linearly. So actually, I actually broke it down by probability of solving a SWE bench issue. Given the number of tokens of the context window at 60 K, it's basically .5. So if you go …”
Alistair Pullen Aug 22, 2024 ▶ 36:26 Is finetuning GPT4o worth it?
LATENT SPACE Disclosure
Cosine's Genie hooks into existing GitHub CI rather than building environments
“The model itself is not in charge of like setting up the code base and running it. So genie sits on top of GitHub. And if you have CI running GitHub, you have GitHub Actions and stuff like that, then Genie essentially makes a call out to that, runs your CI, se…”
Alistair Pullen Aug 22, 2024 ▶ 40:04 Is finetuning GPT4o worth it?
LATENT SPACE Disclosure
Cosine injects synthetic AST errors into training data to teach error recovery
“And that was in sort of two parts. We synthetically generated runtime errors where we would Intentionally mess with the AST to make stuff not work or index out of bounds or refer to a variable that doesn't exist or errors that the foundational models just make…”
Alistair Pullen Aug 22, 2024 ▶ 46:15 Is finetuning GPT4o worth it?
LATENT SPACE Disclosure
Cosine refuses to publish SWE-bench trajectories to prevent competitor model distillation
“For the moment, as a closed source company, like fighting for an edge, we've decided not to publish that information for that exact reason. I don't want someone basically taking my tragedies and then taking a model that's suing them in GA and just distilling i…”
Alistair Pullen Aug 22, 2024 ▶ 50:35 Is finetuning GPT4o worth it?
LATENT SPACE Assertion Supported
Cosine's Genie achieved a state-of-the-art 43.8% on SWE-bench Verified
“We got 219 out of 500, which is 43.8%, which is To my knowledge, at least right now, state of the art also”
Alistair Pullen Aug 22, 2024 ▶ 52:38 Is finetuning GPT4o worth it?
LATENT SPACE Disclosure
Cosine built a version of Genie fine-tuned on its own codebase
“We have a version of Genie that is fine-tuned on our code base. So we basically, it's the base Genie, and then we run the same data pipeline that we run on, like, all the stuff that we did to generate the main data set on our repo.”
Alistair Pullen Aug 22, 2024 ▶ 55:12 Is finetuning GPT4o worth it?
20VC Assertion Not checkable as stated
Social network Genie failed because Facebook expanded into users over 30
“The thesis behind Jeannie was that there would be a social network for people above 30. Unfortunately, Facebook Expanded into that segment. So that didn't work out.”
George Zachary Oct 12, 2020 ▶ 16:39 20VC: CRV's George Zachary on His Relationship To Money and How it has Changed Over Time, Why The Best Founders Have Often Experienced Parental or Home Instability and The Stories Behind Investing in Unicorns; PillPack, Yammer and Udacity

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.