Factorio Learning Environment

7 statements across 1 episodes · 3 bullish · 2 bearish · 1 people on the record · first statement Apr 27, 2025 by Jack Hopkins · said 3 times in 1 episodes since 2025 · across every show →

Mentions by year

brought up most by Alessio Fanelli (3)

tap a year for its mentions
0021312025episodesmentions
0112025episodes it came up in
001.50.5312025episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Factorio Learning Environment, oldest first

Apr 27, 2025 positive
Assertion Supported
Claude scored nearly twice as high as the next best model
“So we see that Claude right here is almost got twice the score of the nearest best model.”
Jack Hopkins Apr 27, 2025 ▶ 26:32 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Apr 27, 2025 neutral
Assertion Contradicted
Untrained AI models exhibit a 100x competency gap versus human players
“It took models something like eight hours or so to get to the point where they have a kind of working factory that could make a few things, a few let's say iron gear wheels or electric circuits, or maybe some science and maybe start progressing through the tre…”
Jack Hopkins Apr 27, 2025 ▶ 24:05 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Apr 27, 2025 positive
Assertion Supported
Only Google models defined reusable code in the Factorio AI benchmark
“Only Google models tend to do this which is quite interesting.”
Jack Hopkins Apr 27, 2025 ▶ 16:17 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Apr 27, 2025 neutral
Assertion Supported
Claude 3.5 wrote fire-and-forget code while GPT-4 used defensive programming
“Claude, for instance, the Sonnet 3.5 was very much fire and forget. It would write code in a kind of Pythonic way, just like, let it fail. Don't be careful about it. Whereas GPT four would use defensive programming, use self assertions.”
Jack Hopkins Apr 27, 2025 ▶ 16:27 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Apr 27, 2025 positive
Insight
Dual reward signals prevent AI agent behavioral collapse in Factorio
“So we have these kind of two reward signals that compliment each other to and the reason why this is necessary is to avoid certain, I guess, behavioral collapses where a model might choose, for example, to mine coal and mine a billion or a trillion coal. And t…”
Jack Hopkins Apr 27, 2025 ▶ 7:15 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Apr 27, 2025 negative
Assertion Supported
Factorio benchmark results show reasoning models underperform expectations on extended planning
“One thing we have found in preliminary results is that the reasoning models don't seem to do as well as you'd expect in this setting. And I think that's probably because the way we set this up, it's a bit like we're already making it do reasoning traces over a…”
Jack Hopkins Apr 27, 2025 ▶ 12:12 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Apr 27, 2025 negative
Assertion Not publicly verifiable
Providing agents with RAG factory blueprints yielded zero benchmark score improvement
“When you try and move that into a benchmark setting with already pre-trained models, just using in-context learning, it's just not that helpful. A thousand lines of Python telling you how to make this kind of factory unit, which it may not be directly applicab…”
Jack Hopkins Apr 27, 2025 ▶ 22:00 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.