Exploit Bench
other on 1 show · 0 statements across 0 episodes
1 statements about Exploit Bench, every show
OpenAI GPT-5.6 Matches Anthropic Mythos on Exploit Bench at Double Efficiency
“So they ran GPT, 5.6 through a test called exploit bench and the cap, the capabilities were basically on par with mythos. And this is obviously a cybersecurity test, but mythos is much less token efficient. So mythos do something like 300,000 tokens for the fo…”