Coding Benchmarks
topic on 1 show · 1 statements across 1 episodes
1 statements about Coding Benchmarks, every show
Shaw: Current AI coding benchmarks rely on redundant bespoke test harnesses
“In fact, every single benchmark that gets released at least I would say maybe all coding benchmarks that get released at this point are some form of instruction container tests with some bespoke harness that was coded up That feels very analogous to every othe…”