Eval Benchmarks
topic on 1 show · 1 statements across 1 episodes
1 statements about Eval Benchmarks, every show
Kevin Hou: Code search requires multiple needles, unlike standard benchmarks
“Benchmarks like the one that I showed you before heavily skew towards this idea of needle in a haystack. It's the idea that you can sift through a corpus of text and find some instance of something that is relevant to you. Note, it is only one single needle. S…”