Code Execution Evals
topic on 1 show · 1 statements across 1 episodes
1 statements about Code Execution Evals, every show
Ground-truth evaluations without code execution provide significant mileage for AI
“So code execution, like checking the correctness of solutions and running tests automatically can help, although not, although you can get a lot of mileage out of evals that don't have code execution in them that just compare against ground truth.”