Tay: Zero-shot benchmark scores at 1B model scale are random chance
Yi Tay · The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka · Jul 5, 2024 · at 1:46:43
Yi Tay (co-founder of Reka and former Google Brain researcher) discusses why academic efficiency papers reporting gains on 1B parameter models often demonstrate statistical noise rather than genuine improvements.
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance performers, right? And they will be like, ok, I get like 50 versus 54, I'm better, but like, dude, that's all random chance, right?”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →