Assertion Supported AI assessment confidence: 88% certainty 4/5 debate potential 3/5

Tay: Zero-shot benchmark scores at 1B model scale are random chance

Yi Tay · The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka · Jul 5, 2024 · at 1:46:43

Yi Tay (co-founder of Reka and former Google Brain researcher) discusses why academic efficiency papers reporting gains on 1B parameter models often demonstrate statistical noise rather than genuine improvements.

0:00 / 0:19exact quote · 19.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Every time some people propose like this, they run like some zero-shot score on like some LM event harness or something like that, and you know like at one B scale, all the numbers are random, basically. Like all your bull kill, they're all like random chance performers, right? And they will be like, ok, I get like 50 versus 54, I'm better, but like, dude, that's all random chance, right?”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Yi Tay

Opinion
Yi Tay: Gap Between Closed AI Labs and Open-Source Is Increasing
“I think the gap is definitely increasing.”
Yi Tay Jan 23, 2026 ▶ 53:49 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Opinion
Yi Tay: IR and RecSys research lags significantly behind NeurIPS and ICML
“Also the IR community and the retrieval community is also like always behind the mainstream. And then now it's just probably gotten even more worse because of ILM and stuff. So, okay, I'm getting into Hottick territory, but it's just, like, certain conferences…”
Yi Tay Jan 23, 2026 ▶ 1:17:23 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Insight
Tay: Frontier AI researchers cannot maintain standard nine-to-five work-life balance
“You cannot be, like, checking out on, like, Friday, Saturday, Sunday, and, like, work at, like, nine to five if you want to, like, Make progress, or like, some people are just so good at detaching, like, ok, like, you know, like, eight pm, I'm not going to, my…”
Yi Tay Jul 5, 2024 ▶ 38:49 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Opinion
Tay: Long context architecture is the future of AI over RAG
“And, yeah, I mean, I think long context is definitely the future, rather than rec. But I mean, they could be used in conjunction, like,”
Yi Tay Jul 5, 2024 ▶ 1:40:05 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Disclosure
DeepMind Abandoned AlphaProof to Run Gemini End-to-End for IMO Math
“We wanted to try to, like, use, actually use Gemini as an end-to-end model. Basically, no, no second system with alpha proof. No second system. In, text out.”
Yi Tay Jan 23, 2026 ▶ 13:38 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Prediction Not checkable as stated
Tay: Most specialized tools will be subsumed directly into model parameters
“Then the most I can see in the future is there'll be a model then that, that is, there's something that really cannot be subsumed by a model. Then you just use a tool or something, right? But my prediction is that I think most things can be subsumed by the mod…”
Yi Tay Jan 23, 2026 ▶ 19:31 Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.