“We've got very little rewards right now, but pretty quickly over the next year or two, you're going to start to see much more meaningful and long horizon rewards.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Sholto Douglas
PredictionNot checkable as stated
Douglas: Generalist models will obsolete specialized fine-tuned models
“I really do think that similar to how we saw with large pre-trained models before with small fine-tuned models made it like, had gains over the sort of GPT-II era, but then were obsoleted by GPT-IV being generally good at everything. I think, to be honest, you…”
Douglas: Frontier Labs Avoid RL Training Directly on ARC-AGI
“And I mean, I think if you are old on Arc AGI, then it would, you'd probably get superhuman at it pretty fast. But I think we're all trying not to RL on it so that it functions as like an interesting held out.”
This entire site, over 500 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.