Continuous Pre Training
topic on 1 show · 1 statements across 1 episodes
1 statements about Continuous Pre Training, every show
Howard: Pre-training data mixes should be continuous per-batch functions, not discrete phases
“So the point at which they're doing proper continued pre-training is the point at which that becomes a continuum rather than a phase. So the only difference with what I was describing last time is to say, like, oh, they should, you know, There's a function or …”