base model fine-tuning
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Apr 29, 2025 by Roger Jin · across every show →
Everything said about base model fine-tuning, oldest first
Apr 29, 2025 positive
Jin: Token-level RL enables mixing instruct and base model fine-tuning
“Another thing you can do, like, with, in, in, like, the token world is, like, the trainer is now, like, agnostic to, like, chat versus instruct model. And what that means is, like, you can do all, like, the cool, like, R-one-zero kind of style, like, experimen…”