reinforcement learning environments
1 statements across 1 episodes · 0 bullish · 0 bearish · 1 people on the record · first statement Apr 29, 2025 by Roger Jin · across every show →
Everything said about reinforcement learning environments, oldest first
Apr 29, 2025 neutral
Jin: Nous RL environments return literal tokens instead of parsed text
“Another, like, kind of weird quirky thing about our design is that at least for text, the thing that's returned by each of these environments is, like, the literal tokens. So it's not, like, it's not text, it's not, like, messages, it's the tokens.”