RL Environments
topic on 9 shows · 14 statements across 14 episodes
BG2 Pod
Latent Space
Lenny's Podcast
the Neon Show
No Priors
Invest Like the Best
the a16z Podcast
TBPN
20VC
14 statements about RL Environments, every show
Movva: Frontier AI labs now spend much more on RL environments than data
“They used to spend much, that much on data, now they spend a lot more on RL environments, and these environments absolutely capture that relationship of recursive self-improvement on a verifiable task.”
Mercor shifts human data services to RL environments across modalities
“The data types also change very frequently. So we're, we've moved from supervised fine tuning to preference ranking to Rubric based annotation to now RL environments across a whole bunch of different modalities.”
AI reinforcement learning requires a 20% to 40% task success rate
“If you have tasks that are too easy for the model, there is no learning signal.
If it is too difficult, there is no learning signal.
So there is a sweet spot of complexity that you'd want the RL environments to be at.
Usually like, 20 to 40%, something in that…”
Catanzaro: RL Environments Are Just a Temporary Industry Fad
“So, I know I'm going on record on this, and like, I'm actually okay to be wrong, but I think RL Environments is just a fad.”
Goyal: Average AI companies cannot hire expertise to prevent reward hacking
“You need to have like a pretty specific expertise to design the RL environment in a way that's not vulnerable to reward hacking. And I think that either you'll end up with some fixed number of very well engineered RL environments, or you need to somehow employ…”
Scale AI has built RL environments for AI agents for over a year
“There's these things called RL environments that effectively are sandboxes for AI agents to play in to accomplish a goal so that they can learn how to accomplish that goal. We've been doing this for over a year.”
Fedus: Periodic Labs uses physical experiments as RL reward functions
“And what we're doing, and by having the lab, is we create a physically grounded reward function. That becomes the basis on which we're optimizing against. And so, If a simulator has some deficiencies or some issues, we always error correct, because for us, the…”
Foody: AI evals and RL environments share the exact same data type
“There's not actually a nuance in the data type. It's more just a different semantic way of what describing what it's being used for. But ultimately it's just some stasis point for like, how do you measure what good looks like?”
Mercor commands a 50% to 60% market share in RL environment data
“There's the new data types that everyone's moving towards called RL environments where we'll call it, rough estimate is like, 50 to 60% of the market, and so doing quite well on that and expanding market share quickly.”
Wu: Short on reinforcement learning environment startups
“RL environments I think are really big right now as well.
Unfortunately, I'm very short on those.
not really I don't really see a lot of potential there.
See a lot of potential and reinforcement learning and applying it, but I think the startup space around R…”
Chen: Synthetic RL Environments Alone Won't Meet Future Frontier AI Demand
“Like, I don't think our environments alone will suffice just because, I mean, it depends on how you think about our environments, but oftentimes these are very, very rich trajectories are very, very long. And so it's almost like inconceivable that a single rew…”
Mann: AI models can recursively self-improve by generating RL environments
“And then on the data side, RL environments are really important these days, but constructing those environments Has traditionally been expensive. Models are pretty good at writing environments, so it's another area where you can sort of recursively self-improv…”