RL Environments

topic on 9 shows · 14 statements across 14 episodes

BG2 Pod Latent Space Lenny's Podcast the Neon Show No Priors Invest Like the Best the a16z Podcast TBPN 20VC

14 statements about RL Environments, every show

INVEST LIKE THE BEST Assertion Not checkable as stated
Movva: Frontier AI labs now spend much more on RL environments than data
“They used to spend much, that much on data, now they spend a lot more on RL environments, and these environments absolutely capture that relationship of recursive self-improvement on a verifiable task.”
Neil Movva Aug 25, 2026 ▶ 38:37 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
20VC Disclosure
Mercor shifts human data services to RL environments across modalities
“The data types also change very frequently. So we're, we've moved from supervised fine tuning to preference ranking to Rubric based annotation to now RL environments across a whole bunch of different modalities.”
Osvald Nitski Jul 25, 2026 ▶ 40:43 Mercor Head of Product on Revenue Concentration from Frontier Labs
NEON SHOW Insight
AI reinforcement learning requires a 20% to 40% task success rate
“If you have tasks that are too easy for the model, there is no learning signal. If it is too difficult, there is no learning signal. So there is a sweet spot of complexity that you'd want the RL environments to be at. Usually like, 20 to 40%, something in that…”
Jonathan Siddharth Jun 18, 2026 ▶ 12:39 Why Coding is the Fastest Path to AGI | Turing CEO Jonathan Siddharth
Catanzaro: RL Environments Are Just a Temporary Industry Fad
“So, I know I'm going on record on this, and like, I'm actually okay to be wrong, but I think RL Environments is just a fad.”
Sarah Catanzaro Dec 30, 2025 ▶ 23:42 [State of AI Startups] Memory/Learning, RL Envs & DBT-Fivetran — Sarah Catanzaro, Amplify
Goyal: Average AI companies cannot hire expertise to prevent reward hacking
“You need to have like a pretty specific expertise to design the RL environment in a way that's not vulnerable to reward hacking. And I think that either you'll end up with some fixed number of very well engineered RL environments, or you need to somehow employ…”
Ankur Goyal Dec 7, 2025 ▶ 24:35 The Great Evals Debate — Ankur Goyal & Malte Ubl
Scale AI has built RL environments for AI agents for over a year
“There's these things called RL environments that effectively are sandboxes for AI agents to play in to accomplish a goal so that they can learn how to accomplish that goal. We've been doing this for over a year.”
Jason Droege Oct 9, 2025 ▶ 19:20 Scale AI CEO on Meta’s $14B deal, scaling Uber Eats to $80B, & what frontier labs are building next
a16z Disclosure
Fedus: Periodic Labs uses physical experiments as RL reward functions
“And what we're doing, and by having the lab, is we create a physically grounded reward function. That becomes the basis on which we're optimizing against. And so, If a simulator has some deficiencies or some issues, we always error correct, because for us, the…”
Liam Fedus Sep 30, 2025 ▶ 5:01 Building an AI Physicist: ChatGPT Co-Creator’s Next Venture
Foody: AI evals and RL environments share the exact same data type
“There's not actually a nuance in the data type. It's more just a different semantic way of what describing what it's being used for. But ultimately it's just some stasis point for like, how do you measure what good looks like?”
Brendan Foody Sep 18, 2025 ▶ 14:25 Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
20VC Disclosure
Mercor commands a 50% to 60% market share in RL environment data
“There's the new data types that everyone's moving towards called RL environments where we'll call it, rough estimate is like, 50 to 60% of the market, and so doing quite well on that and expanding market share quickly.”
Brendan Foody Sep 15, 2025 ▶ 58:58 Mercor CEO & Co-Founder, Brendan Foody: How They Grew from $1M to $500M in 17 Months · 20VC with Harry Stebbings
BG2 Opinion
Wu: Short on reinforcement learning environment startups
“RL environments I think are really big right now as well. Unfortunately, I'm very short on those. not really I don't really see a lot of potential there. See a lot of potential and reinforcement learning and applying it, but I think the startup space around R…”
Sherwin Wu Sep 11, 2025 ▶ 46:01 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
NO PRIORS Prediction Not checkable as stated
Chen: Synthetic RL Environments Alone Won't Meet Future Frontier AI Demand
“Like, I don't think our environments alone will suffice just because, I mean, it depends on how you think about our environments, but oftentimes these are very, very rich trajectories are very, very long. And so it's almost like inconceivable that a single rew…”
Edwin Chen Jul 24, 2025 ▶ 16:59 No Priors Ep. 124 | With SurgeAI Founder and CEO Edwin Chen
NO PRIORS Insight
Mann: AI models can recursively self-improve by generating RL environments
“And then on the data side, RL environments are really important these days, but constructing those environments Has traditionally been expensive. Models are pretty good at writing environments, so it's another area where you can sort of recursively self-improv…”
Ben Mann Jun 12, 2025 ▶ 17:54 No Priors Ep. 118 | With Anthropic Co-Founder Ben Mann
Jin: Open source needs a standard to scale RL environments
“And so, like, there just needs to be some kind of, like, standard for open source developers to all, like, work together to, like, build up this, you know, to scale environments up to, like, millions and millions of environments.”
Roger Jin Apr 29, 2025 ▶ 8:03 What is an RL environment? w/ Nous Research's Roger Jin
TBPN Insight
Foody: RL environments make model customization far more data-efficient
“And I think a big reason for this is that it's now much more data efficient to customize models with RL environments. And a lot of this, like New kind of data versus fine tuning data that people would do historically.”
Brendan Foody Mar 24, 2025 ▶ 7:55 Brendan Foody on using AI to predict job performance and scaling from $1M to $100M in 11 months

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.