Everything Kevin Wang said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Kevin Wang: Cross-entropy trajectory classification enables scalable deep reinforcement learning
“I think it's because we're fundamentally shifting the burden of learning from something like, Q-learning or, like, regressing to, like, TD errors, which we know is quite spurious and noisy and biased, to fundamentally, like, a classification problem. We're try…”
Wang: Deep Self-Supervised Method Beats SOTA on Goal-Conditioned RL Significantly
“We do achieve state-of-the-art performance on goal-conditioned RL and Jack's CCRL by a significant amount.”
Kevin Wang: Scaling RL network depth unlocks effective batch size scaling
“We notice that we see that scaling width actually also improves performance, and we also find that actually by scaling depth, we actually unlock the ability to scale along batch size as well.”
Kevin Wang: 1000-layer RL networks can train on single 80GB H100
“The nice thing is that all of our experiments, even the thousand layer networks, can be run on one single, 80 gigabyte, each 100 GPU.”
Wang: AI in 2025 is earlier in evolution than early-2010s mobile
“I think there's some interesting parallels, say, to where like AI is today, where I think that AI is probably a few quarters earlier than that, and people still don't know what's, what's going to happen, and there's a lot more change going on, so it's probably…”
Wang: Thinking about competition under $5M ARR is generally distracting
“And even I would say up through maybe two million to five million ARR, I think that thinking too much about competition is generally distracting for two reasons. One is that you don't really know the identity of your company and also you don't know You are not…”
Wang: Traditional value-based reinforcement learning fails to scale
“And so what we did is that we know that traditional RL, like let's say like value value-based RL doesn't really scale, right? This is pretty clear from the literature.”
Wang: 64 layers saturate performance in most reinforcement learning tasks
“Within our paper, like, for most environments we are able to, like, saturate, like, get to, like, almost perfect performance within just, you know, we don't even need to get to, like, a thousand layers. Like, maybe just 64 layers, for example, is sufficient.”
Wang: Early startup peers were way smarter than big consulting colleagues
“This is very, very different from big consulting, I would say. The people are way smarter, and this is a very much more scrappy environment.”
Wang: Marketing creates power-law outcomes and intense competitive selection
“These marketing buyers did have, there are a lot of network effects of marketing, because if I'm running much better marketing than you are, and we have a very similar sort of brand, I'm going to win and I'm going to win really big. I'm going to win sort of li…”
Wang: GPU environments collect hundreds of millions of RL timesteps hourly
“With these, like, GPU accelerated environments, we can collect hundreds of millions of time steps of data within just a few hours”