Reinforcement Learning

topic on 23 shows · 247 statements across 156 episodes

the Y Combinator Startup Podcast More or Less BG2 Pod Innovators & Investors Cheeky Pint American Optimist the Knowledge Project Latent Space Lenny's Podcast the Neon Show No Priors WTF is with Nikhil Kamath Invest Like the Best Sourcery Capital Allocators Catalyst Top Founders the MAD Podcast the a16z Podcast Big Technology All-In TBPN 20VC

The latest 60 statements about Reinforcement Learning, every show

Kantrowitz: Reinforcement learning layers add ruthlessness to AI models
“What we have now is that the reinforcement learning type of AI technology has been put on top of the self supervised learning to get these AI models working better, which has added a level of ruthlessness to them. Because one of the things we know about RL is …”
Alex Kantrowitz Sep 7, 2026 ▶ 41:22 GPT-6 & OpenAI’s Comeback, Hugging Face Attack Debate, Ballmer’s Scandalous Legacy
BIG TECHNOLOGY Assertion Supported
Kantrowitz: OpenAI paused some reinforcement learning on new model training
“We are starting to see some of the labs do things like opening. I, for instance, paused some reinforcement learning for a bit on the training of its new models.”
Alex Kantrowitz Sep 7, 2026 ▶ 43:08 GPT-6 & OpenAI’s Comeback, Hugging Face Attack Debate, Ballmer’s Scandalous Legacy
LENNY'S PODCAST Prediction Not checkable as stated
Enterprises will split work between cheap open-weight and expensive frontier models
“I think what we're going to see is a split between job functions that demand kind of mid IQ intelligence, and those will often be open weight, sort of biased with reinforcement learning, you know, things that make the models even cheaper, more performant for a…”
Anish Acharya Sep 6, 2026 ▶ 23:21 Why companies are becoming a series of loops | Anish Acharya (a16z)
a16z Insight
Litt: Reinforcement learning struggles to reward intermediate mathematical theory building
“I think what is definitely true is that, like, the skill of, like, developing a theory or, like, building your understanding of some poorly understood object is, like, a fuzzier one. So it might be harder, you know I guess you can try, you can tell it, you kno…”
Daniel Litt Sep 1, 2026 ▶ 27:08 Can AI Learn Mathematical Intuition?
Jeffrey: Base model AI startups struggle most bridging research to product
“And then on the model side, you're looking for like world expert, reinforcement learning people who have a scientific insight on something that what you're really trying to find there is how you compare the researcher's intuition With like product thinking and…”
Corey Jeffrey Aug 27, 2026 ▶ 16:17 Building High-Growth Tech Ventures in the AI Era with Kory Jeffrey of Inovia Capital
a16z Insight
Acharya: Domain-specialized open models outperform general models through RL and reasoning traces
“If you actually have a problem that you can specialize the model around with your reasoning traces, You can start to create this compounding advantage in your domain for your customer base, where you're able to kind of shape the intelligence to be better than …”
Anish Acharya Aug 26, 2026 ▶ 9:19 The State of AI: Models, Moats, and the Consumer Renaissance
MAD Assertion Supported
Wolf: Frontier AI Training Has Shifted From RLHF to Pure RL
“What we know though, is we moved from this pure, like human data, you know, that was first just pre-training on human data and then also aligning with like human preferences that was called RLHF, where we had a lot of human in the loop and human data. To like …”
Thomas Wolf Aug 6, 2026 ▶ 29:08 “OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf
Y COMBINATOR Assertion Not checkable as stated
Chaubard: CPU-based environment simulation bottlenecks on-policy reinforcement learning rollouts
“The, it's amazing how much of simulators, when you call environment.step, is still run on the CPU, and so that's usually the bottleneck for a lot of your on-policy rollouts”
Francois Chaubard Jul 29, 2026 ▶ 2:50 Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
SOURCERY Insight
Wu: Reinforcement learning can solve basically any clearly defined benchmark
“We're kind of getting to the point where you can solve basically any benchmark, right? Because what does it mean to have a benchmark? It means you've already defined the task. You've clarified what success or failure looks like. You've given a bunch of example…”
Scott Wu Jul 27, 2026 ▶ 30:12 Inside the Fastest-Growing Category in AI: Scott Wu, CEO of $26B Cognition · Sourcery with Molly O'Shea
MAD Disclosure
Feldman: Cerebras serves second-tier AI labs for model training
“We do RL and we do traditional training too. Not for the largest models, for the largest lab, but for the next tier.”
Andrew Feldman Jul 23, 2026 ▶ 45:43 Cerebras CEO: Why GPUs Can't Do Fast Inference
LATENT SPACE Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Eiso Kant Jul 22, 2026 ▶ 6:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
LATENT SPACE Prediction Not checkable as stated
Kant: Reinforcement learning will move earlier into LLM pre-training
“I have I would say a not commonly held opinion that reinforcement learning will move earlier and earlier into pre-training.”
Eiso Kant Jul 22, 2026 ▶ 45:09 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Kant: AI coding models perform best in their creators' proprietary harnesses
“No doubt it's going to be better in your own harness. And it's just because of like, where are you putting your reinforcement learning compute, right? You're putting your RL and your synthetic data. You're putting it to your own harness because it's the one th…”
Eiso Kant Jul 22, 2026 ▶ 1:01:03 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Kant: RL compute cannot scale like pre-training due to task batch constraints
“And RL is batch size constraint, right? So like you are ultimately in your batch size constraint because you don't have infinite tasks, right? When you've got the entire web, you can be much more flexible in scaling up your batch size because you've got the en…”
Eiso Kant Jul 22, 2026 ▶ 1:39:07 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Beam: Nature and scientific experiments are ultimate verifiers for RL
“But what at Lilo we believe is that actually science running the scientific method and using nature and experiments as verifier is like the ultimate version of that. And so what we're building, we'll talk about these things that we call AI science factories. T…”
Andy Beam Jul 16, 2026 ▶ 6:59 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
LATENT SPACE Assertion Supported
Beam: Reinforcement learning achieves only 5% to 6% GPU FLOP utilization
“And for reinforcement learning, it's always somewhere, like, around five to, like, six percent. So, said differently, that means that we're getting, like, five percent of the actual GPU computing power that we're paying for.”
Andy Beam Jul 16, 2026 ▶ 1:38:42 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Perszyk: Task-specific reinforcement learning fails to produce generalizable intelligence
“You can use things like reinforcement learning to get them really good at specific tasks that we might care about, but you do that for one task and you, it is not good at another task or it doesn't generalize.”
Danielle Perszyk Jul 11, 2026 ▶ 18:32 Why AI Agents Don't Actually Understand You — Danielle Perszyk, Amazon AGI Lab
Bubna: RL Rollouts Are Extremely Bursty and Can Require 100,000 Sandboxes
“RL is insanely bursty. Like when you're doing rollouts you sometimes need a 100,000 sandboxes.”
Akshat Bubna Jul 8, 2026 ▶ 15:42 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
Bubna: Transferring RL weights is fundamentally an OS memory problem
“Like the way you move around your KV cache and how efficiently you can do it, how efficiently you move your weights from your training GPUs to your inference GPUs in RL is, there's a lot of degrees of freedom, and it is basically a systems problem of Moving me…”
Akshat Bubna Jul 8, 2026 ▶ 31:40 The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
BIG TECHNOLOGY Assertion Not checkable as stated
Bosworth: Reinforcement learning plays a far bigger role in AI than predicted
“Reinforcement learning is playing a huge, a much bigger role in today's kind of AI than people had maybe predicted two or three years ago that it would.”
Andrew Bosworth Jul 8, 2026 ▶ 38:01 Meta CTO Andrew Bosworth: Our Path To Frontier AI, Renting Models, Consumer AI's Struggles
Kutylowski: Task-focused reinforcement learning outperforms broad multi-task training in specific domains
“And if you run this reinforcement learning step on too many different tasks, the model will be able to do all of that. But once again, it's going to be very, very, very broad. And if you focus on making sure that the model understands and knows that it needs t…”
Jarek Kutylowski Jul 7, 2026 ▶ 7:15 Why Specialized AI Models Are Challenging the Frontier Labs — With DeepL CEO Jarek Kutylowski
OpenAI's Chen: Reinforcement learning struggles in subjective, hard-to-grade fields
“RLs traditionally had headwinds when it's come to fields that, you know, it's more kind of, Subjective than objective. So if you kind of think of, you know, one kind of, you know example of this is creative writing, where, you know, you could take two pieces o…”
Mark Chen Jun 25, 2026 ▶ 5:54 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
LATENT SPACE Disclosure
OpenAI's three research pillars are pre-training, RL, and alignment
“At the very highest level, right, we have an org that focuses on pre-training, right, which is, you know, giving models a lot of world knowledge. We focus on RL, like, teaching the models how to reason with that knowledge, how to chain the little insights toge…”
Mark Chen Jun 25, 2026 ▶ 14:03 Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen
LATENT SPACE Prediction Not checkable as stated
Zaharia: Customizing AI models will get significantly easier over time
“My feeling is, like customizing models is actually going to get way easier over time. That's what we're finding, because The base models are smarter, so they generate better traces in RL already, and then RL is about learning from your own past traces, and the…”
Matei Zaharia Jun 24, 2026 ▶ 1:03:14 The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
Malde: Standard reinforcement learning is broken for continual learning
“RL, it's still taking all of this kind of Useful information from the real world, like I mentioned, all the corrections and everything, and putting it into just one number. Which is really broken.”
Ronak Malde Jun 21, 2026 ▶ 19:13 ⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
NEON SHOW Insight
Siddharth: Non-binary knowledge work requires rubric-based AI evaluation
“Like with code or with math, it's relatively more binary, easy to verify. But how do you verify the quality of a board deck? Yeah. It's a, you have to be, you have to have like a good rubric based evaluator.”
Jonathan Siddharth Jun 18, 2026 ▶ 16:27 Why Coding is the Fastest Path to AGI | Turing CEO Jonathan Siddharth
NEON SHOW Insight
Siddharth: Verifiability makes coding ideal for reinforcement learning improvements
“Coding is one of those areas where, because it's verifiable, I think that there is a good path to using reinforcement learning to improve coding models quickly.”
Jonathan Siddharth Jun 18, 2026 ▶ 21:44 Why Coding is the Fastest Path to AGI | Turing CEO Jonathan Siddharth
NEON SHOW Assertion Supported
Siddharth: AI compute is shifting significantly toward post-training reinforcement learning
“In the past, it was a lot of the compute went into pre-training. Now a lot of compute goes into reinforcement learning in post-training as well. Especially after O-one came out and DeepSeek came out.”
Jonathan Siddharth Jun 18, 2026 ▶ 52:36 Why Coding is the Fastest Path to AGI | Turing CEO Jonathan Siddharth
MAD Disclosure
OpenAI plans to increasingly rely on reinforcement learning to scale intelligence
“When you have a lot of compute, you want to turn that compute into intelligence in a way that's useful, and RL is one way of doing it, and we just started doing it then, and we're going to do a lot more of it now.”
Dan Roberts Jun 4, 2026 ▶ 25:35 OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
MAD Insight
Roberts: Powerful pre-trained models are necessary for effective RL and reasoning
“If you have a powerful enough pre-trained model, then it can start to do well at RL. It can start to like think at use test time compute to for instance, solve, solve math problems that it wouldn't otherwise be able to do.”
Dan Roberts Jun 4, 2026 ▶ 27:15 OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
MAD Prediction Not checkable as stated
OpenAI will release reinforcement learning products for consulting, banking, and legal
“I definitely think OpenAI will have amazing products that will be relevant in those domains, and some amount of RL will play a role in there.”
Dan Roberts Jun 4, 2026 ▶ 37:01 OpenAI's Dan Roberts: Why AI Can Now Make Discoveries
Hong: Lean and Rust yield superior reinforcement learning convergence over Python
“If you want proof to be informal math, It's very annoying, because then that's, like, just makes objective function. Your code is something like Python, your proof is, say, natural language, math proof. You will not have very strong RL kind of performance, rig…”
Carina Hong Jun 3, 2026 ▶ 30:00 Scaling Past Informal AI - Carina Hong, Axiom Math
CATALYST Assertion Not checkable as stated
AI reinforcement learning energy demand probably already exceeds traditional pre-training
“And this is an area that is becoming huge in terms of energy demand. It'll, it will, The probably already is bigger than what we have historically considered training, you know, pre-training”
Garth Sheldon-Coulson May 28, 2026 ▶ 43:57 Building inference data centers on the high seas
LATENT SPACE Prediction Not checkable as stated
Burazin: RL workloads will reach 50% of Daytona's volume this month
“It will be this one 50%, yeah.”
Ivan Burazin May 21, 2026 ▶ 28:22 AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
MAD Disclosure
Dubois: OpenAI expanded RL training from math competitions to real-world coding
“We were able to take many of the tools that we built for these, like, verifiable reward cases, and we were able to use them more generally in on, for reinforcement on, like, real use cases, and I think that's, like, really why we're feeling that right now in, …”
Yann Dubois May 21, 2026 ▶ 3:40 OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
MAD Insight
Dubois: RL allows AI reasoning models to backtrack wrong paths earlier
“Part of it is the model knowing when it's going down the wrong path. But this is also something that we can that the model can be trained for with reinforcement learning is like knowing, okay, like that seems like not a great path. Let me backtrack and let me …”
Yann Dubois May 21, 2026 ▶ 22:59 OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
MAD Insight
Dubois: Starting post-training with RL without SFT is extremely inefficient
“Because if you just started from reinforcement learning, it would be very inefficient. Because the problem with reinforcement learning is that you have to stumble across the right answer, basically.”
Yann Dubois May 21, 2026 ▶ 36:48 OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
MAD Insight
Dubois: RL becomes effective once base models possess strong world priors
“It seems that after crossing a certain scale of models that know basically everything about the world, and what we call, like, good priors about the world, It seems that reinforcement learning just started to work, and this is not only with LMS. Robotics seems…”
Yann Dubois May 21, 2026 ▶ 40:08 OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
MAD Insight
Dubois: Agentic RL training suffers from sparse reward credit assignment
“When we are training more agentic systems, you only know whether you're correct at the end of your very long rollout. So you get very little information per token of whether you were correct or not. And it's hard to say it's hard to basically do attribution. I…”
Yann Dubois May 21, 2026 ▶ 41:15 OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
a16z Disclosure
Caldwell: Mariana Minerals uses reinforcement learning to automate mineral refineries
“We're making a big bet on autonomy and refineries, where we use reinforcement learning to actually remove humans from the loop in determining how refineries operate.”
Turner Caldwell May 13, 2026 ▶ 0:28 The Founders Who Left Tesla to Rebuild America | a16z
Rao: More efficient inference directly increases reinforcement learning efficiency
“If we're doing reinforcement learning on the model, it's basically inference within a sandbox with a reward function, right? And so if the model's better at more efficient inference, that RL is more efficient as well.”
Krishna Rao May 13, 2026 ▶ 9:44 Inside Anthropic's $100 Billion Al Compute Commitment | CFO Krishna Rao · Invest Like The Best
MAD Assertion Not checkable as stated
Zico Kolter: Reinforcement learning is now the foundation of all AI post-training
“RL is now the foundation of really all post training. It's all done by RL.”
Zico Kolter May 7, 2026 ▶ 1:04:27 OpenAI Board Member Zico Kolter: Modern AI Is Just 200 Lines of Code
BIG TECHNOLOGY Assertion Not checkable as stated
Kantrowitz: Scale AI now does most of its training via reinforcement learning
“Scale AI, Alexander Wang's company, they told me recently that most of the training that they're doing is reinforcement learning, where you build environments for the bots and they go and they try to figure out what to do.”
Alex Kantrowitz Apr 27, 2026 ▶ 42:52 Apple After Tim Cook, OpenAI’s New Mojo, Meta’s Internal Tracking Escapade
Patel: RL Simulation Environments Run on CPUs, Not GPUs or ASICs
“So the environments can get more and more complex, and those environments run on CPUs. They don't run on GPUs. They don't run on ASICs. The ASICs run the model,”
Dylan Patel Apr 23, 2026 ▶ 38:44 The Supply and Demand of AI Tokens | Dylan Patel Interview · Invest Like The Best
KNOWLEDGE PROJECT Assertion Not checkable as stated
Brockman: OpenAI's 10-year roadmap focused on RL, unsupervised learning, then complexity
“We came up with what I would Really say is almost the technical plan that we have pursued for the past 10 years. Number one, solve reinforcement learning. Number two, solve unsupervised learning. And number three was gradually learn more complicated, in quotes…”
Greg Brockman Apr 22, 2026 ▶ 4:00 Ai Goes Parabolic | OpenAI Co-Founder Greg Brockman
Physical Intelligence aims to fuse generative AI prior knowledge with reinforcement learning
“So, I think the big challenge, and this is kind of what I'm leaning up to, and what I hope to, ah, that we'll figure out here at Physical Intelligence is how to combine those threads. How to bring in all of that knowledge that you get with generative AI, but a…”
Sergey Levine Mar 31, 2026 ▶ 16:38 World's Top Researcher on AI, LLMs, and Robot Intelligence · Invest Like The Best
Levine: Physical Intelligence trained espresso-making robot using repeated RL practice
“And for example, we had this demo on, ah, making espresso. That system practiced making those espressos many, many times and used that to improve robustness, improve speed, improve throughput.”
Sergey Levine Mar 31, 2026 ▶ 18:30 World's Top Researcher on AI, LLMs, and Robot Intelligence · Invest Like The Best
Levine: Robots Surpass Human Speed by Editing Out Cognitive Pauses
“It turns out to be like pretty straightforward to go in and like find all those pauses and remove them. And you can speed things up further, so you can get a task where a person demonstrates what it means to succeed, and then you can have the robot practice th…”
Sergey Levine Mar 31, 2026 ▶ 38:54 World's Top Researcher on AI, LLMs, and Robot Intelligence · Invest Like The Best
Lample: Long-horizon RL trajectories require new algorithms beyond GRPO
“GRPO, for instance, it doesn't really work with any bit of policy, which was okay initially, because you are solving math problems that can be solved in like a few thousand tokens, so the model can actually generate them pretty quickly, so when you do your upd…”
Guillaume Lample Mar 30, 2026 ▶ 45:43 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
BG2 Insight
Turley: Quantitative Knowledge Work Will See Agentic Breakthroughs Due to RL Suitability
“I won't be surprised if you see this happen for other forms of sort of quantitative knowledge work, just because it happens to have the properties that code has. It's testable. You know if it worked or not. It's very RL friendly.”
Nick Turley Mar 15, 2026 ▶ 19:08 ChatGPT – The Super Assistant Era | BG2 Guest Interview · Bg2 Pod
BIG TECHNOLOGY Assertion Supported
Kantrowitz: Scale AI shifted majority of training to reinforcement learning
“I just did the story for with about scale AI saying that the majority of their training has moved to reinforcement learning where they train models to act in specific environments like filling out forms, and then they baked those capabilities back Into the mod…”
Alex Kantrowitz Mar 9, 2026 ▶ 28:43 AI Revenue Explodes, Dario’s Memo, McDonalds’ CEO’s Baby Burger Bite
MAD Disclosure
Axiom Math focuses on post-training reinforcement learning to achieve performance gains
“And I think that we shouldn't do pre-training. We shouldn't try to just only train from scratch. I think we're kind of focusing on post-training reinforcement learning can potentially get us better performance gain.”
Corinna Hong Feb 26, 2026 ▶ 17:26 AI That Can Prove It’s Right: Verification as the Missing Layer in AI — Carina Hong
Welling: Diffusion Models Share Exact Mathematics With Non-Equilibrium Stochastic Thermodynamics
“It turns out that the mathematics that we use for diffusion models, but even for reinforcement learning, for Schrodinger bridges, for MCMC sampling, has the same mathematics as this theory, this physical theory of non-equilibrium Systems.”
Max Welling Feb 25, 2026 ▶ 4:59 🔬Max Welling: Materials Underlie Everything
LATENT SPACE Prediction Not checkable as stated
O'Laughlin: Tech industry may face a CPU shortage from AI coding and RL
“You feel like we might actually be seeing a CPU shortage partially because of this refresh cycle, but partially also because like I legitimately believe the cloud code Cloud code is increasing software creation and then on top of that, there is real demand fro…”
Doug O'Laughlin Feb 24, 2026 ▶ 1:56:07 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
WTF Prediction Not checkable as stated
Amodei: Static training data is becoming less central than dynamic RL data
“Static data is becoming less important and what we might call like dynamic data that the model creates itself is, you know, for reinforcement learning is becoming more important. So, you know, I don't think data is, is, is quite the most central thing anymore,…”
Dario Amodei Feb 24, 2026 ▶ 1:00:37 The AI Tsunami is Here & Society Isn't Ready | Dario Amodei x Nikhil Kamath | People by WTF · Nikhil Kamath
Dean: Applying RL to non-verifiable domains would dramatically improve AI models
“How do you get RL to work for non-verifiable domains? I think it's a pretty interesting open problem because I think that would broaden out the capabilities of the models, the improvements that you're seeing in both math and coding if we could apply those to o…”
Jeff Dean Feb 12, 2026 ▶ 42:58 The AI Frontier: from Gemini 3 Deep Think distilling to Flash — Jeff Dean
TBPN Insight
Coogan: Git commit history makes reinforcement learning uniquely effective for coding
“Long context reinforcement learning has been very, very successful in the coding world because Git has a complete history of every line of code that's been written, every comment, why it happened.”
John Coogan Feb 6, 2026 ▶ 18:17 Anthropic’s Trust Nuke, OpenAI’s new releases, Google Claims AI Crown | Diet TBPN
SOURCERY Opinion
Das: Reinforcement learning is an inefficient paradigm requiring massive sample sizes
“One is RL's kind of a shitty paradigm to learn. Karpathy obviously talks about this a lot. It takes a lot of samples to learn some very basic stuff because you only get a reward at the end. You don't actually understand things as it's happening.”
Didi Das Feb 5, 2026 ▶ 51:23 How Anthropic’s $100M Anthology Fund Works | Menlo Ventures · Sourcery with Molly O'Shea
INNOVATORS & INVESTORS Prediction Not checkable as stated
Toeman: AI quality will improve via reinforcement learning on elite content
“If we want to have AI give us amazing caliber stuff. It needs to be trained on amazing caliber stuff. So, and that will change over time because we'll start, you know, we'll, we'll start that reinforcement learning around better, better quality content.”
Jeremy Toeman Feb 3, 2026 ▶ 33:43 Scaling AI Video Innovation with Jeremy Toeman of JWX | The Innovators & Investors Podcast
White: Writing bulletproof RL verifiers is far harder than supervised training
“Pre-training or training transformers, you know on just data, like just supervised training where you just have the inputs and the outputs directly, very nice, relaxing, you know, like things are always robust, you know, things go pretty smoothly. When we do t…”
Andrew White Jan 28, 2026 ▶ 1:12:03 🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.