The Ledger, every show

Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.

shows every show 44 of 44
every show
clear all ✕
Corbitt: GRPO is likely a dead end due to parallel rollout constraints
“The big downside, the huge downside of GRPO, and I think actually the reason why GRPO actually is likely to be a dead end, and we probably will not be continue using it indefinitely. The fact that you need to have these parallel rollouts in order to train on i…”
Kyle Corbitt Oct 16, 2025 ▶ 22:46 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
LATENT SPACE Assertion Not checkable as stated
Corbitt: Prompt optimization methods like JEPA failed OpenPipe's agent benchmarks
“It didn't work on the problems we tried it on. It just didn't. It got like a minor boost over the sort of like more naive prompt we had and was just like, it was like, okay, Just kind of like our naive prompt with our model gets maybe like 50% on this benchmar…”
Kyle Corbitt Oct 16, 2025 ▶ 37:08 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: Fine-tuning offers poor ROI for 90% of unconstrained use cases
“I would say for 90% of use cases where you aren't forced to a smaller model, then it's still not a good ROI, and you probably shouldn't invest in it today.”
Kyle Corbitt Oct 16, 2025 ▶ 12:49 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
LATENT SPACE Prediction Not checkable as stated
Corbitt: 55-60% chance RL becomes the standard pattern for deploying scale agents
“I think that the chances that like everyone should be, or, you know, everyone who's deploying an agent at scale should be doing RL with it, either as part of sort of like a, you know, like pre-deployment or even like continuously as it's deployed, that that's …”
Kyle Corbitt Oct 16, 2025 ▶ 18:18 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
LATENT SPACE Assertion Supported
Corbitt: OpenPipe Beat Frontier Models Using a Qwen 32B Judge
“One of the results we published was we used Quen 2.5 14 B as the model we're training, and as the judge we used Quen 2.5 32 B, which is, like, Not, I mean, it's fine, but it's like not a, it's much worse than any frontier model. Right. And even with that combi…”
Kyle Corbitt Oct 16, 2025 ▶ 53:18 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: AI inference could be 10x larger if reliability issues are solved
“I think that there is today, like. 10 times as much AI inference that could exist than is existing right now, just Purely with projects that are like sitting in the proof of concept stage and have not been deployed because there's like huge bucket of those. An…”
Kyle Corbitt Oct 16, 2025 ▶ 1:04:50 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
LATENT SPACE Assertion Not checkable as stated
Corbitt: GPU Cloud Fine-Tuning Offerings Failed Due to Poor Usability
“I did not see the competition ever really materialize from the Neo clouds, from the GPU providers. Everybody had an offering in fine tuning. When we were talking to customers, nobody used them because they just were really hard to use.”
Kyle Corbitt Oct 16, 2025 ▶ 7:04 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: No downside to using LoRAs for task-specific model customization
“For the types of training runs that we're interested in, where it's like, hey, I'm doing a relatively lightweight customization of an existing model for a specific task, there's really no downside to using Allura, and there's a lot of, like, upsides from an, l…”
Kyle Corbitt Oct 16, 2025 ▶ 10:34 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: Agent RL requires real runs inside highly realistic environments
“For RL to work, you have to be looking at real runs, ideally of your actual agent in its current state across within an environment as real as possible.”
Kyle Corbitt Oct 16, 2025 ▶ 30:10 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: Building RL training environments is currently a services-heavy business
“It seems to me like that definitely is a services heavy business at the moment as it, as it's presently constituted.”
Kyle Corbitt Oct 16, 2025 ▶ 31:36 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: Generic LLM-as-a-Judge Models Won't Beat Frontier Labs
“I'm pretty bearish on like Hey, this is a model that is trained as an LMS judge, but it's a generic LMS judge that can be used to judge anything. I just don't think you're going to beat the frontier labs on that.”
Kyle Corbitt Oct 16, 2025 ▶ 56:35 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: RL reward hacking is easily detected as models repeat the exploit
“Reward hacking is quite easy to detect once it starts happening, because once the model's found some hack, it just starts, like, doing it all the time.”
Kyle Corbitt Oct 16, 2025 ▶ 1:05:41 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: Ambitious startups benefit more from long-term vision than fast YC shipping
“If I do another startup, like I would like, I think at least some points I probably would have done better to be like heads down and execute on my vision for longer and like, kind of like go for the more ambitious thing, but that would take longer to sort of l…”
Kyle Corbitt Oct 16, 2025 ▶ 1:07:45 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
LATENT SPACE Assertion Not checkable as stated
Corbitt: Fine-tuning compute runs cost only $5 to a few hundred dollars
“The dollar cost, I would say, is basically never a factor. It's just so much less than the time, the amount you're spending this engineer to do the work that it's not, I mean, it's, you know, each of these runs is between five and a couple of hundred dollars.”
Kyle Corbitt Oct 16, 2025 ▶ 14:21 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: PPO enables training purely on real production traces without simulation
“And PPO, now in practice, a lot of times when you're training with PPO, you also will use an environment like that because it lets you do a bunch of runs and be more data efficient. But at least in principle, you have the option with PPO, you can actually, lik…”
Kyle Corbitt Oct 16, 2025 ▶ 23:39 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Corbitt: LLM user simulators lack the diversity needed to train robust agents
“If you're just purely training on kind of like an LLM user simulator, it's going to have its own idea of, like, what the correct way to answer is, and the breadth of, like, a way a human might respond in this situation is wider, and your agent just may not be …”
Kyle Corbitt Oct 16, 2025 ▶ 25:31 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
LATENT SPACE Disclosure
Corbitt: Realistic sandbox environments almost universally do not exist in enterprises
“When we talk to enterprises almost universally, that's like not something that really exists. So there are some startups, like there's some companies we've talked to that do have it and we can just like use that, but it's a very, very small number that, that a…”
Kyle Corbitt Oct 16, 2025 ▶ 26:10 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
LATENT SPACE Assertion Not checkable as stated
Corbitt: OpenPipe Hit $1M ARR Within Eight Months of Launch
“And so we got our first three customers after launching probably within a month, and we were doing significant revenue. Over the next six months, we actually got to a million in ARR over about a eight month period following that launch.”
Kyle Corbitt Oct 16, 2025 ▶ 4:52 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
LATENT SPACE Assertion Not checkable as stated
Corbitt: Weights & Biases founders drove CoreWeave's OpenPipe acquisition
“So that was driven by actually mostly the weights and biases founding team. Lucas and Sean, particularly. So they, had recently been acquired by CoreWeave and CoreWeave was looking to continue growing up the stack. And so, yeah, they approached me and were lik…”
Kyle Corbitt Oct 16, 2025 ▶ 1:00:24 Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.