Oct 16, 2025 · 1h 8m · latent-space

Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)

Kyle Corbitt · 39m spoken Shawn Wang · 15m spoken Alessio Fanelli · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Latent Space, OpenPipe co-founder Kyle Corbitt discusses the evolution of model adaptation from prompt distillation and LoRAs to task-specific reinforcement learning, culminating in OpenPipe's acquisition by CoreWeave.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 37.5% of the talking time here. How this is scored →

The hosts as informed peer 5.7 Guest teaching 4.2 Guest disagreement 2.0 The hosts pushing back 2.7
05100:0015:0030:0045:001:00:002:56–7:45 · The hosts as informed peer 5/10 Founding OpenPipe and the GPT-4 Distillation Era Swyx demonstrates market knowledge by analyzing how distillation startups were squeezed between frontier lab price cuts and neo-clouds offering fine-tuning. Kyle clarifies that neo-cloud developer experience was too poor to be real competition.7:45–14:42 · The hosts as informed peer 6/10 The Evolution of Fine-Tuning: Mistral, LoRAs, and ROI Swyx cites recent research from Thinking Machines and John Schulman regarding LoRAs. Kyle explains the infrastructure benefits of LoRA multiplexing and details when fine-tuning makes economic sense versus using frontier models.14:43–24:02 · The hosts as informed peer 5/10 Pivoting to Reinforcement Learning: PPO vs. GRPO When Swyx complains about mathematical complexity in RL papers, Kyle pushes back noting the equations are intuitive when written in code. Kyle then educates the hosts on why GRPO requires deterministic parallel rollouts, predicting it may be a dead end compared to PPO.24:02–29:05 · The hosts as informed peer 5/10 The Deterministic Sandbox Bottleneck for AI Agents Swyx asks why sandboxing is difficult if you can just capture inputs. Kyle delivers a masterclass on simulating failure modes, complex backend state, and the failure of LLM user simulators to capture real human distribution width.29:05–35:00 · The hosts as informed peer 6/10 Enterprise Tool-Call Environments and Compliance Workflows Alessio draws on portfolio company Various to discuss enterprise tool-call telemetry and compliance guardrails. Kyle responds that deterministic compliance workflows are poor candidates for RL compared to long-horizon agent tasks.35:00–44:33 · The hosts as informed peer 6/10 Beyond GRPO: Prompt Optimization (JEPA/DSPy) vs. Online Evals A spirited debate ensues over automated prompt optimization (JEPA/DSPy). Kyle bluntly states JEPA failed to produce results compared to RL, while Swyx defends prompt optimization as automating human lab system prompt iterations.44:33–50:22 · The hosts as informed peer 7/10 Macro AI Economics: Open vs. Closed Models and Compute Subsidies Kyle prompts the hosts on open vs closed model economics. Alessio and Swyx take the floor, breaking down Claude Code token subsidies, margin realities at Anthropic, and Stargate compute financing.50:22–1:00:11 · The hosts as informed peer 6/10 Ruler Library, LLM-as-a-Judge, and World Models Kyle explains OpenPipe's Ruler library and how relative group ranking solves reward modeling even with weak judge models. Swyx and Kyle then contrast pre-training code world models with execution simulation world models.1:00:11–1:08:12 · The hosts as informed peer 5/10 CoreWeave Acquisition, Serverless RL, and Continual Learning Kyle shares the backstory of the CoreWeave acquisition via Weights & Biases and pitches serverless RL for continual agent learning. Swyx and Kyle reflect on YC's advice regarding rapid shipping versus long-term conviction.2:56–7:45 · Guest teaching 3/10 Founding OpenPipe and the GPT-4 Distillation Era Swyx demonstrates market knowledge by analyzing how distillation startups were squeezed between frontier lab price cuts and neo-clouds offering fine-tuning. Kyle clarifies that neo-cloud developer experience was too poor to be real competition.7:45–14:42 · Guest teaching 3/10 The Evolution of Fine-Tuning: Mistral, LoRAs, and ROI Swyx cites recent research from Thinking Machines and John Schulman regarding LoRAs. Kyle explains the infrastructure benefits of LoRA multiplexing and details when fine-tuning makes economic sense versus using frontier models.14:43–24:02 · Guest teaching 6/10 Pivoting to Reinforcement Learning: PPO vs. GRPO When Swyx complains about mathematical complexity in RL papers, Kyle pushes back noting the equations are intuitive when written in code. Kyle then educates the hosts on why GRPO requires deterministic parallel rollouts, predicting it may be a dead end compared to PPO.24:02–29:05 · Guest teaching 7/10 The Deterministic Sandbox Bottleneck for AI Agents Swyx asks why sandboxing is difficult if you can just capture inputs. Kyle delivers a masterclass on simulating failure modes, complex backend state, and the failure of LLM user simulators to capture real human distribution width.29:05–35:00 · Guest teaching 4/10 Enterprise Tool-Call Environments and Compliance Workflows Alessio draws on portfolio company Various to discuss enterprise tool-call telemetry and compliance guardrails. Kyle responds that deterministic compliance workflows are poor candidates for RL compared to long-horizon agent tasks.35:00–44:33 · Guest teaching 5/10 Beyond GRPO: Prompt Optimization (JEPA/DSPy) vs. Online Evals A spirited debate ensues over automated prompt optimization (JEPA/DSPy). Kyle bluntly states JEPA failed to produce results compared to RL, while Swyx defends prompt optimization as automating human lab system prompt iterations.44:33–50:22 · Guest teaching 2/10 Macro AI Economics: Open vs. Closed Models and Compute Subsidies Kyle prompts the hosts on open vs closed model economics. Alessio and Swyx take the floor, breaking down Claude Code token subsidies, margin realities at Anthropic, and Stargate compute financing.50:22–1:00:11 · Guest teaching 5/10 Ruler Library, LLM-as-a-Judge, and World Models Kyle explains OpenPipe's Ruler library and how relative group ranking solves reward modeling even with weak judge models. Swyx and Kyle then contrast pre-training code world models with execution simulation world models.1:00:11–1:08:12 · Guest teaching 3/10 CoreWeave Acquisition, Serverless RL, and Continual Learning Kyle shares the backstory of the CoreWeave acquisition via Weights & Biases and pitches serverless RL for continual agent learning. Swyx and Kyle reflect on YC's advice regarding rapid shipping versus long-term conviction.2:56–7:45 · Guest disagreement 1/10 Founding OpenPipe and the GPT-4 Distillation Era Swyx demonstrates market knowledge by analyzing how distillation startups were squeezed between frontier lab price cuts and neo-clouds offering fine-tuning. Kyle clarifies that neo-cloud developer experience was too poor to be real competition.7:45–14:42 · Guest disagreement 1/10 The Evolution of Fine-Tuning: Mistral, LoRAs, and ROI Swyx cites recent research from Thinking Machines and John Schulman regarding LoRAs. Kyle explains the infrastructure benefits of LoRA multiplexing and details when fine-tuning makes economic sense versus using frontier models.14:43–24:02 · Guest disagreement 3/10 Pivoting to Reinforcement Learning: PPO vs. GRPO When Swyx complains about mathematical complexity in RL papers, Kyle pushes back noting the equations are intuitive when written in code. Kyle then educates the hosts on why GRPO requires deterministic parallel rollouts, predicting it may be a dead end compared to PPO.24:02–29:05 · Guest disagreement 2/10 The Deterministic Sandbox Bottleneck for AI Agents Swyx asks why sandboxing is difficult if you can just capture inputs. Kyle delivers a masterclass on simulating failure modes, complex backend state, and the failure of LLM user simulators to capture real human distribution width.29:05–35:00 · Guest disagreement 2/10 Enterprise Tool-Call Environments and Compliance Workflows Alessio draws on portfolio company Various to discuss enterprise tool-call telemetry and compliance guardrails. Kyle responds that deterministic compliance workflows are poor candidates for RL compared to long-horizon agent tasks.35:00–44:33 · Guest disagreement 5/10 Beyond GRPO: Prompt Optimization (JEPA/DSPy) vs. Online Evals A spirited debate ensues over automated prompt optimization (JEPA/DSPy). Kyle bluntly states JEPA failed to produce results compared to RL, while Swyx defends prompt optimization as automating human lab system prompt iterations.44:33–50:22 · Guest disagreement 1/10 Macro AI Economics: Open vs. Closed Models and Compute Subsidies Kyle prompts the hosts on open vs closed model economics. Alessio and Swyx take the floor, breaking down Claude Code token subsidies, margin realities at Anthropic, and Stargate compute financing.50:22–1:00:11 · Guest disagreement 2/10 Ruler Library, LLM-as-a-Judge, and World Models Kyle explains OpenPipe's Ruler library and how relative group ranking solves reward modeling even with weak judge models. Swyx and Kyle then contrast pre-training code world models with execution simulation world models.1:00:11–1:08:12 · Guest disagreement 1/10 CoreWeave Acquisition, Serverless RL, and Continual Learning Kyle shares the backstory of the CoreWeave acquisition via Weights & Biases and pitches serverless RL for continual agent learning. Swyx and Kyle reflect on YC's advice regarding rapid shipping versus long-term conviction.2:56–7:45 · The hosts pushing back 2/10 Founding OpenPipe and the GPT-4 Distillation Era Swyx demonstrates market knowledge by analyzing how distillation startups were squeezed between frontier lab price cuts and neo-clouds offering fine-tuning. Kyle clarifies that neo-cloud developer experience was too poor to be real competition.7:45–14:42 · The hosts pushing back 2/10 The Evolution of Fine-Tuning: Mistral, LoRAs, and ROI Swyx cites recent research from Thinking Machines and John Schulman regarding LoRAs. Kyle explains the infrastructure benefits of LoRA multiplexing and details when fine-tuning makes economic sense versus using frontier models.14:43–24:02 · The hosts pushing back 2/10 Pivoting to Reinforcement Learning: PPO vs. GRPO When Swyx complains about mathematical complexity in RL papers, Kyle pushes back noting the equations are intuitive when written in code. Kyle then educates the hosts on why GRPO requires deterministic parallel rollouts, predicting it may be a dead end compared to PPO.24:02–29:05 · The hosts pushing back 3/10 The Deterministic Sandbox Bottleneck for AI Agents Swyx asks why sandboxing is difficult if you can just capture inputs. Kyle delivers a masterclass on simulating failure modes, complex backend state, and the failure of LLM user simulators to capture real human distribution width.29:05–35:00 · The hosts pushing back 3/10 Enterprise Tool-Call Environments and Compliance Workflows Alessio draws on portfolio company Various to discuss enterprise tool-call telemetry and compliance guardrails. Kyle responds that deterministic compliance workflows are poor candidates for RL compared to long-horizon agent tasks.35:00–44:33 · The hosts pushing back 5/10 Beyond GRPO: Prompt Optimization (JEPA/DSPy) vs. Online Evals A spirited debate ensues over automated prompt optimization (JEPA/DSPy). Kyle bluntly states JEPA failed to produce results compared to RL, while Swyx defends prompt optimization as automating human lab system prompt iterations.44:33–50:22 · The hosts pushing back 2/10 Macro AI Economics: Open vs. Closed Models and Compute Subsidies Kyle prompts the hosts on open vs closed model economics. Alessio and Swyx take the floor, breaking down Claude Code token subsidies, margin realities at Anthropic, and Stargate compute financing.50:22–1:00:11 · The hosts pushing back 3/10 Ruler Library, LLM-as-a-Judge, and World Models Kyle explains OpenPipe's Ruler library and how relative group ranking solves reward modeling even with weak judge models. Swyx and Kyle then contrast pre-training code world models with execution simulation world models.1:00:11–1:08:12 · The hosts pushing back 2/10 CoreWeave Acquisition, Serverless RL, and Continual Learning Kyle shares the backstory of the CoreWeave acquisition via Weights & Biases and pitches serverless RL for continual agent learning. Swyx and Kyle reflect on YC's advice regarding rapid shipping versus long-term conviction.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 53.7% · guest 46.3%0:00 · the hosts 53.7% · guest 46.3%3:00 · the hosts 22.6% · guest 77.4%3:00 · the hosts 22.6% · guest 77.4%6:00 · the hosts 46.6% · guest 53.4%6:00 · the hosts 46.6% · guest 53.4%9:00 · the hosts 39.6% · guest 60.4%9:00 · the hosts 39.6% · guest 60.4%12:00 · the hosts 14.5% · guest 85.5%12:00 · the hosts 14.5% · guest 85.5%15:00 · the hosts 7.3% · guest 92.7%15:00 · the hosts 7.3% · guest 92.7%18:00 · the hosts 34.1% · guest 65.9%18:00 · the hosts 34.1% · guest 65.9%21:00 · the hosts 15.6% · guest 84.4%21:00 · the hosts 15.6% · guest 84.4%24:00 · the hosts 17.6% · guest 82.4%24:00 · the hosts 17.6% · guest 82.4%27:00 · the hosts 62.1% · guest 37.9%27:00 · the hosts 62.1% · guest 37.9%30:00 · the hosts 44.7% · guest 55.3%30:00 · the hosts 44.7% · guest 55.3%33:00 · the hosts 47.5% · guest 52.5%33:00 · the hosts 47.5% · guest 52.5%36:00 · the hosts 50% · guest 50%36:00 · the hosts 50% · guest 50%39:00 · the hosts 73.9% · guest 26.1%39:00 · the hosts 73.9% · guest 26.1%42:00 · the hosts 43.6% · guest 56.4%42:00 · the hosts 43.6% · guest 56.4%45:00 · the hosts 83.9% · guest 16.1%45:00 · the hosts 83.9% · guest 16.1%48:00 · the hosts 71% · guest 29%48:00 · the hosts 71% · guest 29%51:00 · the hosts 2.9% · guest 97.1%51:00 · the hosts 2.9% · guest 97.1%54:00 · the hosts 39.2% · guest 60.8%54:00 · the hosts 39.2% · guest 60.8%57:00 · the hosts 32.7% · guest 67.3%57:00 · the hosts 32.7% · guest 67.3%1:00:00 · the hosts 25.2% · guest 74.8%1:00:00 · the hosts 25.2% · guest 74.8%1:03:00 · the hosts 13.3% · guest 86.7%1:03:00 · the hosts 13.3% · guest 86.7%1:06:00 · the hosts 22.1% · guest 77.9%1:06:00 · the hosts 22.1% · guest 77.9%
Sharpest disagreement ▶ 36:49 Kyle dismisses JEPA and prompt optimization

Kyle rejects the premise that prompt optimization competes with weight updates, bluntly asserting that JEPA barely beat a naive baseline while RL reached 96%.

Hardest push from the hosts ▶ 36:31 Swyx defends prompt optimization techniques

Swyx pushes back against Kyle's dismissal of JEPA, arguing that prompt engineering models the genetic evolution big labs use for system prompts.

Biggest teaching moment ▶ 24:08 The reality of building deterministic agent sandboxes

Kyle thoroughly explains why naive input-capture fails in agent evaluation, detailing the difficulty of modeling subtle environment failure modes and realistic human response distributions.

The host holds their own ▶ 45:43 Hosts dissect frontier lab token subsidy economics

Alessio and Swyx demonstrate deep domain expertise analyzing how frontier labs use consumer subscriptions as loss leaders and subsidize compute utilization.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Founding OpenPipe and the GPT-4 Distillation Era 5312 Swyx demonstrates market knowledge by analyzing how distillation startups were squeezed between frontier lab price cuts and neo-clouds offering fine-tuning. Kyle clarifies that neo-cloud developer experience was too poor to be real competition.
The Evolution of Fine-Tuning: Mistral, LoRAs, and ROI 6312 Swyx cites recent research from Thinking Machines and John Schulman regarding LoRAs. Kyle explains the infrastructure benefits of LoRA multiplexing and details when fine-tuning makes economic sense versus using frontier models.
Pivoting to Reinforcement Learning: PPO vs. GRPO 5632 When Swyx complains about mathematical complexity in RL papers, Kyle pushes back noting the equations are intuitive when written in code. Kyle then educates the hosts on why GRPO requires deterministic parallel rollouts, predicting it may be a dead end compared to PPO.
The Deterministic Sandbox Bottleneck for AI Agents 5723 Swyx asks why sandboxing is difficult if you can just capture inputs. Kyle delivers a masterclass on simulating failure modes, complex backend state, and the failure of LLM user simulators to capture real human distribution width.
Enterprise Tool-Call Environments and Compliance Workflows 6423 Alessio draws on portfolio company Various to discuss enterprise tool-call telemetry and compliance guardrails. Kyle responds that deterministic compliance workflows are poor candidates for RL compared to long-horizon agent tasks.
Beyond GRPO: Prompt Optimization (JEPA/DSPy) vs. Online Evals 6555 A spirited debate ensues over automated prompt optimization (JEPA/DSPy). Kyle bluntly states JEPA failed to produce results compared to RL, while Swyx defends prompt optimization as automating human lab system prompt iterations.
Macro AI Economics: Open vs. Closed Models and Compute Subsidies 7212 Kyle prompts the hosts on open vs closed model economics. Alessio and Swyx take the floor, breaking down Claude Code token subsidies, margin realities at Anthropic, and Stargate compute financing.
Ruler Library, LLM-as-a-Judge, and World Models 6523 Kyle explains OpenPipe's Ruler library and how relative group ranking solves reward modeling even with weak judge models. Swyx and Kyle then contrast pre-training code world models with execution simulation world models.
CoreWeave Acquisition, Serverless RL, and Continual Learning 5312 Kyle shares the backstory of the CoreWeave acquisition via Weights & Biases and pitches serverless RL for continual agent learning. Swyx and Kyle reflect on YC's advice regarding rapid shipping versus long-term conviction.

Statements from this episode (24)

Assertion Not checkable as stated
Corbitt: OpenPipe Hit $1M ARR Within Eight Months of Launch
“And so we got our first three customers after launching probably within a month, and we were doing significant revenue. Over the next six months, we actually got to a million in ARR over about a eight month period following that launch.”
Kyle Corbitt Oct 16, 2025 ▶ 4:52
Assertion Not checkable as stated
Corbitt: GPU Cloud Fine-Tuning Offerings Failed Due to Poor Usability
“I did not see the competition ever really materialize from the Neo clouds, from the GPU providers. Everybody had an offering in fine tuning. When we were talking to customers, nobody used them because they just were really hard to use.”
Kyle Corbitt Oct 16, 2025 ▶ 7:04
Opinion
Corbitt: No downside to using LoRAs for task-specific model customization
“For the types of training runs that we're interested in, where it's like, hey, I'm doing a relatively lightweight customization of an existing model for a specific task, there's really no downside to using Allura, and there's a lot of, like, upsides from an, l…”
Kyle Corbitt Oct 16, 2025 ▶ 10:34
Opinion
Corbitt: Fine-tuning offers poor ROI for 90% of unconstrained use cases
“I would say for 90% of use cases where you aren't forced to a smaller model, then it's still not a good ROI, and you probably shouldn't invest in it today.”
Kyle Corbitt Oct 16, 2025 ▶ 12:49
Assertion Not checkable as stated
Corbitt: Fine-tuning compute runs cost only $5 to a few hundred dollars
“The dollar cost, I would say, is basically never a factor. It's just so much less than the time, the amount you're spending this engineer to do the work that it's not, I mean, it's, you know, each of these runs is between five and a couple of hundred dollars.”
Kyle Corbitt Oct 16, 2025 ▶ 14:21
Prediction Not checkable as stated
Corbitt: 55-60% chance RL becomes the standard pattern for deploying scale agents
“I think that the chances that like everyone should be, or, you know, everyone who's deploying an agent at scale should be doing RL with it, either as part of sort of like a, you know, like pre-deployment or even like continuously as it's deployed, that that's …”
Kyle Corbitt Oct 16, 2025 ▶ 18:18
Opinion
Corbitt: GRPO is likely a dead end due to parallel rollout constraints
“The big downside, the huge downside of GRPO, and I think actually the reason why GRPO actually is likely to be a dead end, and we probably will not be continue using it indefinitely. The fact that you need to have these parallel rollouts in order to train on i…”
Kyle Corbitt Oct 16, 2025 ▶ 22:46
Insight
Corbitt: PPO enables training purely on real production traces without simulation
“And PPO, now in practice, a lot of times when you're training with PPO, you also will use an environment like that because it lets you do a bunch of runs and be more data efficient. But at least in principle, you have the option with PPO, you can actually, lik…”
Kyle Corbitt Oct 16, 2025 ▶ 23:39
Insight
Corbitt: LLM user simulators lack the diversity needed to train robust agents
“If you're just purely training on kind of like an LLM user simulator, it's going to have its own idea of, like, what the correct way to answer is, and the breadth of, like, a way a human might respond in this situation is wider, and your agent just may not be …”
Kyle Corbitt Oct 16, 2025 ▶ 25:31
Disclosure
Corbitt: Realistic sandbox environments almost universally do not exist in enterprises
“When we talk to enterprises almost universally, that's like not something that really exists. So there are some startups, like there's some companies we've talked to that do have it and we can just like use that, but it's a very, very small number that, that a…”
Kyle Corbitt Oct 16, 2025 ▶ 26:10
Assertion Not checkable as stated
Swix: RL environment startups sell contracts to major labs for seven figures
“For those who are interested, when you make a reference to our environment startups selling to the big labs, they're selling it for a lot of money. Like at least seven figures.”
Shawn Wang Oct 16, 2025 ▶ 28:36
Insight
Corbitt: Agent RL requires real runs inside highly realistic environments
“For RL to work, you have to be looking at real runs, ideally of your actual agent in its current state across within an environment as real as possible.”
Kyle Corbitt Oct 16, 2025 ▶ 30:10
Opinion
Corbitt: Building RL training environments is currently a services-heavy business
“It seems to me like that definitely is a services heavy business at the moment as it, as it's presently constituted.”
Kyle Corbitt Oct 16, 2025 ▶ 31:36
Assertion Not checkable as stated
Corbitt: Prompt optimization methods like JEPA failed OpenPipe's agent benchmarks
“It didn't work on the problems we tried it on. It just didn't. It got like a minor boost over the sort of like more naive prompt we had and was just like, it was like, okay, Just kind of like our naive prompt with our model gets maybe like 50% on this benchmar…”
Kyle Corbitt Oct 16, 2025 ▶ 37:08
Opinion
Swix: Standalone model routing companies will be commoditized and absorbed
“I think that I'm very bullish on model routing as a feature, but less bullish on model routing companies because of exactly stuff like this, where like, it is just going to get, get absorbed into the model.”
Shawn Wang Oct 16, 2025 ▶ 44:08
Prediction Not checkable as stated
Swix: Open-source token generation share will rise but stay well below 50%
“I think it's going to go up because of the amount of enterprise adoption of open models that I'm seeing. And also there's a lot of demand. Like there's the enterprises would much rather be on open models if they actually could get the performance they're loo…”
Shawn Wang Oct 16, 2025 ▶ 44:52
Prediction Not checkable as stated
Alessio: Open-source token share could reach 15-20% excluding coding
“I think once you take coding out, I think, yeah, it can be like 15, 20%, but I think with coding, it's still gonna be very low because like these max plans are like, So subsidized and so many tokens are being generated”
Alessio Fanelli Oct 16, 2025 ▶ 45:47
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year. And my money is on, they have to do a coin. Like it's, I'm not a crypto guy at all, but like, y…”
Shawn Wang Oct 16, 2025 ▶ 49:13
Assertion Supported
Corbitt: OpenPipe Beat Frontier Models Using a Qwen 32B Judge
“One of the results we published was we used Quen 2.5 14 B as the model we're training, and as the judge we used Quen 2.5 32 B, which is, like, Not, I mean, it's fine, but it's like not a, it's much worse than any frontier model. Right. And even with that combi…”
Kyle Corbitt Oct 16, 2025 ▶ 53:18
Opinion
Corbitt: Generic LLM-as-a-Judge Models Won't Beat Frontier Labs
“I'm pretty bearish on like Hey, this is a model that is trained as an LMS judge, but it's a generic LMS judge that can be used to judge anything. I just don't think you're going to beat the frontier labs on that.”
Kyle Corbitt Oct 16, 2025 ▶ 56:35
Assertion Not checkable as stated
Corbitt: Weights & Biases founders drove CoreWeave's OpenPipe acquisition
“So that was driven by actually mostly the weights and biases founding team. Lucas and Sean, particularly. So they, had recently been acquired by CoreWeave and CoreWeave was looking to continue growing up the stack. And so, yeah, they approached me and were lik…”
Kyle Corbitt Oct 16, 2025 ▶ 1:00:24
Insight
Corbitt: AI inference could be 10x larger if reliability issues are solved
“I think that there is today, like. 10 times as much AI inference that could exist than is existing right now, just Purely with projects that are like sitting in the proof of concept stage and have not been deployed because there's like huge bucket of those. An…”
Kyle Corbitt Oct 16, 2025 ▶ 1:04:50
Insight
Corbitt: RL reward hacking is easily detected as models repeat the exploit
“Reward hacking is quite easy to detect once it starts happening, because once the model's found some hack, it just starts, like, doing it all the time.”
Kyle Corbitt Oct 16, 2025 ▶ 1:05:41
Insight
Corbitt: Ambitious startups benefit more from long-term vision than fast YC shipping
“If I do another startup, like I would like, I think at least some points I probably would have done better to be like heads down and execute on my vision for longer and like, kind of like go for the more ambitious thing, but that would take longer to sort of l…”
Kyle Corbitt Oct 16, 2025 ▶ 1:07:45
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.