The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Kyle Corbitt no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yeah. So you left YC, you spent a year in the, kind of, the wilderness. You went through YC, uh, as 23. Uh, what's that journey like? What's the...

A You know, I was very excited about AI things in general. Um, this was, so I left YC, I guess, uh, beginning of twenty-twenty-two, and I was trying out a bunch of different things. Um, ended up landing on what turned into OpenPipe in early twenty-twenty-three. This was, uh, let's see, so I'd been working, so my, my co-founder is my brother, um, my little brother, which has been a fun journey on its own. We were looking at different ideas, and one thing we realized was we actually started the company immediately after the GPT-IV launch, and what we saw as the opportunity in the market at the time, which has changed since then, was GPT-IV was insanely expensive and extremely powerful, but there was an opportunity to distill, like, specific workflows from GPT-IV down to much smaller, much cheaper models, and there was, like, a very clear value prop there. Given how expensive GPT-IV was, it was hard to deploy in production, But you could sort of like take those abilities and deploy them much more cheaply. So, so that was kind of the first thing we built was this kind of very managed, very clean, um, distillation flow.

AI assessment note: “I left YC, I guess, uh, beginning of twenty-twenty-two, and I was trying out”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Can you double click on why it's hard to do the sandboxing? Because in principle, we just capture all the inputs.

A Yeah. Well, you don't need to just capture all the inputs. You need, you need a system that reacts the same way your production system does. That's, and in many different ways. And, um, so let's say you're, you're Airbnb, right? And I'm bringing this up because this is like an example of one that like, you know, companies have gone out and built sandboxes. Like if you're Airbnb and you're trying to, Um, you want to train an agent to, like, maybe you're not Airbnb, fine. You're, you're a company like us that's trying to train an agent to, like, do really well at operating Airbnb and booking on your behalf, right? Like, you have to build a copy of the Airbnb website that reacts to you as the user the exact same way that the real one does with the same failure modes, right? Because if you don't include the same failure modes and bugs they have, then, like, one of those, when one of those bugs comes up in production, your agent's gonna have no idea what to do with it. It's just gonna fall over. You also need to simulate if this is like a sort of cooperative agent, right, where it's getting human input as well and kind of like working with the human to get something done, which in practice is the way a lot of these are deployed. You also need to simulate the user. And I mean, you can do the naive thing and just say, oh, we're going to have a separate LLM that, you know, with a syste…

AI assessment note: “You need a system that reacts the same way your production system does.”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q like fine tuning the model? Because even the open models were not that great, you know? And so what were maybe the bottlenecks? Like instead of having three to get to like 30 customers, did you feel like in the beginning it was like a matter of like just the market growing, like the open source models not being good enough, like the fine tuning not being simple, efficient enough?

A The pain point, I guess repeating what I said before was the price was too high on the closed models, but you couldn't just drop in an open model and replace them. Cause like you're saying, the quality was quite bad. Especially as you're moving to two smaller model sizes, but larger models, open models weren't even available at that time. So, so that's kind of where the value prop was, was like, Hey, the closed models are too expensive. At least the ones that are performance enough to, to do the job, the open ones are not good enough. We have like a very clear managed flow. Um, the way the flow worked was, was quite simple. You simply put in our SDK. It's a drop in replacement for the open AI SDK. It's capturing, you continue to use GPT four in production for a period of time. We're capturing the requests and responses. And then we had just a very clean managed flow where it's like, okay, at some point you say, Hey, I want to distill this down and you, you train on that. And then, you know, we provided an API that was a direct drop in replacement. You would just change kind of the inference URL and you were using your own model in it at your app continued working.

AI assessment note: “price was too high on the closed models, but you couldn't just drop in an open model”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q Right. Okay. When was the switch to RL? Was it when a one preview came out? You were maybe like, okay, it's time to move on from SFT or?

A Yeah. So that was a big moment for us, um, with, you know, there's all the leaks before that about strawberry and all this. And like, you know, a lot of people talking about, okay, how are they doing it? Um, we realized through that that like, okay, someone's figured out how to make RL actually work. Like with LLMs, which was not a thing. I mean, it was a thing that like some people had played around with before that, but it wasn't like I think many people were thinking about. And so our bet at that point was, yes, let's figure out whether this works for task specifically. And the space we, we just, I think it's important to kind of like tease out different parts of the market. I think with the release of O-one, and this has been like proved out many times with releases since then, I think like there's now like a very strong consensus that like, okay, On the frontier model, like general purpose model side, investments in RL are paying off. I think, I, I don't think most people would argue with that. You're, especially as, as you're getting into these agentic tasks, um, and training them to do that, like, it seems very clear. Well, obviously the big labs are paying like ridiculous amounts of money for these environments and everything, but also like they're actually getting really good results. The, the, the models coming out, you know, we're seeing it, especially on the coding …

AI assessment note: “Yeah. So that was a big moment for us, um, with, you know, there's all the leaks”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q then on top of it, you have helping people trying to build the exact replica of their thing. There's obviously value in like the formally verified ones. We verified that. Do you think there's value in this like RL environment startups that are building like somewhat generic, but test specific environments. And then if none of those work, then what do we do instead of GRPO? I guess the question.

A Yeah, I suspect there is value in that. You know, I think the, you know, the folks buying those environments and training on them in the big labs would have the best knowledge on how well they work. I think they probably work okay. I think they probably also are like, You know, and we'll see maybe with the next generation of models released, like how well they transfer. I would say so far, um, it seems like they don't train well enough. Like if you, if you use, um, you know, open AI's agent interface, it's like, okay. Or if you use the computer use products that, that everybody's putting out, they're like, okay, but like not reliable enough to like, actually like let go, do something interesting unsupervised in the world. And I think if the eight, you know, if the environments they were training in were high enough fidelity, Then they would be good enough in the same way that like coding agents can go much further because I think that in that case we do have environments that are much higher fidelity because it's a much simpler environment in a lot of ways. It's like, it's a code base. It's like maybe running a web browser. Like it's, it's, it's much easier to capture the full realistic environment in that context.

AI assessment note: “Yeah, I suspect there is value in that.”

Partly raw tape D 3 · C 4 · P 3 · Cm 3 3.30

Q Okay. One thing I saw from your, your post was, uh, your North star as the RL team at core weave is to build an old world where every agent learns continually from his real world experience. So you're touching on the hot topic of the moment, continual learning. What else do we need to get there?

A I super believe that. And like, that's basically the vision where I'm like, you know, I've keep talking about these percentages, like if we get to the world where we build that, um, then I think it's just like the advantages are huge. They're clear. Everyone should just deploy their, their agents that way. Um, we want to be like the team that builds the, the software, um, that makes that easy to do. So I talked to a lot of engineers at our customers and they're trying to deploy agents and it's so easy to get the initial prototype and like something that like kind of works well. It is so hard to get from that to something that like you are confident is reliable enough to actually deploy in production. And when you actually look at what those failure modes look like, it's like, oh yeah, like we know if it gets in this situation or if it gets like these kind of like inputs, like it behaves funnily. But then it's like, yeah, you can update your problem to, to address that. But like, that's not scalable because at a certain point it's like going to start breaking other things. You know, you don't know what it's breaking. You really want some way to just like say, okay, look, this thing you did there, that was the wrong thing. Just like adjust this behavior when you get in this and then, you know, otherwise carry on. Right. And that's what we can do with RL. And that's what we can do…

AI assessment note: “we want to be like the team that builds the, the software”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.