The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Eiso Kant argument clarity score 4.3/5 from 11 exchanges on raw tape · average scores: directness 4.5 · coherence 4.5 · precision 4.1 · compression 3.7 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
14exchanges match
11on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q AWS. I think AWS is building one for them. I guess that's what you were saying. You need a tenant in that case. The hyperscaler is willing to do it because they have Anthropic as a tenant. So as an indie lab, for lack of a better term, like, you have to own your own destiny, and that involves building your own data center, right? Just to play it back.

A Yeah, I think as a foundation model company, you can go two paths, right? You can choose to deeply partner with a hyperscaler. And kind of, you know, become, you know, have them become an owner in you and really go all the way. And I think that's one direction. But at the same time, the world is getting to a point where I don't think anyone has any doubts anymore that we're now on track to, to reach human level capabilities and intelligence. And in that world, 29 trillion dollars of knowledge work rewrites itself, right? Scientific progress starts pushing beyond levels that we've ever seen. And all of it is intelligence on compute. And so we are far more in a, in a foundation model company is far more in a physical infrastructure business than most people realize. Because the ability to scale that compute, right, and, and bring it to end users, and frankly do so cost effectively, right? The cost of your tokens are going to matter more and more. As our, as our intelligence has all become closer to each other, the ability to scale this up and, and do so to end users where, you know, the cents per token are kind of a determinant if someone's going to buy it is, is critical. So we already started several years ago asking ourselves the question, like, what would it take to move towards this? And then over time, you know, we learned more, we observed more, and then we started acting …

AI assessment note: “Yeah, I think as a foundation model company, you can go two paths, right?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q On that topic of geography, you guys started mostly in Europe, I believe, but it feels like over the last couple of years, you've re-centered the company in large part towards the US. Is that a fair sentiment, and if so, what drove it?

A Yeah, so, We've always been an American company, right? We've been incorporated, uh, headquartered out of the U.S. since day zero, and I think at any given moment, the balance of people would have always sat somewhere between forty-sixty or sixty-forty, one way or another, like, throughout the life of the company. So building it as a global company was critical for us. But the decision we made early on, and I think we spoke about it on the podcast two years ago, is to hire researchers outside of the Bay Area, and particularly originally very centered on Europe. And still today, if we look at where, like, researchers for pool sites sit, they sit predominantly, you know, 95% outside of Silicon Valley. And this offered an incredible opportunity for us. It offered the opportunity to find highly capable, like, highly motivated people who were not in the echo chamber that is the Bay Area. It's one of the most valuable echo chambers in the world. I personally love it. I spent a fair amount of time there. But you've seen and probably realized that there's a bunch of things we've done That we're quite contrarian to the belief of, of that echo chamber. And I think that was partially unlocked by the fact that we sat outside of it and continue to do so. So we hire all across the United States. We hire all across Europe. Uh, we're hiring increasingly more in Asia as well, uh, as we're scali…

AI assessment note: “We've always been an American company, right? We've been incorporated, uh, headquartered out of the U.S.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q news in the last A couple of weeks. First of all, there is a rumored large fund raise up to two billion dollars, where NVIDIA reportedly would be investing up to a billion dollars at twelve billion pre, fourteen billion post. It's a very large round. That seems to be very tied to another big news that you guys did formally announce, which is Project Horizon. What is Project Horizon?

A So Project Horizon is us building out one of the largest data center complexes in the United States. And this comes back to something I said earlier. We talked about intelligence becoming a commodity, right? Our view for pretty much since starting the company over two and a half years has been that there's three layers of the stack that fundamentally are going to matter. It's energy, it's compute, and it's the intelligence built on top. And within this world, if you think that intelligence is going to become less distinguishable between the companies building it, And becomes a commodity, a commodity probably more like oil or cloud compute than like bread at the bakery, is there's two things that matter, your ability to scale it and the cost at which you deliver it to your end user. Now, the ability to scale it was frankly the primary motivating driver for Project Horizon, but cost as well. And it's important to think about what you could do 12 months ago versus what you could do today. At the scale of compute that we were talking about, you know, two years ago in our industry, You could call up a data center colo and say, hey, I want this much compute in six months, and there would be space, like physical space where you could deploy it, or there'd be someone who has the capacity available to you. Today, we're talking about a scale of compute that gets counted in the hundreds o…

AI assessment note: “Project Horizon is us building out one of the largest data center complexes in the United States.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q reinforcement learning to all sorts of different tasks. Yeah. So that's option A. What you're saying is that you're, you're working on an option B, or maybe that's option A, and the first one was option B, but you're working on a, on a different approach where you're going back to internet data, reverse engineer the, the thoughts that went in reaching that conclusion in the first place. Is that?

A It's a perfect way of framing it, and, and we don't think they're mutually exclusive, right? We think that you want to have a model learn how to think and reason as early as possible in its training. And then you want to actually have it learn in the environments to sharpen its skill sets, right? And it's not unlike us, right? Like it's, you will have learned a lot coming out of university, but then we're put in the job, and we're actually doing it, and we're learning from experiences. So the learning from experiences, uh, the reinforcement learning from environments, right, and from experience is effectively like a renewable energy. The density of information in those tokens is not as high as like in a physics book is, Right? We're in a huge amount of density, like, within a small amount of tokens. In experience, it's less, but it's highly valuable. So our view is, is that, you know, we pre-training and predicting the next token on the web is an incredible bootstrap of understanding language and helping us get, you know, to a level of intelligence. Reinforcement learning to learn, RL to L, internally we also refer to as the Bondi techniques, kind of our code name for it. We think we'll push models to a level of reasoning and thought That will happen far earlier in their training than it does today. And then you have reinforcement learning from code execution feedback and then …

AI assessment note: “It's a perfect way of framing it, and, and we don't think they're mutually exclusive”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q mentioned you're building and training your own LLM. And as I was prepping for this, I read that part of the idea is to allow, and I'm quoting, allow your LLM to improve by completing millions of tasks in tens of thousands of real world software projects. And you call this approach reinforcement learning from code execution feedback. Do you want to get into this and explain what that means?

A Yeah. So we, we can kind of look at, uh, the training of our model as, as, as two parts. One is the, you know, the actual training of the base models, everything that we do there, and we can dive into that. And then after we've trained our base models, what do we do to, to really improve them specifically for being more accurate at going from instructions to, to code that is useful to the end end user. We throw this under the umbrella term of RLCF reinforcement learning from code execution feedback. But it really comes down to, to an idea that even started back in, in sourced in, in 2016 or 20 17, which is that it's not enough to show these models, uh, essentially have them learn how to code from showing them massive amounts of examples. And I want to emphasize one thing. Our models are not only trained on code. Our, our data split is about 50, 50 between language and code. Uh, we process web skill data, but our point of view is, is that, well, if, if you treat Pre-training the training of the base model as kind of the, the model reading the textbooks. When does it do the exercises at the end of the textbook?

AI assessment note: “When does it do the exercises at the end of the textbook?”

Answered raw tape D 5 · C 4 · P 5 · Cm 4 4.55

Q And you'll have benchmarks for those. Do you have benchmark?

A We do have benchmarks, yeah. So, so if you think about our benchmarks today, um, And we primarily search it on privately with the organizations. We'll get the commercial part that we work with. But if you think about the Malibu agent as a coding agent, for instance, right now, it sits at a level of like sweet bench verified, for instance, where Gemini two and a half pro was when it came out. Uh, and so it's a much smaller model, uh, but it's really pushed those capabilities because of our work and reinforcement learning. And so we should do a demo after, I think we're actually doing a demo in New York, Uh, publicly because we've only done it privately with enterprises in a couple of weeks at the AI engineer summit. Uh, and so anybody's interested in come and see it there. And I think that will be recorded as well.

AI assessment note: “We do have benchmarks, yeah.”

Answered produced feed D 4 · C 5 · P 5 · Cm 4 4.55

Q And speaking of engineers and switching texts a little bit, uh, one of the things that caught people's attention in the base of both sides, in addition to all of the above is that, uh, you guys decided to build the company largely out of, uh, Europe and France, although I'm sure you have, uh, engineers elsewhere. Can you talk about what, what is the reasoning there?

A Yeah. So in the very early days of the company talking, you know, we started April, 25th officially. And so we were just, you know, a couple of people and we went out and started recruiting and we were primarily recruiting in the Bay area in the first weeks of the company. And one of the things that we had done and our very first hire was a head of people, someone I've known for a very long time. She was an absolute machine at what she does and together with her and one of the early founding engineers, We built this very large list of talent we found relevant, right? This was just like doing the work, like looking at GitHub's research papers, LinkedIn's like all ourselves, you know, and, and we went through about 3200 or 3200 people that we ended up putting on the list. And we had this location column, like think genuinely Google spreadsheet, you know, like where's someone based. And at some point we looked at all these locations and we're like, we should try to group this between Europe and North America. Uh, and now let's filter it on all the people we considered good and all the people we considered incredible. That was kind of, it was like reject good or incredible was, it was our criteria. And then we found a fifty-fifty split between Europe and the US or Europe and North America. And that was surprising to us. And Jason at the time, you know, I think was the one who said,…

AI assessment note: “we found a fifty-fifty split between Europe and North America”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Last question on this, and then we'll transition over to the more familiar software and AI side of the conversation. The, uh, environmental, Impact. That's obviously a question that the whole data center industry grapples with. How do you guys think about it?

A Let's first be honest about this, right? We're using natural gas for the generation of power. And in the world that we are right now, it's a skill that data centers are being built out. The sun, you know, doesn't shine all day, the wind doesn't blow all the time, and so renewables only, in combination with the battery capacity that you would need to power a data center if you want renewable only, it's just not economically viable yet. But the combination of, like, natural gas, uh, with renewables, with batteries, we've got a big best being built out, uh, on the site, so big battery system, uh, allows for kind of an, an optimal point in the middle. Uh, now, when you talk about natural gas generation, you want to make sure that you have your emissions in check. So what you're adding is SCRs. Uh, and so you're.

AI assessment note: “when you talk about natural gas generation, you want to make sure that you have your emissions in check”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q If you will, talk about, uh, that fourth stage a little bit, like, Continuous learning is something that people may have heard about that keeps coming back as one of the next avenues for the whole AI systems to progress. What is it, and how does that work currently, and how does that work in the context of Bullseye?

A Today, there's one thing about foundation models that we have yet to really optimize almost at all, like overall at this time, which is the ability to learn from a single, you know, sample of data. Foundation models today can't do it, and we don't yet have a path In them successfully doing so in a way that, like, really improves it. Internally, we call it the hot stove problem. If you're a kid and you touch a hot stove once, you're never touching it again. Single sample and you're good, right? Foundation models, because of the underlying technology gradient descent, right, is, it's just, it's such a data hungry algorithm. It requires so many samples to be able to navigate that higher dimensionality space to a place where it does something more optimal. And so we've got this Big hole towards AGI of like, how can we get a model to learn continuously from a smaller number of experiences like you and I can. Now, there's a question if that's required to reach AGI. I would actually tend to say that if we take the definition of AGI to be that, you know, foundation models are as capable as you and I to do the vast majority of economically valuable tasks, maybe first behind a laptop and over time embodied in robotics, I would tend to say that it might not be necessary. Uh, but it of course is an incredible thing to add to intelligence because it will massively make it more compute effic…

AI assessment note: “how can we get a model to learn continuously from a smaller number of experiences”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q uh, mentioned. I think, uh, I and, and people may be curious about the current reality of, of poolside, both from a, Model standpoint, product standpoint, and commercial. So starting with a model, so you have, you have three products on your website. You have Malibu, you have Point, and then you have Assistant. What's the current status of those products, and what do they do with products slash models?

A Yeah, so if you go back two and a half years ago, right, one of the first things was Point, which was like the code completion models. Today there are table stakes. We have them, but it's not where the intelligence sits, right? So our Malibu was our first big family of models. And so within the Malibu models, they were really oriented towards originally to be very capable coding assistants. Now they've become incredibly capable coding agents, and now they've also become incredible knowledge work agents. And so within that, those models are right now for their, their size and weight class, best in class. They have become incredibly capable of coding, but they're not yet at what I would define as the frontier. So Frontier today is OpenAI, Anthropic, Google, and XAI. And that's why we have all of this compute coming online to scale up those model sizes. That's our upcoming family of models, Laguna, where we'll have a small, a medium, and a large. So small is finishing training in a couple of weeks, medium actually starts training this week, uh, and, and large starts training as our 41,000 GB 300 come online.

AI assessment note: “Malibu models, they were really oriented towards originally to be very capable coding assistants.”

Partly raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q all of this, fascinating in terms of ambition, it's also fascinating in terms of like what that means for ultimately a young software AI company to become, you know, a major physical infrastructure company. So the CoreWeave partnership Helps, but presumably you have to hire all sorts of people. That's a whole different skill set. How did you think about that? And, um, also the, the financial aspect of it.

A We followed an algorithm. I would probably say most of my career by now. Uh, at least I remember going this back as far as nine years is you have to, by all accounts, avoid the Dunning Kruger effect as a founder, right? Uh, because when you get into something new, like it has been on the physical infrastructure side, you start with learning. You start with reading, you start with like meeting the experts, learning and learning, and as you learn more about a topic, you fall on this point where you start thinking you really know it, but you don't really know it because you haven't hit the real world, or you haven't done it before. And what I found to be one of the best advice I got a very long time ago, whenever you find yourself in that situation, well, this is when you want to start finding the experts to join you, but how do you find the experts? Because you yourself aren't. And here's when you just start interviewing hundred plus people, For every role. And you build a distribution. You start understanding who sits on the right tail end outlier of that distribution. And it turns out, you know, even just with relatively, I would, I wouldn't say our knowledge at this point was shallow because it's been years, like, you know, understanding the space and, uh, getting close to it, but with definitely not the level of it, like an expert, we got to the point where we started being a…

AI assessment note: “here's when you just start interviewing hundred plus people, For every role.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q Part of the idea is that both side will be the anchor tenant, but you'll be renting the facility out to others as well. What is going to be the revenue stream?

A We are not yet at a place where we can be signing 15 year leases on Massive, multi-billion dollar, you know, like, build-outs. Uh, and so, but we are ramping up rapidly where we see a path to being able to scale up compute, uh, at those levels. And so we found a really great partnership here with CoreWeave that was announced. CoreWeave, frankly, is second to none in terms of operating NVIDIA's compute, right? First, I have GB 200 online at skill, and now coming with the GB 300, and, and I've been in this great partner, like, for us in the space. And so, we found a really great hybrid solution with them. So, they are the anchor tenant on our data center. Jointly, we can scale up our compute inside that data center, and any capacity that we choose not to take or not able to take in our space, they can bring to other customers. And this was really an important thing, because we're, we're not a hyperscaler who can make, you know, multi-billion dollar bets for two or three years out. So, we needed to find our way of, of having great partnerships that were kind of win-win situations, Where we could scale up our compute, but didn't need to make a lead time decision years ahead of it. And this has been a fantastic setup because it kind of brings the best of both worlds to the site.

AI assessment note: “any capacity that we choose not to take [...] they can bring to other customers”

Answered raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q You're very focused on software development. Uh, how do you think about the beta lesson? Is that something that you guys worry about?

A If you go back to our first podcast, or we put a post on our website on day zero of the company, we always set the path for AGR runs through software. It is not software. And we laid out this three-step master plan, which was step one, assist people in coding. It's kind of very early days. Step two, allow people, anyone to build software. I think the world is clearly there right now. And then step three, generalize to all domains. And we're now really in that step three moment. And so already our models today have become generally capable across the board. But because we had kept our blinders on for those first couple of years on really software development capabilities as a proxy for intelligence, it allowed us to go really far. It allowed us to push the things that mattered, and our view what was really missing was pushing reasoning. Like improving knowledge, you can improve knowledge outside the model. You can give it access to the right sources of information, but improving this kind of complex reasoning is what was missing. So we're kind of converged on the same point. So today I use our models to, the other day I had to write a sci-fi book while I was reading it, which was a lot of fun, right? Uh, and, uh, my brother uses it in his, in his growth marketing job and prefers it over, you know, using, uh, using other models. And so we're already in that domain now, but softwa…

AI assessment note: “we always set the path for AGR runs through software. It is not software.”

Redirected produced feed D 1 · C 3 · P 2 · Cm 2 2.00

Q So let's jump right into the overall vision for the product. What is it that you guys are building and why did you choose that specific problem to go after?

A So I think it's worth taking a step back and kind of where Jason and I found ourselves at the beginning of this year. We like to jokingly say amongst ourselves that poolside exists because of an open AI induced existential crisis. And we are incredibly deeply grateful to them for it. I felt a very long Help, belief, and you can find all YouTube talks of me, you know, talking about this, that, that it's very like that in our lifetime, neural networks will become capable of learning anything and everything that we are capable of as humans. And earlier this year, Jason and I found ourselves in long conversations, seeing that we were, you know, on a trajectory with everything that had happened, which had GPT and subsequently GPT-IV, which I think was even more impressive. That was putting us on a trajectory where This is very likely to occur in our lifetime and, and probably rather sooner than later. I'm notoriously careful with putting timeline predictions on such big events. While others might say five or 10 years, I like to reserve a little bit right for it for seeing, you know, what's, what's going to happen. But there is something that, you know, a lot of people will, will hear this and say, okay,

AI assessment note: “I think it's worth taking a step back and kind of where Jason and I found ourselves”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.