The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

855exchanges match
855on raw tape
42redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q in that this market would be really important, and you more than others, right, since you actually started a company in it. But then it took some time for the market to really expand to the point where, uh, to your point now, it's, it's this massive use case. People really care about speed of inference and other things. Um, what gave you the conviction back then to do this?

A Combination of, of vision, um, the right co-founders, and a little bit of arrogance, a little bit of luck. You know, we, we saw AI on the horizon as a new workload. And as computer architects, new workloads are opportunity, right? It's very, very hard to, to, to enter in the x-eighty-six world, right? Where there's not, nothing new is happening there, and nothing has happened for generations. But You know, when graphics emerged, you got the discrete GPU and you, you, you got, uh, Nvidia and, and when, uh, when the mobile, uh, compute hit, you, you got arm. And it was interesting that, that not Intel, not AMD, not all sorts of people who you would have thought have been really well positioned to win in that business. They all got no share. And so we knew that, that this new workload would eat a lot of compute. It would require Uh, a new architecture, dedicated architecture, and that ought to be very different. The architecture could not be a derivative of what's existing. Those were our big bets, and they were a hundred percent contrarian, and they turned out to be dead right.

AI assessment note: “Combination of, of vision, um, the right co-founders, and a little bit of arrogance”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Is this set up as a special economic zone in the Philippines or can you tell us more about the details beyond, um, sort of the, the legal side that you mentioned?

A Yeah, absolutely. So, um, it, so right now the, there are two phases to the plan. Um, the first phase is, uh, the State Department taking into custody the, the zone. Uh, we, we are, we are referring to it as an economic security zone because, uh, it is a very unique type of arrangement. The State Department has authorities to take in, um, Land and property into custody, sort of how foreign governments gift the State Department counselors and councilates and embassies. It's very unique to, uh, do a gift of 4000 acres, but fortunately there's no statutory limits on how big or small property can be. And so that's phase one. So right now it's actually diplomatic property, um, that is effectively, you know, uh, governed by the same laws as, as our embassies are. Um, phase two will be the long-term development and build out of the land, and so we are gonna spend, ah, we have two years, a two-year window to negotiate the details with our, ah, Filipino counterparts on the investor protections that will apply to the land, the tax, taxation regimes, and, ah, and all of the different, you know, legal safeguards that investors will be able to benefit from, from for the long term, and the goal is within that two-year window, To actually have a long-term framework that will be multi, um, multiple decades.

AI assessment note: “we are referring to it as an economic security zone because, uh, it is”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q I think there's a lot of actual concern about, like, is the vision we had for the company in terms of its durability or its value, uh, still as valuable as we thought five years ago? And for Roblox, perhaps because you were You know, pointing at something so far away to begin with, it's like pretty clear, you know, I think we have crisped that up over the years.

A I would say two years ago, we were a little less crisp. We would say we're going for a billion DAUs. That's a good spec with this spec I just mentioned. It's a little harder to operationalize around product management teams and that. So we, we took a step, a stepping stone and said, we want to get to 10% of global gaming content. Um, on the way to that. It's actually a remarkably good step. It's about three hundred million DAUs. It's about twenty billion. It's about three X where we are. It can be forensically torn apart. We can see what part is the USA market. We actually have a higher target for us than 10%. Um, and so that has been very clarifying. I think for a software CEO, Always knowing both three X and 10 X is very helpful, because if you don't know 10 X, then it's hard to sleep at night. If you don't know three X, it's hard to plan forensically how we're going to operationalize that, how quickly we're going to get there.

AI assessment note: “I would say two years ago, we were a little less crisp.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q So, uh, genai.mil, I think was like a surprise and how quickly, um, it happened. Uh, uh, can you tell us the story of that and what the, like the goal is?

A So genai.mil, because if you're on a, on a, Department of War network. You, you know, even the unclassified networks are secure, right? And, and different levels of security as you go up. So you still have to figure out how to architect using AI in, into those networks. It's not as simple as going to chatgbt.com and, and signing up because those things are sort of generally prohibited because you have to have restrictions on how to use them. So we had to figure out what are the, the policies, how do you architect it in the network? So nothing goes Back into the pool of data that ChatGPT or Claude or any of these companies have, because we certainly don't want, no one in this country wants our data getting out into the general public, right? Into these models. So you have to architect a different, uh, data flow. So, but we moved really fast. I had some like Databricks engineers, former Databricks engineers, former Meta engineers, former AWS people, Tiger team, 60 days, Uh, good, great collaboration from Gemini, uh, from Google's Gemini, who already works at Department of Warsets, some familiarity with our architecture and our systems, and got that launched to three million people. And we've had over a million people Unix use it in the last 30 days, which is kind of awesome. We've got one third of the enterprise on one model. That's 2.5, not three point oh, three point oh is comi…

AI assessment note: “Tiger team, 60 days, Uh, good, great collaboration from Gemini”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q for or procured. And, you know, that's somewhat different from how you think about innovation in terms of something new comes up. You want to rapidly be able to iterate, to launch it, to deploy it. And so if, how do you think about either flexible budget spend or reallocation of budget over time and how does that relate to the legislative process and how does that impact entrepreneurs ultimately?

A Yeah, it's been, it's been a problem, such a problem that they've given like an awful name to it called the valley of death, which I, which, uh, you'll hear eight ways to Sunday here. And the ways, you know, we've come at it and I've come at it is we have this defense innovation unit, which is rapid contracting, has a billion dollars, has like a reasonable amount of money to do things really fast to get startups off the ground and get them through. We have several other programs that have that same capability, one called AFIT that takes companies that have developed the product, but now need to scale their manufacturing. So there's a different line in the, you know, different point in the process. I have the Office of Strategic Capital, which has two hundred billion dollars in lending authority, um, low cost loans, and that both sends a demand signal to these companies that other private capital crowds around it, equity capital often, and it's, Low cost loans, right? It's treasuries plus a hundred bips. And in the last, I don't know, four or five months, we've done five critical minerals deals, like really fast. Cause that's an, that's a target area and we'll have other target areas. So if you're a company that's in one of the super critical areas, um, we have this huge lending authority there too. So I'm trying to collapse the valley of death by crowding capital around in diff…

AI assessment note: “we're reserving a bit of the budget every year so that we can make in-year budget decisions.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q So there's a breakthrough, uh, result in CHI-II. Can you give us a sort of lay person's explanation of what the result was and, and the model itself and, and, um, what you think is the most valuable part?

A Sure. CHI-II is our latest series of models, which are state of the art across a number of different tasks, but specifically the one we're most excited about is design. And what we've shown is that we can design a class of molecules known as antibodies, Which are some of the most therapeutically interesting molecules as well. These account for close to 50% of all recent drug approvals, and seven of the top 10 best-selling drugs out there are actually, actually antibodies. And so what we've shown with CHI-II is really the ability to design antibodies against targets that one wants to go after in just a small, what we call a twenty-four-well plate, in just 20 attempts. What this means is that we take a target, run our models, ask the model to design a antibody. We then ship that antibody to the lab. We have about a two week validation cycle in the lab, and two weeks later, we see that roughly close to 20% of these antibodies actually bind their targets in the intended way. So Chaito is a major breakthrough for the field. When we set out on this project, Uh, we were actually only targeting a success rate of one percent. That was the company-wide goal for the entire year. Uh, and the reason we set that goal of one percent is that previous attempts at this problem are maybe successful around .1% or even lower of the time. And that's, those are the computational techniques. If you lo…

AI assessment note: “CHI-II is our latest series of models... what we've shown is that we can design”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q browsing and agents, they land on the same, like, two, three transactional use cases that I actually don't think are particularly inspiring, right? So it tends to be, like, Order a burger on DoorDash or something like that, or, uh, I feel like ordering flowers is also like a really common one. Why do you think you came up with, like, such a different set of goals for the agent?

A Yeah, so I think before we focused on taking right actions, which those are examples of taking right actions, we wanted to get really good at synthesizing information from a large number of sources and mostly read-only tasks. That was for a number of reasons. Firstly, just A huge number of knowledge work professions mostly do that, so it would be quite useful for those groups of people. Secondly, I think the overall goal for OpenAI is to create an AGI that can make new scientific discoveries, and we kind of felt that a prerequisite to that is to be able to synthesize information. You know, if you can't write a literature review, you're not going to be able to write a new scientific paper, so felt very in line with the, um, company broader goals.

AI assessment note: “we wanted to get really good at synthesizing information from a large number of sources”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q more about the technology that you're using? Obviously you mentioned you're mixing sort of older school data, modern Uh, uh, image-based data, et cetera, and then you have to kind of data mine it or extrapolate where these potential deposits are. What sort of models are you using? What approaches are you using? How do you think about overall what you're building from a sort of AI and data perspective?

A For sure. So Cobalt's technology is a full stack system for guiding exploration decision making. So there are dozens of different products that work together. And they fit on three, three themes. The first one is sensors, hardware that we have developed that collects new kinds of data about the earth. The second is the data system for taking all of the, all the data that we're collecting, all the historic data in structured data from many different kinds and a huge corpus of unstructured data and getting this all in one system so that we can interact with it systematically. And rather than hunting and pecking through this, we can interact with the whole corpus of data at the same time. Uh, LOMs and other technologies are very powerful for being able to interact with all of these different types of information. And the third theme are models, dozens of different models for making better predictions about where and how to look. And so these models, uh, they, again, they operate at many different length scales. So there's models trained on satellite imagery, uh, or our proprietary hyperspectral airborne imagery. And you've got some rock samples on the ground. And so we can predict from imagery what types of rocks we're going to find at the surface and what the properties of those rocks are going to be. And then what's really exciting is that it's not just that we have a model or a…

AI assessment note: “Cobalt's technology is a full stack system for guiding exploration decision making.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q So first in history opportunity for, um, scientists and entrepreneurs to go work on this dataset and create these virtual cell models. How do you tell the quality of one of these models?

A I mean, the, I mean, the core idea is it's what's its predictive ability, right? And so, you know, you, you take a cell, you perturb it, you can, you can do that either by, you know, from a genetic perspective, you can, you can suppress or, or, or, uh, upregulate genes and then, or, or apply drugs, and then you look at the response. And so the, the measure of the model is how well it predicts the, what we call the, the differentially expressed genes. Um, the, the reality is today the, the best models are, Um, very poor at this. Um, like the, the, the predictive ability of, of the DEGs as we call them is, is in the order of 10%. Um, and one of the conjectures.

AI assessment note: “the measure of the model is how well it predicts the, what we call the, the differentially expressed genes.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Yeah, one of the, I guess one of the other key threads of your life has really been public service, and I believe that you've served under three different mayors in LA over time. You were the runner-up in the twenty-twenty-two LA mayoral race. You were former president of LA's police commission. You've had a variety of roles. Could you tell us a little bit about your public service engagements?

A Yeah, I appreciate it. I'm a big believer in public service. I enjoy it, and I just think it's an important thing to do. So at a very young age of 26, uh, Tom Bradley tapped me to be on the board of the Department of Water and Power, and I then became president. Um, after that, Dick Reardon asked me to go back on the board, so I was at DWP for about 13 years. And then after that, Jim Hahn, uh, the mayor at the time, asked me to head up the police commission because L.A., Back then was having really rising crime. We were losing a lot of police officers, very similar to what's happening today, and to come in and turn LAPD around, and, uh, I did. I brought in Bill Bratton as the chief of police at the time, and we were able to get crime down to levels not seen since 1950. So, I, I'd been very fortunate to have been involved in government service, and reason I decided to eventually run for mayor was because I saw what could be done Especially if you're not beholding to worry about getting reelected, if you're just focused on doing the right thing and becomes a very powerful, powerful mindset to have.

AI assessment note: “Tom Bradley tapped me to be on the board of the Department of Water and Power”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q There's cigarette butts, there's arsons. So I'm a little bit curious how you think about, you know, what do you actually do to prevent things from starting in the first place?

A I think, um, the first place to start in my mind is the utilities. So utilities cause about 11% of ignitions, but those ignitions cause 50% of the damage. And why is that? It's because the same sort of high wind conditions that cause massive fire and negative fire conditions also cause utility equipment failures and lines to fall and stuff like that. For utilities, there's things that they can do. They can do vegetation management, basically trimming trees along power lines. You know, our fund has an investment in a company called Overstory that helps utilities, um, prioritize where they're going to trim these trees, uh, for sort of maximum effect. There's an acronym called PSPS or ESPS, which is public safety power shutoffs or enhanced safety power shutoffs. These are basically de-energizing lines in the, in advance of a, a red flag event, having more sensitive breakers on the, on the circuit so that if there's a short, it de-energizes quickly. They're not very popular with people. People don't like to have their power shut off, particularly before a big wind event, but Um, it can be a big, uh, improvement for safety, and then there's undergrounding and, and things like that as well, but those are very expensive. You know, undergrounding can cost three to four million dollars a mile, and so that, you know, manifests in, in larger, uh, electrical rates. If you can fix just util…

AI assessment note: “the first place to start in my mind is the utilities. So utilities cause about 11%”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q You did mystify one other combination issue between set of corporations and regulators. Um, there's been a bunch of anger about, like, it's very hard to get fire insurance in large parts of California. I think there is a perspective of, like, the insurers are evil for not insuring, um, large areas. You just started an insurance company. Like, can you explain this?

A Insurance rates in the admitted market are regulated, right? And so you have to basically file a rate sheet with the regulator. They say, yep, you're charging an appropriate amount, not too much, not too little, and you go sell that, you know, policy out in the market. But you were not allowed in those rate applications to include the cost of reinsurance, and you had to use historical models, not forward looking models. And so those are like two massive issues because the cost of reinsurance in California has at least doubled, maybe tripled over the last five or 10 years. If you have a cost that just tripled and you cannot pass it through, I mean, that's like sort of a, Crazy thing, right?

AI assessment note: “Insurance rates in the admitted market are regulated, right? And so you have to”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q If you were to do one thing as a homeowner to protect your house going forward, what would you do?

A Um, I'd probably just defensible space. There's sort of this concept of zone zero, which is the first five feet around your, your structure. That should be clear of any flammable materials. It is not in most houses. There's landscaping and azalea bushes or whatever. Um, and, and, but if you can remove that, um, that has a pretty significant impact on your home's Uh, propensity to burn. Um, probably the next thing would just be, uh, a good water source. So that again, it can be attractive for, um, a fire department to, uh, to actually defend your home. But if you have, if you have defensible space and a good water source, um, you're, you're in the top quartile.

AI assessment note: “I'd probably just defensible space. There's sort of this concept of zone zero”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And is it, is it dominated by a specific use case? Like, were there customers that you feel like really represent the pinecone use case well?

A Yeah, a hundred percent. Uh, first text is probably most of what we see. Uh, nowadays models are really good at images and so on, but, uh, text is still the predominant data type. Notion Q&A now runs on, on Pinecone, and they serve essentially question answering with AI, uh, to tens of thousands and probably hundreds of thousands of, of their own customers. Uh, Gong does the same thing with sales calls. Again, serves all of their use cases for all of their customers, and so on. So one of the most common patterns is companies that themselves become trailblazers and innovators with AI, and they themselves hold a lot of their own users' or customers' text, and they want to search over it or generate information on top of it. That ends up being an incredibly common pattern.

AI assessment note: “Notion Q&A now runs on, on Pinecone, and they serve essentially question answering”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q You mentioned you have some favorite applications that have already, like, these sort of assistant capabilities. Can you talk about some of them?

A The primary ways we're finding, um, LLMs useful today at Stripe in user-facing applications is first, automating the writing of code, um, and then second, accelerating information retrieval. And both are proving Really powerful for our users. So on automating code, um, Radar Assistant and Sigma Assistant are two new products that are in beta and rolling out to all users soon. Radar Assistant is really about generating custom fraud rules from natural language. So most folks listening probably have heard of Stripe Radar. It was one of our first non-payments products. It's an ML powered product. It helps identify and block fraudulent transactions. Um, but then in addition to the core radar product, which works generally under the hood without any user provided direction, we have radar for fraud teams, which is about letting users write custom rules. So maybe, you know, you don't have any customers in a given country and you want to block any transactions, um, from IP addresses in that geo. To generate these rules, employees that are users used to have to code up the rules themselves, but radar assistant lets them use natural language to write those rules. Um, it's a little thing, but speed matters a bunch in fighting fraud. You have to work faster than the fraudsters. And with Radar Assistant, a whole range of people in an organization from fraud analysts all the way to less techn…

AI assessment note: “Radar Assistant and Sigma Assistant are two new products that are in beta”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Maybe we can talk about each of those pieces. So, 10 years in machine learning before you were a co-author on the Chinchilla scaling laws paper. You worked on, um, uh, the sort of mixture of experts ideas early. Can you talk a little bit about what your research directions were at DeepMind?

A Yeah. So I come from an optimization background, so my focus has always been, uh, for the last 10 years to make algorithms more efficient and to Uh, use better the data that we have to make, uh, models with good prediction performances. Uh, and so when I arrived at DeepMind, uh, I joined the LLM team that was 10 people at the time, and very quickly I started to work on retrieval-augmented models. Uh, so with a paper called Retro, uh, that I co-led with my friend Seb Borjo, who is still at DeepMind. The point there was to use Like very large databases during pre-training so that we didn't force knowledge into the model itself. And we would tell the model that it would have access to an external memory anyway. And so it was working quite well. We could actually lower the perplexity. Let's say that's what you work on when you make LLMs. Uh, there were some limitations that we, that, that I think the community has started to address quite, quite well. And that was at the time when retrieval methods weren't really Uh, mainstream, uh, now they've become completely mainstream. So that's the first project I did. I worked on, um, on sparse mixture of experts, uh, also quite quickly, uh, because that was related to my topic of postdoc, which was optimal transport. So optimal transport is the setting where you have, I guess, tokens, you need to assess them, assign them to devices, and you…

AI assessment note: “when I arrived at DeepMind, uh, I joined the LLM team that was 10 people”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q Simfluence, which tracks how much smarter your model gets after consuming each example. Can you talk a little bit about the motivation for this work? And then broadly, I think the area of, like, I mean, intersecting with a bunch of your other work, but explainability with large language models is an interesting one. Can we solve that? This seems like something that people generally approach, you know, post hoc.

A Yeah, so the Sinfluence paper is part of an emerging area of research called training data attribution. So the problem there is not specific to large language models. It's really general to any machine learning model. The question that it tries to answer is, given that my model has done a certain thing, which training examples taught it to do that thing? This question goes back all the way to statistics and linear models, but for large language models, it's especially difficult to answer the question because of the Kind of long training process that it goes through and the many forms of generalization that these models have. So in this influence paper, we take a somewhat new approach to the problem. So we say, okay, how would you really justify that this training example was important for doing a certain thing? Well, if you had infinite compute, what you would do is you would take that example and remove it from the training set and retrain your entire model and see if the behavior changed. That's obviously too expensive for most people to do, so what we come up with is a very lightweight model that simulates the training process. It observes a few training runs you've done, tries to estimate the effect of each of the individual examples, and then you as a developer can sit down and say, imagine that I had done this training run, what would the result be? Any sort of simulation…

AI assessment note: “Sinfluence paper is part of an emerging area of research called training data attribution.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q of closed and open source? Do you think the most cutting edge models, the, the giant language models are going to be both? Or do you think like capital will eventually become such a large obstacle that it'll make, um, the private world more likely to drive progress forward? And I know you have plans in terms of how to offset that, but I just love to hear about those.

A The reality is we have more compute available to us than Microsoft or Google. So I have access to national supercomputers and I'm helping multiple nations build exascale computers. So to give you an example, we just got a seven million hour grant on Summit, one of the fastest supercomputers in the US. And like I said, we're building exascale computers that are literally the fastest in the world. Private companies don't have access to that infrastructure. Uh, cause governments, thanks to us are realizing that this is infrastructure of the future. So we have more compute access. We have more cooperation from the whole of academia than all of them do because their agreements tend to be commercial. There's no way that private enterprise can keep up with us. And our costs are zero as well. When you actually consider that, whereas they have to ramp up tens of billions of dollars of compute. So my take is that foundation models will all be open source for the deep learning phase. Cause we're actually got multiple phases now. The first stage is deep learning. That's creating of these large models and we will be the coordinator of the open source. The next stage is the reinforcement learning, the instruct models, flan palm or instruct GPT or others that requires very specified annotation. And that's something that private companies can excel in. The next stage beyond that is fine tuning…

AI assessment note: “my take is that foundation models will all be open source for the deep learning phase.”

Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q That's really cool. And on the computational biology side, I guess there's a few different areas. There's things like protein folding, and then to your point, there's things like MedPalm. Are you thinking of playing a role in both of those types of models in terms of both the medical information? Yes.

A We will release an open MedPalm model. Well, med stable GPT. Um, and then protein folding, we are the, one of the key drivers of open fold right now. So we just released a paper on that, um, much faster ablations than alpha fold. We're doing as well, um, DNA diffusion, uh, for predicting the outcome of DNA sequences. We have BioLM around taking language models for chemical reactions. And that's an area that we will aggressively build because there's a lot of demand from the computational biology side for some level of standardization there. There have been initiatives like Melody and others looking at federated learning, but there is a misalignment of incentives in that space that I think we could come in and fix. And I think that's where we really view ourselves. How can you really align incentive structures and create a foundational element that brings people together? And I think that's where we are most valuable because private sector can't do it that well. Public sector can't do it that well. A mission oriented private company that has this broad base and all these areas could potentially.

AI assessment note: “We will release an open MedPalm model... and then protein folding, we are... key drivers”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Uh, how has that landed with the rest of your customer base?

A So one of the things that we were very careful about was getting, making sure the ecosystem was on board with this, because we do CPU IP, which is really only as good as the ecosystem. The ecosystem of chip people and the ecosystem of software folks and people who build around that. So we talked to just about everybody who were customers and said, you know, how do you feel about this as direction we're going? And surprisingly, we got a lot less pushback than I, than I thought. And the reason for that was the more software that's available in the wild, whether it's proprietary and or open source, benefits the broader ecosystem and the customers themselves. So whether it was Nvidia, Amazon, Microsoft, Google, All people who build ARM-based server chips, they were all on board, and I think the ultimate proof point was when we announced the product last March, we had Jensen, we had Ronnie Boker, we had Amin, we had James Hamilton, you know, all the folks from those customers I mentioned, all saying, congratulations, it's a great thing. So, uh, it's been okay.

AI assessment note: “surprisingly, we got a lot less pushback than I, than I thought.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Amazing. And you took a, I guess, somewhat unconventional startup path or founder CEO path to get here. Um, certainly think of you as founder. Um, what, uh, how did you end up owning the business?

A Well, we didn't start by owning a business. We just started by owning a domain name, uh, that we got from a bankruptcy auction in 2005. Someone else owned the domain name. They were building chess software, but they, um, you know, couldn't pay their, their debts and, and the money they'd raised a little bit of money to try to do that, um, and ended up with a failed business. And so, uh, uh, a friend and I bought the chess.com domain name for 56,000 dollars, um, And we were going to build a chess community. We wanted to be like my space and build a chess community where people could kind of come and, and, and express themselves and learn and share. We bought the domain name and started building.

AI assessment note: “a friend and I bought the chess.com domain name for 56,000 dollars”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Do you think of like, when you look forward, do you think of chess.com as like a game company or an education company or a social network, or, you know, you have like creators and influencers, like a media company. Like, how do you, how do you frame the business besides the mission of get everyone to play chess?

A Well, we're just definitely a chess company. I'll say that, um, the, the, the easy answer is we all love chess. We love playing it. So we're, we're a chess company. But that does have many facets to it. Um, you know, we have said like, we're a gaming company. We're a content company. We're a social network. We are all of those things because chess, the game encompasses all of those things. Um, more recently we have branched out, not that we're like, oh, we're now a games company, but a lot of us do love a lot of different games. Um, so even though chess is like, you know, 99.99% of what we do, uh, we recently launched gambit at gambit.com, which is our poker Um, site and you can play gambit on your phone as well, but we're doing poker in a different way too. We're bringing the chess playbook, which is like poker for ratings. How good are you really? Not just, can you buy the most chips or can you use a bot or, you know, can you steal the most money? Whatever it is, but like, how good are you really? And we're going to roll out other games as well. Um, because we like the playbook of what we've done. So we're cooking up a bunch of other classic games. Um, and we're gonna be launching those out there. We have our take on how we feel like gaming should be done and can be done in like a pure and wholesome and great learning and educational way. So you'll see us rolling out more gam…

AI assessment note: “we're just definitely a chess company... We are all of those things”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I feel like that's a fair statement, right? Um, how do you think about like the timing and sequencing of these very long-term bets and like we just, from a capital allocation perspective, like, When you can invest in these things?

A Yeah, I think it's, it's probably the same with how we invest in a lot of things at DoorDash is everything start out as experiments. I mean, in a way that's a, that was a founding story behind DoorDash. DoorDash was a Stanford college, like dorm room experiment. It started out as a website called politodelivery.com with eight PDF menus and a Google voice phone number. And it was only once we figured out, okay, there's something here, let's turn this into company. And, and, and that's basically, we've kind of taken that philosophy Throughout the past 13 years, and, and we've kind of applied it to autonomy as well, AI as well. I mean, when we first started in 2018, the intention wasn't, hey, let's go spin up this giant robotics program. Let's hire a roboticist, go build hardware. It was really, we put together, it was me and half an engineer's time. It was a skunkworks project. It was an experimentation to go, let's go explore like what's out there. Like we don't even know what autonomy looks like. How robotics is going to impact our space, but let's go explore. Let's go form partnerships. Let's go learn. Let's go experiment. Um, and, and, and, and I think in the beginning, the intention wasn't to build our own robot. Actually, we didn't think we needed to build any of this technology ourselves. We thought, okay, we can just partner up with a bunch of folks. Like, you know, back …

AI assessment note: “everything start out as experiments... it was me and half an engineer's time.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Given the model evaluation cycle and the fact that performance does not Asymptote for many tasks over, um, quite a long period of time. What do you do about that issue? The fact that some of the evals that you would want to run are both beyond the scope of budget or time that's reasonable given the current model release cycle?

A I mean, I think for things like cyber, we've seen, and actually the AISI, um, in their evaluations has shown that the models continue to improve, um, at A hundred million tokens. You know, if you run them for a hundred million tokens, they're still improving at beyond that point. And that can take a very long time to run, but you also do see that like the performance is, is it's not just like a discontinuous jump. It's actually like, you can see the slope of improvement over those hundred million tokens. And so you could, you could probably do some kind of, uh, evaluation up to a certain budget and then just say, okay, well, this is, What we project the performance to look like. And I think this, this, there hasn't been a lot of research on this yet. I actually think this would be a great paper to publish if there's any academics out there looking for something to research. Can you predict what the performance looks like at an inference budget of, let's say, 10,000 dollars only using inference budgets up to 10 or a hundred, or 10 or a hundred dollars?

AI assessment note: “you could probably do some kind of, uh, evaluation up to a certain budget and then”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Let's talk about the larger implications of You know, needing to evaluate these models relative to, um, let's say, like, speed of their reasoning or efficiency versus, you know, token volume, right, um, or dollar budget or whatever, whatever your scaler is. Can you describe some of the larger implications in your essay, including around, um, like, safety evaluations?

A Yeah, the safety evaluations thing, um, it's, it's a bit of an inconvenient truth thing, where Okay, so I guess for background, a lot of the, all the labs have these things called either responsible scaling policies, preparedness frameworks, they go by various names. But the idea is that whenever a model is released, they go through a series of evaluations to measure, are there dangerous capabilities? Um, could these models do things that we're, we, we wouldn't want, um, a bad actor to do? And if the model isn't very capable, then it's no big deal. But if it is very capable, if it could be used, for example, to make bioweapons, then you want to put in mitigations against that. But the question is, okay, well, how do you evaluate whether the model is capable of that? And they have, like, various protocols about, like, how they do these valuations. But a lot of these frameworks were developed around the era of ChatGPT, either before or after, when test time compute scaling was not really, uh, as much of a thing. And it made sense. Like, with GPT-III, you couldn't scale test time compute. Like, if you gave it a budget of ten million dollars and said, okay, well, let's see what GPT-III can do, it really can't do that much, more than what you could do with, like, 10 dollars or one dollar. The preparedness frameworks and responsible scaling policies, they don't really account for the…

AI assessment note: “preparedness frameworks and responsible scaling policies, they don't really account for”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I think I see, I need some guardian, guardian spirits around it. Um, given your deployments today and talking to large enterprises, what is the state of deployment, right? Uh, like how much do you see that's within these more scoped, like studio-like platforms versus, uh, you know, uh, free, free riding coding agents, you know, how, how much are you actually seeing in large enterprises and in different sectors?

A Yeah. So I think right now in our typical enterprise, we're going to see if we break it down to three categories. So we break it down to various SaaS platforms that are typically more low code and where people build agents in this drag and drag way. And they're not really autonomous agents, right? They're kind of the same kind of, I would think of them more as the automations. And then there are, um, first party agents, people are building in their cloud, Potentially because it's an application they want inside the company or even a product they're planning to release to the customers that is agentic. And then the third category is very autonomous coding agents and assistants. Of these categories, I would say roughly at this point, over 50% is the autonomous, uh, coding agents and assistants in the average enterprise. Then probably 45%, uh, is, uh, is those, uh, Uh, low code automations. And the last two percent are really the first party ones that they're building themselves because obviously it's much harder to build effective agents. So, and it's much easier to adopt agents off the shelf or, or build them with low code. So, and that's what we're seeing. And we do see that the autonomous users are also the fastest growing category. So it used to be that only developers and we would see cloud code growing like fire in our customer base. And now we're seeing a cloud cowork grow…

AI assessment note: “over 50% is the autonomous, uh, coding agents and assistants in the average enterprise”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Can I ask you as an aside, actually, just because you, you have for more than a decade believe that this revolution is going to happen. Uh, how much is all of this, um, AI generated coding relevant for Cerebris internally?

A Hugely. I would say that, that, you know, eight months ago, we weren't spending a thousand dollars in engineer on tokens, and we're probably at 25 or 30,000 right now, and it's ripping. I, I think it's not useful for everybody. I, I think that's the truth. I, I think there are some, some people who have sort of the perfect mindset for it, right? And you, you, they are running eight or 10 agents, seven by 24. They've moved their coding, Style to being one in which they govern agents, whether they think about how to QA, so they've got a QA agent running, they think about how to sort of remedy some of the weaknesses in the coding models, right, they're often verbose, they often cut out comments, so they've really thought about, and it's a type of puzzle that's the perfect fit for their mind, and they've gone from being sort of 10 X guys to being hundred X guys. I think the rest of us, myself included, we're sort of Limping along. We, we, we're trying to figure out how, how we can make it work for, for our different jobs, for being the CEO, for being the CFO, for being accountants, for being in marketing. Um, but for a, a small number, it is such a tool. And then the rest, we try and, try and show them what, what, what others are doing, what best practices are.

AI assessment note: “Hugely. I would say that, that, you know, eight months ago, we weren't spending”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And your engineers are just viewing your employees as their customers in some sense then. Is that correct?

A That's right. Our, our team views our employees and our team members in the field as the customer and that feedback loop internally. That's the, the other point is we have a much tighter feedback loop. So, you know, the, the old skunk works thing of you want the engineers and the factory to be co-located so you can have more innovation. That's what we have at Long Lake. So our team members and our engineers are together in the field all the time. I think there's our engineering team. They're probably in 20 different states right now, sitting with team members across our architecture business, across our HOA business, across our HR services, or You know, specialty tax business. And so they're sort of, um, and there's a deep amount of change management that's involved. So this is a lot of, you know, sitting with the team members, understanding their pain points. And so there's a real like solutions orientation of how do we take the pain point? We, and then we build a tool within Nexus to solve it. And that feedback loop is really important. So you get to better outcomes this way.

AI assessment note: “That's right. Our, our team views our employees and our team members in the field”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What do people do with it? What are the main use cases of USDC?

A The, the conception of this obviously is, like, a general protocol for dollars on the internet. And in fact, the whole design is this is like a general purpose, general architecture money, and we actually see it used, you know, from at the very smallest end, like someone who's paying, you know, 25 cents for a digital object in a digital game that's built on a blockchain, that would be like one end, or even now we're starting to see, and we'll come back to this topic, I'm sure, you know, AI agents that are paying for, uh, the output of essentially the AI tokens of another AI agent, and they're, you know, spending, again, Just, you know, uh, a dollar, 50 cents, 20 cents, et cetera. So super tiny transactions at one end, all the way to the largest electronic trading firms in the world that do huge amounts of, of capital markets activity who are, you know, settling multi-hundred million dollar transactions. And the powerful thing is it's all the same. Just like, you know, if I send you an email, uh, you know, the, and my email's like, hey, this is what I had for breakfast. The payload of that is the same as if I sent you an email that had, like, a CIA dossier attached to it. Like, USDC doesn't care. You know, so as a, as a general architecture, it can be used across a huge range of things, and we have everything from merchants in Stripe and Shopify that are using it, to Visa actual…

AI assessment note: “merchants in Stripe and Shopify that are using it, to Visa actually using it”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And what did you work on specifically at OpenAI?

A Well, so the goal was we need to come up with some productionization of GPT-IV. So we, OpenAI had GPT-IV. It was pre-trained, and there were some, like, um, post trains on it. And there's questions about, like, how do we turn this incredibly powerful model into products? And we're all spitballing ideas, like writing bot, uh, coding bot. You know, very natural at the time. Some of our least interesting ideas were a meeting bot, so it would just sit in a Google Meet, take notes, and then send out, like, to-dos after. But John Schulman was very opinionated. He's like, we think we should keep it very general. Let's do a chat bot. And that became a large part of the effort, um, for those few months.

AI assessment note: “we need to come up with some productionization of GPT-IV”

page 1 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.