The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

855exchanges match
855on raw tape
42redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q The obstacles to that right now are some new model advancements. Is it just building out some other core technology? Is it just the UI? Like what, what is keeping that from happening right now?

A Yeah, I think the first thing is the model, uh, the full O three model that's not available yet, that OpenAI showed, um, uh, as part of the ship miss, uh, uh, right before the holidays. We're going to see, you know, improved reasoning, and I think it's as the models get better in reasoning, uh, we're going to get closer to a hundred percent of this VBench, which is that benchmark out of 12 repos, um, uh, open source Python repos, um, a team in Princeton identified, uh, 2200 or so, uh, issue pull request pairs. Effectively, all the models and agents are measured against. And so that's number one, you know, the, the model and the agent combination. I think the second piece is just, uh, figuring out what's the right user interface flow. Um, if you think about the workflow of a developer, right, you, you have an issue that somebody else filed for you, you know, user or product manager or something that you filed yourself. Now, how do you know whether you should assign co-pilot to this, um, the agent to it, um, or, or whether you need to refine the issue to be more specific, right? It's crucial that the agent is predictable, that you know that this is a task that the agent can solve. If not, then you need to steer it. So steerability is the next thing you need to either, you know, extend the, The definition, um, uh, or the agent needs to come back to you and, and ask you additional …

AI assessment note: “I think the first thing is the model... I think the second piece is”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q If, uh, you all could propose a magical adoption tactically of some policy or action to the current administration, what is the first step here? It is the, you know, we will not build a super weapon and we're going to be watching for other people building them too.

A And so I've sort of been alluding to throughout this whole conversation, like what would the companies do? Like not that much. I mean, add some basic anti-terrorism safeguards, but I think this is like pretty technically easy. This is unlike refusal for other things. Refusal robustness for other things is harder. Like if you're trying to get it like crimes and torts, that that's harder because it's, it's a lot messier. It overlaps with typical everyday interaction. I think likewise here, the, the asks for states are not that challenging either. It's just a matter of them doing it. So one would be the CIA has a cell that's doing more espionage of other states' AI programs. So that way they have a better sense of what's going on and aren't caught by surprise. And then secondly, maybe some part of government, like let's say cybercom, which has a lot of cyber offensive capabilities, um, gets some cyber attacks ready to, um, disable, um, other data centers in other countries if they're looking like they're doing something, running a, or creating a destabilizing AI project. That's it for the deterrence for nonproliferation of, of AI chips to rogue actors in particular. I think there'd be, um, some adjustments to export controls. In particular, just knowing where the AI chips are at reliably. We want to know where the AI chips are at for the same reason we want to know where our fissi…

AI assessment note: “one would be the CIA has a cell that's doing more espionage”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q want to change tax for our last couple of minutes and talk about evals. Um, and it's obviously very related to, uh, safety and understanding where we are in terms of capability. Can you just contextualize where, where you think we are? Uh, you came out with a triggeringly named humanity's last exam eval, and then also Enigma, um, like why are these relevant and where are we in evals?

A Yeah, yeah. So for context, I've been making evaluations to try and understand where we're at in this, um, in AI for, uh, I don't know, about as long as I've been doing AI research. Uh, so previously I've done some, um, data sets like MMLU and the math data set. Before that, before ChatGPT, there's, Things like ImageNet-C and, and other sorts of things. So Humanity's last exam was basically an attempt at getting at what's the, what would be the, um, end of the road for the evaluations and benchmarks that are based on exam-like questions, ones that test some sort of academic type of knowledge. So for this, we asked professors and researchers around the world to submit a really challenging question And then we would add that to the, the data set. So it's a big collection of what professors, for instance, would encounter as challenging problems in their, in their research that have a definitive closed ended objective answer. With that, I think the genre of here's a closed ended answer where it's, you know, multiple choice or a simple short answer. I think that genre will roughly be expired when performance on this data set is, uh, near the ceiling. So, and when performance is near the ceiling, I think that'd basically be an indication that, like, you have something like a superhuman mathematician, um, or a superhuman STEM scientist for, in many ways, for when they're, when closed-…

AI assessment note: “Humanity's last exam was basically an attempt at getting at what's the... end of the road”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Flagship, just in terms of ambition and scope, doesn't just do therapeutics. You've done ventures in nutrition and agriculture and climate in these different areas. Um, how do you think about opportunities to extend into beyond medical biotech?

A We're very careful in that regard, but, but where we have a core advantage, whether it's intellectual property we've created, or maybe a daring that comes from not knowing enough about the space, we will kind of venture into it and You know, it's, it's for us. The first thing we do in a space informs the next five things we do. And then if all five things don't succeed, uh, we'll kind of say, you know what, maybe we can't get, you know, paid for the innovation in this space. Our methodology is all about trying to bring to life today what might otherwise exist five years from now. Uh, not everything that will exist five years from now will be valuable. So on top of that, we've got to come up with something today that's also going to be valuable. And that is not for every sector. So for example, we worked for many, many years in renewable energy. Turns out that, you know, one of the most advanced ways we could make carbon neutral liquid fuels was to engineer photosynthetic bacteria that usually grows in the depths of oceans. Really cool bacterium that sees a little bit of photons. We engineered these things to make diesel. It could literally consume CO two. And, and make diesel fuel and secrete it. And people thought it's impossible and we did it. But, and then we created reactor systems, dirt cheap to do this in. Next thing you know, we go out and this was in the 2008 to 2012 ti…

AI assessment note: “where we have a core advantage, whether it's intellectual property we've created”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q three clinical asset type company. How do you think about this trade-off between early investment and platform versus clinical assets? Because, you know, in a, in a different field, as you and I were talking about before, investor climate and the, the macro changes around you too. Not every biotech survives. And so, you know, how do you think about that having been through several cycles of that investor climate?

A Let me answer how, how we think about it at Flagship, and then let me answer it about how one should think about it because, you know, we, we of course only represent one subset. How we think about it is that because we go to far out places looking for undiscovered value, the notion that you do that to come back with one asset is the definition of insanity. Because if you're going to do that, You might as well bet on well-known proven technologies, just a slightly different version. And you could then bet on an asset and hope that your lottery ticket gets, gets pulled. Sorry for being a bit rash, but that's how I view it. But if you're going to go and do RNA for the first time, DNA for the first time, gene writing, gene editing, computational proteins, you need to diversify that because you don't know which ones of these are going to get knocked out for reasons that have nothing to do with the underlying technology. Every single of the 110 companies we've been involved with for 25 years are a platform. Every single one. There's no exception to that. So we embraced from day one, but for a different reason, namely we wanted to go beyond adjacencies, beyond kind of the reasonable zone into unreasonable things. So that's why we do it that way. Now, if you said, well, why does, why does everybody not do it? Or most people not do it that way, even though they want to, as you said. An…

AI assessment note: “Every single of the 110 companies we've been involved with for 25 years are a platform.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q You did amazing work on Cruise, and then you decided to start another company, which I always think is a really brave endeavor because anybody who's been through multiple startups knows how painful and terrible it is. Could you tell us a little bit more about the impetus behind the bot company and what you were doing there?

A Yeah. Well, we talked about this a little bit when I was, uh, making that decision, what, what to do next. And, uh, I did some soul searching and determined that I'm, I'm just a builder. I'm, I like building things and I was sitting on the sidelines or helping other entrepreneurs or doing something else, um, I think would be fun, but not quite scratching that, that same itch. Uh, I'm 39. I feel like I got at least one more startup in the tank. So the question became of like what to do. And I look back on my career. This is my, I guess, Depending on how you count, like third major startup. And the first one was, you know, Twitch and Justin TV, and that was straight out of college. And that was just doing anything, doing a startup and trying to make it work, um, was the priority. And that ended up being video games and entertainment. The second time around for a cruise, after doing entertainment, I decided I want to focus on impact. So like, what's, what's something where we can use technology to meaningfully improve people's lives? Self-driving cars, they save lives, they give you tons of time back. That was like squarely in the, you know, impact category. Third time around, um, I, I definitely care about impact, but also fun. So it's like working with people I like on problems I like, uh, really challenging technical problems and, and building amazing products. And so we're bui…

AI assessment note: “we're building home robots. And, uh, the impact side of that is”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q severity and frequency of fires across the city and state have been escalating, with this one really becoming, uh, one of the costliest in history on multiple levels, a human level, as well as the financial level. Uh, what is behind this trend? Are wildfires getting worse due to climate change? Is it other factors like Policy, deterring infrastructure, negligence, like what, what do you think are the causes here?

A All the above. I mean, listen, I don't, I don't think you can doubt that there's an impact to climate change, so let's just say that's a given. But, uh, the Palisades fire was fueled by 40 years of brush that was never managed, and so you have an enormous amount of brush in those hills behind the Palisades, and when you had the winds coming up, uh, the city was not prepared adequately to deal with it. The fire department was not properly deployed to deal with it. To be in the second largest city in the United States and to run out of water and have fire hydrants empty is, Completely insane. So there was a series of things, but it really starts with, frankly, incompetent leadership that wasn't prepared for something like this, like clearing brush, like making sure the reservoirs were full. There's a whole reservoir that was empty. That was drained because they wanted to have repairs on it. Well, we had a fire in Malibu, 10 minutes from the Palisades fire, three weeks before. That's probably a pretty good sign that we needed to be prepared, and we knew the winds were coming. So unfortunately, I think it's border negligence. Um, there's no doubt in my mind, but it's clearly bad planning, bad leadership, and maybe you couldn't have prevented the fire. I'm completely convinced you could have substantially mitigated the fire. We've lost the equivalent of two Manhattans in terms of la…

AI assessment note: “All the above. I mean, listen, I don't, I don't think you can doubt”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q bass, and, um, you know, you talked a little bit about water availability in the reservoir. Uh, can you tell us a little bit more about what happened there? Because I've heard arguments that, you know, the house is burned down, and therefore there was, uh, open pipes, and water was just leaking out, and that was a reason that there wasn't enough water. Do you think that's just... No.

A You know, I just got off the phone with an elected official in Congress that was finding every excuse in the world Uh, why it wasn't anybody's fault, which I think elected officials are really good at. I think they have a PhD in that. But no, listen, I was on the phone with my team. A senior member of the rapid response team that we have was up there and embedded in the command staff. And at about, I think it was a little bit after 10 o'clock or 10 30, I get the call. The, the hydrants are empty. We're not getting water. It had nothing to do with broken pipes. It had everything to do with the reservoirs not being filled. Everything up there is gravity flow. So those reservoirs draining are coming into the hydrants, and there's protocols should be in place that's keeping those reservoirs full, but the largest reservoir was empty. And that's, that was a huge problem.

AI assessment note: “It had nothing to do with broken pipes. It had everything to do with the reservoirs”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q was a 17 and a half million dollar budget Terry cut for the fire department, uh, by the incoming mayor. And then there was some claim around a lack of training, lack of personnel, and other issues as well, in addition to a lack of water. And so it sounds like it was sort of this multi-pronged issue beyond just access in. What do you think about those different factors?

A Yeah, I mean, certainly water was an issue. Um, you know, there's been a lot of like kerfuffle on Twitter where people are claiming that because of the water bill that didn't, wasn't acted on, that that somehow was related to this. I don't think that is actually true at all. Like really, The water issues were related to there's three, one million gallon tanks, um, as houses burned down, the plumbing opens. If you've got like a five eighth or half inch plumbing connection into a house, that's going to run at 20 or 30 gallons a minute. You have 500 of those going, you've got really significant water flow and pressure drop. And so those water systems are just not really built for sort of these urban firestorm, uh, environments. So, you know, there's plenty of water in the state of California, at least sometimes, um, and there was here, but, um, You know, I think the water issue was much more about the scale of the incident. Access was an issue. In terms of staffing, that's not my area of expertise. You know, in the scale of the LA Fire Department budget, I don't think seventeen million dollars is really all that much. Um, in California, we have some of the best resourced fire agencies in the world. Cal Fire's budget is, I think, about triple what the US Forest Service budget is. In my opinion, we are relatively well resourced, at least compared to most of the other areas. I think …

AI assessment note: “in the scale of the LA Fire Department budget, I don't think seventeen million dollars is really all that much.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q it may even be worth also defining like, what are the products that Sierra provides today for its customers? And then where do you want that to go? And then maybe we can feed that back into like, what are the components of that? Obviously, folks are really emerging as a leader in your vertical, but it'd be great just for a broader audience to understand what you focus on.

A Yeah, sure. I'll just give a couple of examples to make it concrete. So if you buy a new Sonos speaker or you're having technical issues with your speaker, you get the dreaded flashing orange light. You'll now chat with the Sonos AI, which is powered by Sierra to help you onboard, help you debug, whether it's a hardware issue, a wifi issue, um, things like that. If you're a SiriusXM subscriber, their AI agent is named Harmony, which I think is a delightful name. And, uh, it's everything from Upgrading and downgrading your subscription level to, if you get a trial when you purchase a new vehicle, speaking to you about that. Broadly speaking, I would say we help companies build branded customer facing agents. Um, and branded is an important part of it. It's, it's part of your brand. It's part of your brand experience. And I think that's really interesting and compelling because I think just like, you know, when I go back to the proverbial, 1995, you know, your website was on your business card. It was the first time you had sort of this digital Presence and I think the same novelty and probably we'll look back at the agents today with the same sense of, oh, that was quaint. Uh, you know, I remember if you go back to the Wayback Machine, you look at early websites, it was either someone's phone number and that's it, or it looked like a DVD intro screen with like lots of graphics. …

AI assessment note: “we help companies build branded customer facing agents.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So maybe that's a good segue into how you got started working on this, because you've had this idea for a long time. As you said, you started on, um, uh, space applications earlier. Can you talk a little bit about your background and, you know, the original scientific idea and how you thought it would be applied?

A Sure. Uh, so my background is in material science and electrical engineering. I obtained a PhD in electrical engineering with a minor in material science and device physics from Cornell. Uh, and in my PhD, Uh, I focused on bringing together very dissimilar materials, uh, in such a way that one plus one equals 10, okay? Uh, and, and, you know, for example, silicon, very well known, ubiquitous material that's ushered in the current modern era that we have today, uh, but then there are other materials, plastics, other types of semiconductors that don't actually Do as well as silicon, but they have their own strengths. Um, and so for my PhD, I looked at ways of trying to bring together say the optics world with electronic silicon and merging them together such that the overall system is incredibly powerful. That philosophy I've brought to Akash when I started Akash with Ty in 2017 to try to do the same thing. Um, I often found personally Uh, actually, I think it's a very good, good metaphor for, for, for humans, how we interact together. When you bring different people that have different strengths, the combination can be incredibly powerful in ways that excel and exceed the simple summation of the parts, and that's exactly what we do at Akash, where we bring, uh, artificial diamonds, well known as the most thermally conductive material ever grown in nature, Or ever to occur in nat…

AI assessment note: “my background is in material science and electrical engineering. I obtained a PhD”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Ty, are you the, are you the silicon or are you the diamond here?

A Um, I'm actually a little bit of both. I'm the silicon carbide guy. Uh, my PhD was on silicon carbide. What that taught me when I went into the business world, uh, working for a company Cree and Woolspeed, is that, uh, Cree and Woolspeed developed very good silicon carbide materials level technology, and applying this technology to, uh, any sorts of, uh, systems like radar systems, or a power electronic system for EVs, or light emitting diodes, One of the things you learn is that when you have a materials level advancement, as the person who has that material level advancement, you really have to make the system to convince people that you have the solution. If you just go to someone who is making, let's say, a car, and you say, hey, I've got this great silicon carbide, uh, diode, a shocky diode, or a MOSFET, they'll say, all right, great. You know, I already using a, uh, a silicon IGBT. Um, if you meet their price, I'll put you in. And you say, look, if you, if you put my part in, I think I can increase your range, 200 miles. I can increase it 40%. They'll say, okay, yeah, sure. However, if you make the car, okay, you find a partner to get your part into the car, uh, or you make the box, you actually make the MOSFET, um, you make the module, the power module, the farther you go in the system, the better chance you have of convincing the customer that you have the solution. So …

AI assessment note: “I'm actually a little bit of both. I'm the silicon carbide guy.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I'm gonna ask a silly question, but you've mentioned it several times. Um, uh, you're growing diamonds. How does that process work for the form factor that you want? Assuming, you know, the vast majority of our audience has only ever heard of the, the concept of growing diamonds in the, uh, you know, realm of like jewelry.

A It's really no different than, uh, than growing other, uh, other semiconductor, uh, materials. Uh, if you're growing silicon or, um, or silicon carbide or gallium arsenide or indium phosphide, any of these, uh, electronic substrates, uh, you start with a seed crystal and then you use a, uh, typically a, you know, some sort of process chemical vapor deposition, uh, to grow out From that crystal to grow, you know, perfect single crystal material out from that crystal. And that's the same way. Uh, that's the same way you grow diamond. Um, diamond is just carbon, right? So, uh, so you take a seed crystal of perfect carbon of diamond and, uh, and you, you, you use a plasma to, to grow the diamond in a, in a reactor. You know, it tastes, uh, very high, very high temperatures, you know, very high pressures to do this. But, uh, but it's essentially a similar process to growing, uh, to growing silicon or silicon carbide vapors.

AI assessment note: “you start with a seed crystal and then you use... chemical vapor deposition”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q There's a lot of code generation tools out there. You can do this, you know, directly, um, in the core model products as well. What do you think people are finding special about Bolt?

A Yeah, totally. Yeah, what's special about Bolt, and it kind of comes to the origins of our company, but, you know, in short, we've written an operating system in WebAssembly that can, like, run in your browser, and that's really important, because if you want to run dev environments, uh, You need to be able to install arbitrary packages and run different tool chains, right? Whether it's Next.js or Veed or anything else. It's very complicated and expensive to typically do this if you're going to use servers. So it's very valuable to like do it in the browser because it's extremely fast. There's no latency. You're not paying by the minute, you know, for some cloud. What we've done, um, has kind of married these frontier models with this technology we've been making. Um, and when you kind of look at the other stuff in the market, there, there's, uh, you know, like a cloud artifacts is, you know, Probably one of the first things that, uh, that hit the market that did a really good job of this, where you could say, Hey, build me a UI and it will like do it. The problem comes when you actually want to build stuff that's more meaningful. Like it's very good if you're saying, Hey, like, yeah, I use that, you know, Claude, uh, you know, every week you're just kind of generating graphs based on numbers or whatever. Very good for that sort of use case, but if you want to say, hey, create …

AI assessment note: “we've written an operating system in WebAssembly that can, like, run in your browser”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What's your favorite use case? What are people building that's cool?

A What's really cool to me is Folks are actually able to build real world products. And then so, you know, we've been online for just under two months now, and we've already had the first startups launch out of this thing. You know, they've used Bolt to build their startup and are making money, like, you know, charging on Stripe or whatever have you. Um, so a couple of examples, uh, just off the top of my head. Um, one is from, uh, this gal in Thailand. Uh, she's a PM in a software banking company and, uh, her company is, uh, viralhooks.ai. And so she launched this project, um, you know, by herself. Um, on the side, just moonlighting it. And the product is actually pretty cool. So the general idea is, uh, you know, when you make like a TikTok or something, I'm not a TikToker, but you know, I've had aspirations. Um, when you make a TikTok, you need, you need to have like a viral hook to kind of get people to keep watching. Right. And so she's actually, uh, trained up some, uh, uh, models, you know, for open AI or whatever have you to actually help you write, uh, great viral hooks for your videos and kind of reverse engineer how they're great creators have done that. So as you can go check it out, viralhooks.ai. And so what, what was kind of mind blowing and it's a beautiful site, like awesome product. And, um, what was mind blowing about this is a week before we launched Bolt, uh,…

AI assessment note: “one is from, uh, this gal in Thailand... her company is, uh, viralhooks.ai.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q was working on developing their point of view and trying different ideas to refine it for five years, and then they made Notion. Um, and you guys have been working on this for, you know, five plus years as well. Like, uh, Can you talk a little bit about like the, the origin of StackBlitz and, and you and Pi and like when, uh, when you decided to do both?

A Yeah. So I, I co-founded StackBlitz with one of my childhood best friends. His name's Albert Pi. He and I grew up in a settler of Chicago together. Um, and when we were 13, uh, we, we, we, we had ideas. We and I were always very interested in computers. We were building PCs and we wanted to learn how to, to write web applications. This is like the mid 2000. Um, so like for our 13th birthdays, we asked, For, you know, the O'Reilly books. Cause they're like 200 bucks a pop instead of an Xbox, you know, and then we learn how to code together. Um, and, and really, and it was, it was painful. I mean, get to, to, at that time, there's not like Code Academy and all the stuff that's for free online. Um, there wasn't really online communities around these things, but he and I really wanted to be, we thought we had cool ideas for products or whatever. And, and we really wanted to build them and launch them. I think, and, and that's really, I think that's, that's why, and, you know, we've been building stuff together for You know, 1520 years. Um, it's been about that, you know, coding was really a way to, you know, a necessary part of, of how you bring these things to life, you know? Anyways, um, fast forward, Albert and I have done a couple of different startups over the years. Um, but back in 2016, 20 17, we had this realization that browsers had gotten really powerful. Like we've been …

AI assessment note: “I co-founded StackBlitz with one of my childhood best friends. His name's Albert Pi.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What are people trying to do? Can you give us a flavor of some of like the biggest use cases you see in the enterprise?

A It's super broad. Uh, so it spans pretty much every vertical. I mean, the common things are like Q&A. So speaking to a corpus of documents, for instance, if you're a manufacturing company, you might want to build a Q&A bot for your engineers or your workers who are on the assembly line, uh, and plug in All of the, the manuals of the different tools and diagnostic manuals for common errors in parts, and then let the user chat to that instead of having to open up a thousand page book and try to find what they need. Similarly, Q and A bots for the average enterprise worker. So plugging in your IT FAQ, your HR docs, all the things about your company, and having a centralized chat interface onto the knowledge of your organization so that they can get their questions answered. Those are some of the common ones. Beyond that, there are kind of specific functions that we power. Um, a good example might be for a healthcare company. They have these longitudinal health records. Of patients. And that consists of every interaction that that patient has with the, the healthcare system from visits to a pharmacy, to the different labs or tests that they're getting, uh, to doctors visits, and it can spend decades. And so it's a huge, huge record of someone's medical history. And typically what happens is that patient will call in and they'll ring up the receptionist and be like, my knee hurts. I…

AI assessment note: “the common things are like Q&A. So speaking to a corpus of documents”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q One of the things that was an interesting tidbit from your announcement was, you know, there's a series of problems solves one problems within minutes, three days to solve other problems. Um, is there a way you can characterize like the search space for math overall, or what alpha proof is better at, um, in terms of domains and others, like types of reasoning within math?

A The search space in math is quite large compared to like something like chess or, or other board games. So, you know, uh, Some people might think of writing math proofs as like picking from a bag of known tricks, uh, at each step. But in fact, like, um, there are many proofs where like you've got to come up with some non-trivial constructions, like you've got to invent a function out of thin air, or you've got to like, um, come up with like some way to manipulate. Even if you're just rewriting an expression, uh, you can rewrite it in infinite ways. And, uh, there's only like a few ways you could rewrite it to actually make progress on the proof. Sometimes thinking about some novel problem requires like decades of theory building and approaching it from an entirely new perspective. Uh, to arrive at, um, at the angle that helps you solve it. Uh, and so in that's, in that sense, uh, you know, the reason why there are like a lot of, uh, math problems that are like very simple to state, but have stood, you know, stood the test of time in that they've been unsolved for centuries even, is that, um, the search space is not sort of easy to navigate, um, in, in most cases. I think what AlphaProof is good at amongst the IMO categories is it's, it's largely good. So the IMO problems have come in four categories. So there's, uh, algebra, Uh, number theory, combinatrix, and geometry. Um, the…

AI assessment note: “The two that it's strongest at are algebra and number theory.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q conversations, you'd talk to them and they'd say, oh, we don't need this. And then three months later they'd call and say, okay, we really need this. And it was always roughly the same timeframe. Are you seeing any common patterns today in terms of, okay, companies that are now a year or 18 months into their journey using LLMs, like they always have Have the same thing come up?

A There's a couple things. So one is companies that are fairly deep into their journey. They, they have like one or two North star products that are pretty mature and they're trying to figure out how to get those products to the next stage. Probably the most consistent thing I've seen is companies kind of walking back from the, uh, illusion that totally free form agents will solve all of their problems. So I think maybe like two or three months ago, Many of the pioneering companies went way down the agent rabbit hole. Um, and they kind of realized like, wow, this is actually not, this is, this is not, not the right, um, approach. It's so hard to control performance. Um, the error rates are really high and they compound really quickly. Um, and so, you know, most of those companies have kind of walked back and, um, tried to, to, to build a different architecture where the control flow is, is actually managed deterministically by their code. Um, but they, Um, make LLM calls kind of like throughout the entire, uh, architecture of the product. Um, and so that's, that's probably the biggest thing that we're seeing now is, um, I, I don't, I don't know if there's a good term for it yet, but maybe this kind of pervasive, uh, AI engineering throughout a product rather than trying to shove everything into the, you know, um, while loop of an agent.

AI assessment note: “most consistent thing I've seen is companies kind of walking back from the, uh, illusion”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Can you describe some of the key challenges in, like, taking, you know, the capabilities of Foundation Monster and then making them work in, like, the company agent context?

A One of the techniques, and I think you all probably talked about on your, your podcast that's very common today is it's what's called retrieval augmented generation. Um, and essentially what that means is, uh, you take a large language model and rather than using the model and its innate knowledge from the pre-training process to, um, emit answers, you combine that model with a database of content and you say, Use the content as a source of truth, and you ask the model to summarize, uh, selected content from that, that database. And that's kind of a roundabout way of saying if you can ground the agent and knowledge that you provide it, but also you can take off the shelf models and, and integrate it with proprietary business data. So it's a really popular technique right now. I would say that's a really exciting area, but what we found in practice is that broad category of technology investment is woefully Uh, insufficient for almost any meaningful customer experience. If you think about, you know, all of the interactions you've had with brands that you care about, what percentage of those conversations were asking questions? Probably none of them. It's all about taking action, right? It's upgrading or downgrading a subscription. It's returning an order. It's a warranty exchange. It's a, you know, filing a claim with an insurance company. All of those are not only not simply an…

AI assessment note: “All of those are not only not simply answering questions, but also taking action”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q piece of it is if you're actually running all your application, all the data, everything else on one of these cloud providers, pinging out to a third party service just adds latency. So you add the round trip, you add a second sort of buying behavior around, um, approval, budget, Uh, security, et cetera. So do you think it's just going to roughly consolidate around the clouds plus or minus?

A I do think it will roughly end up the cloud of providers in partnership with the big research labs, which is roughly the current, uh, you know, landscape. I'm not sure I completely agree on the security and latency front. It's possibly true. It was interesting. I think that, you know, most companies, most large enterprises now use multiple cloud providers. Uh, most of them use software as a service and don't necessarily care where it's hosted as long as the security and reliability requirements are met. And there's obviously some exceptions to this, but I think thanks to 20 years of software as a service, people sort of evolve their expectations to not ask, you know, where do you get your power? And just say, what is your, you know, SLA for this service? And I think that's probably a positive trend. So I do think there's probably meaningful latency and security issues to overcome, but I, that all being in the same substrate, I'm not sure I make that leap. I might be wrong. I just, you know, I, I view the evolution of software as a service having evolved that, but going back to my history rhyming point, I think you'll have a relatively small number of foundation model builders and doing pre-training. I think there'll be a market of tools companies, um, you know, well, great one in AI might be scale AI. Um, you know, Snowflake was a great example in cloud that might also be an ex…

AI assessment note: “I do think it will roughly end up the cloud of providers in partnership”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And I think, you know, it seems like the one other thing or transition that's happened is basically a move from a lot of, um, sort of, uh, edge case designed heuristics associated with it versus end to end deep learning. And that's what other shift that's happened recently. Do you want to talk a little bit about that and sort of what that?

A Yeah, I think that was always like the plan from the start, I would say at Tesla, as I was talking about how the neural net can like eat through the stack, because when I joined, there was a ton of C++ code, and now there's much, much less C++ code in the test time package that runs in the car, because, uh, there's still a ton of stuff in the, in the backend, uh, that we're not talking about. The neural net kind of like takes, uh, takes through the system. So first it just does like a detection on the image level. Then it does multiple images. It gives you prediction. Then multiple images over time give you a prediction, and you're discarding C++ code, and eventually you're just giving steering commands. And so I think Tesla is kind of eating through the stack. My understanding is that current Waymos are actually, like, not that, but that they've tried, but they ended up, like, not doing that is my current understanding, but I'm not sure because they don't talk about it. But I do fundamentally believe in this approach, um, and I think, um, that's the last piece to fall, if you want to think about it that way. And I do suspect that The end-to-end systems for Tesla in, like, say, 10 years, it is just a neural net. I mean, the videos stream into a neural net and commands come out. You have to sort of build up to it incrementally and do it piece by piece. And even all the intermedi…

AI assessment note: “how the neural net can like eat through the stack, because when I joined”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And do you think that's just like a quality and quality control issue for those systems? Or do you think it's just some form of complexity with some failure rate per component that's rising? Or like, what do you think is sort of the driver of that?

A I think it's more of the complexity has grown, right? And we're in a different regime now, you know, I think that it's fair to say, so maybe stepping back again to definitions, we throw the term large around a lot in the ecosystem. I guess one question is what does large mean? And one useful definition of large that I think roughly corresponds to what people mean when they invoke the term is that a large language model, you enter the large regime when the, essentially the amount of compute necessary to contain even just the model weights starts to exceed The capacity of even a state-of-the-art single GPU or single node. I think it's fair to say you're in the large regime when you need multiple, you know, state-of-the-art servers from NVIDIA or from someone else to even just contain the model, you know, just run the, basically run the training or definitely even just to contain the model. That's definitely the large regime. And so the key characteristic of the large regime is that you have to somehow orchestrate a cluster of GPUs To perform a single synchronized calculation, right? And so it becomes a bit of a distributed systems problem. I think that's one way of characterizing the large regime. Now, a consequence of that is that you have many components that are all kind of collaborating to perform a single calculation, and so any one of these components failing can actually p…

AI assessment note: “I think it's more of the complexity has grown, right?”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Can you, um, can you effectively, are you thinking of, like, fine-tuning a model against certain accounting terms, or doing, you know, like, I'm sort of curious anything about problem solving, or is it just wait for future generations of models to come out, or?

A We did spend some time fine-tuning in many places, and then we very quickly found out that our time was worth way more, uh, and that we should just, like, wait, wait for, like, other generations of model. What we've gotten really good at, though, is hardening our infrastructure so that we can easily switch when we need to, and we can quickly evaluate The different models on the, on the sub-tasks that we care about. Like, I was asked this question by, by, by one of our investors recently. It was like, with, like, the GPT, GPT-Forum Mini, how, how has that changed things for us? Has it brought costs down, and how are we thinking about it? And my answer is like, oh yeah, it's, it's already in production, and for like, 90% of tasks that we're running, it's good enough, so that was a quick switch. And within a day, we can know very quickly, like, that, yes, this is, like, good enough, and we have the right evals.

AI assessment note: “we should just, like, wait, wait for, like, other generations of model.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q original, um, very, you know, simple experience to something still simple, but much more powerful. Um, the company as well has just become much more enterprise facing in recent years. Um, like how did that evolution happen? Like what did you have to change most to support that? And when did you, when did you decide it was time to like go do that if there was a decision point?

A So that was always part of the, the master plan. And we actually wrote this like vision deck and, and kind of like business plan or as close to it as we got back in, when we, we started working on this, that laid this out. And we said, look, like generally it's probably harder to start with a really complicated product. Like you're not looking at SAP and saying, okay, over time, they're going to make it simpler. Whereas it is very common or at least, um, you know, more intuitive to start with a very, very simple product and then kind of make it more powerful and customizable, complex over time. Right. So, uh, and actually I, I, um, I think I got this, this terminology originally from Mike Krieger, but, um, you know, we, we like this idea of like, let's start with a really low floor, get the floor as low as possible. So we really are coming in and undercutting all the existing local ad platforms entirely. We're undercutting Salesforce service. Now we're undercutting, you know, like these old school products, like quick base and so on. And it's just going to be so much easier to use, but then over time we can improve The ceiling, right? Um, and initially we're going to get some like, you know, lightweight, medium weight use cases, but over time we want to improve the data scale. So, you know, actually literally just making it possible to score hundreds of thousands of, of, uh, ro…

AI assessment note: “So that was always part of the, the master plan.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Could we actually get into that? I'd love to hear sort of what you view as the consensus definition of AGI today. What's wrong about it? And then what do you think is the right way to measure or calibrate against that?

A Yeah. The sort of consensus definition that I think is most popular in sort of the AI industry right now is that AGI is a system that can do like the majority of economically useful work that humans can do. I, I think Vinod, uh, gets credit for joining this one. And, um, I, you know, I think it's a useful definition actually, uh, you know, look, I spend my day job building application and there is legitimate economic value that is sort of unlocked by the current regime with language models. Um, however, I don't think it's a good EGI definition though. Um, you know, I think it's a good definition of systems that are useful and economically useful, but, you know, I kind of joke that like, I think it says more about what many humans do for work than it does about actual general intelligence. And, uh, Francois definition, which is the one that I think is the right one is, uh, this definition that general intelligence is a system that can effectively, efficiently acquire new skill. That's, that's it efficiently acquiring new skill and being able to solve these open-ended problems with that ability. And here's sort of the simple, like maybe, um, argument in this line of thinking is, you know, we've had AI systems over the last 1015 years that can now, uh, you know, win at poker, uh, fold proteins, drive cars, win at chess. And yet I can't take any system that was like trained to beat…

AI assessment note: “The sort of consensus definition... is that AGI is a system that can do”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Is there anything you can share in terms of Um, adoption or metrics or usage by Zapier users or customers of, of your AI products?

A Yeah, we've got, um, at this point, over fifty million AI tasks have run on the platform to date over the last year and a half or so since we started tracking. So this is like, you know, think of a Zap, right? Where it's like you've got a trigger instead of actions where one of those actions is an AI step. Dominantly, this is open AI or a chat to PT step. Where, you know, users doing content generation or feature extraction or summarization, um, using AI in the middle of a workflow, uh, is, is kind of the dominant way people are adopting AI today. Um, over the last couple of months, we've introduced, uh, uh, other products in our AI space. So we're using AI basically across the entire product. We've, we launched a new product called Zapier, um, central, which are effectively these AI bots that, um, you don't have to build effectively. Uh, you know, the classic way I think most people experience Zapier is you have to build In the editor, right? You go have to, you know, do lots of configuration and click, click, click in order to get your zap set up and just tuned to the way you want. And one of the cool things with these new AI bots is you program the natural language. And we're not actually even doing natural language to structure mapping. It is a pure inference based engine, interpreting the user's instructions of what they want the bot to do and getting access to the, all th…

AI assessment note: “over fifty million AI tasks have run on the platform to date”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Um, I think, uh, you know, Sophia has an opportunity to be really, um, really impactful. Uh, you started a company last year taking leave from Stanford. Um, given your work has been like theoretical, but with practical applications, like what drove you to do that?

A I think I came to Stanford partly because there's a, um, very strong industry connection here at Stanford. Compared to some of the other universities. Um, and, and also probably entrepreneurship is just part of my, um, uh, my career plan, uh, anyways. And, uh, in terms of the timing, I felt that this is the right timing in the sense that, um, the, the technologies are more and more mature so that it seems that the commercialization is the, is the right timing right now. So for example, I think, um, uh, one, one story I have is that, you know, I, I look up some of my, Um, slide stack, uh, for my, uh, lectures at Stanford CS two and nine, seven years ago, uh, when I started to teach at Stanford. Um, at that point, machine learning, uh, we have a lecture with Chris Ray and the machine learning, uh, on applied machine learning. So how do you apply machine learning industry? And there are seven steps there. So, um, the first step is you define your problem. The second step is you collect your data, um, and you choose the loss function, you train it and you iterate so and so forth. So it's pretty complicated at that point. Um, and now the foundation model, um, uh, arrives to power and, uh, and in a new foundation model era, the only thing you have to do is that you have to, um, you know, someone will tune a foundation model for you, and then, uh, you tune a prompt and you add, uh, uh…

AI assessment note: “that's why I felt that this is probably the right time to commercialize”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q in a more general way. And so the application of, um, of AI in industry is just much, much cheaper, right? Cause you only do, you know, the last few steps and, or different set, but last few steps in essence. So maybe you can talk about like the, the, you know, just given wide range of research, the problem you focus on with voyage that you saw with customers.

A Yeah. So with, with Voyage, I think we are mostly building, uh, these two components, uh, Rerank and Embeddings for improving the quality of the retrieval or the search system. So the reason why we focus on this is because we talked to so many customers and we found that, uh, right now, um, uh, for implementing Rack, um, the bottleneck seems to be that, you know, it's, it's not very hard to implement it, right? You can just connect the components and have your Rack system ready very quickly. But the bottleneck seems to be the quality. Of the response and the quality of the response, uh, is heavily affected or is kind of almost, almost bottlenecked by the quality of the retrieval part. If the large language model, uh, see very relevant documents, then they can synthesize very good answers. Uh, even like a Lama can do that very well.

AI assessment note: “we talked to so many customers and we found that... the bottleneck seems to be”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q is obviously going to be needed. And it's the questions are, are to me are really like, you know, does, does efficiency matter both from a cost perspective and a speed, like a latency perspective, right? And how much can you push the context window? And like, you know, does hallucination management matter? And so I, I think there are lots of arguments for like RAG being very persistent here.

A Yeah, yeah, exactly. And just to add a little bit on that. So, uh, one million tokens, five books, right? So, but many companies has a hundred million tokens. That's a hundred X difference, right? So a hundred X, you know, for cost is a, is a big difference. That could be just, uh, um, you know, a hundred K dollars versus like ten million dollars, right? Ten million dollars is Unacceptable, but a hundred K sounds okay. Yeah. I think that's probably what's going to happen. Like, so, so from, at least for many of the companies, right? So right now, if they have a hundred million tokens, I don't think they can use long context transformers at all because it's way too expensive.

AI assessment note: “a hundred X, you know, for cost is a, is a big difference.”

← previous page 3 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.